Tutorials

How to Extract Structured JSON with Firecrawl and Python

Ten lines of Python, real output from firecrawl-py 4.50.0, and the one check that stops a 404 page becoming a record

Ryan Mitchell11 min read
Share
FIRECRAWL TO JSON: on a dark charcoal background with an orange glow, cluttered website cards pass through a glowing flame machine beside a small code window and come out as a stack of four clean data cards, the top one with a check mark.

You point a scraper at a product page, ask for JSON, and back comes a tidy record: title, price, stock status. It passes your schema. It goes into the database.

The page was a 404. We know because it happened in our own test run, and Firecrawl billed it at the full five credits.

Getting structured JSON out of a website with Firecrawl and Python takes about ten lines. Getting JSON you can trust takes one more check, and it isn't the one the docs lead with. It comes right after a listing-page trick that cuts the bill by 95%.

We ran each scrape example on 11 October 2026 with firecrawl-py 4.50.0, on Firecrawl's keyless free tier, against books.toscrape.com (a practice site built for scrapers). Outputs are pasted as printed; prices and limits come from Firecrawl's pricing page and docs, checked the same day. The batch example needs an API key, so it wasn't run.

What you'll have at the end, and what each page costs

A function that takes a URL and hands back a validated Python object, or a clear reason it skipped the page. Ours returns a Book with a title, a price and a stock count. Swap the model and it returns whatever your pages hold.

Every page costs 5 credits: one for the scrape, four more for JSON format, per Firecrawl's billing docs. On Hobby, Firecrawl's $19-a-month plan with 5,000 credits, that's 1,000 extracted pages, or about 2 cents each. The same plan would scrape 5,000 pages as plain Markdown.

That multiplier is why the rest of this guide cares about wasted calls. Our Firecrawl review covers the other endpoints and plans; here we stay on the Python side.

You might be thinking this test site is ten lines of free BeautifulSoup. It is, and for one static site you know well, write the parser. Firecrawl earns its 5 credits when you have fifty sites with fifty layouts and one schema to fill.

Install firecrawl-py and get a key (the keyless shortcut fails in Python)

The SDK is firecrawl-py on PyPI. Version 4.50.0 was the latest when we checked, it needs Python 3.8 or newer, and it pulls in Pydantic 2. It moves fast: six releases landed between 6 and 8 October, so pin the version you tested.

PyPI page for firecrawl-py 4.50.0 with the install command pip install firecrawl-py and Key dates showing Released Oct 9, 2026, latest release
The SDK we ran. PyPI shows the release date in your own time zone (20:35 UTC on 8 October). Source: pypi.org/project/firecrawl-py, captured 11 October 2026.

Firecrawl's docs say you can skip the key and scrape on a free keyless tier, capped per IP address. The API does allow that. The Python client doesn't.

In 4.50.0, Firecrawl() with no key stops with ValueError: No API key provided, because an older v1 client bundled inside it still demands one.

So get the free key. It comes with 1,000 credits a month and no card, which is 200 JSON pages. Then:

pip install "firecrawl-py==4.50.0" pydantic
export FIRECRAWL_API_KEY="fc-your-key"
import os
from typing import List, Optional

from firecrawl import Firecrawl
from pydantic import BaseModel, Field

app = Firecrawl(api_key=os.environ["FIRECRAWL_API_KEY"])

Describe the fields once, as a Pydantic model

The schema is the whole trick. Firecrawl sends the page to a language model along with your field names, types and descriptions, and the model fills them in.

class Book(BaseModel):
    title: str
    price: str = Field(description="Price exactly as shown, with currency symbol")
    in_stock: bool
    stock_count: Optional[int] = Field(None, description="Copies available. Return null if not shown.")
    rating: Optional[int] = Field(None, description="Star rating, 1 to 5. Return null if not shown.")

You can pass the class straight into a JSON format. The SDK calls model_json_schema() for you, inlines nested models, and checks the result before sending it. The docs' Book.model_json_schema() style works too.

Treat each description as a short instruction. "Return null if not shown" comes from Firecrawl's own tips, and it's meant to stop the model guessing. Remember it when we get to the rating.

One page, one call, five credits

JSON mode is a format on the ordinary scrape call. The schema goes inside the format object; the old jsonOptions parameter from v1 no longer exists, which is why older tutorials break.

URL = "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"

doc = app.scrape(URL, formats=[{"type": "json", "schema": Book}])

print(doc.json)
print(doc.metadata.status_code, doc.metadata.credits_used)

It printed:

{'title': 'A Light in the Attic', 'price': '£51.77', 'in_stock': True, 'stock_count': 22, 'rating': None}
200 5

Title, price and stock count match the page. The call took about a second, and credits_used confirms the 5-credit price. Notice that doc.json is a plain dict, not a Book: the SDK converts your model going out but not coming back.

Now look at the rating. The page shows three stars. We got None.

Why did the rating come back empty?

Because the stars aren't text. In the page's HTML, the rating lives only in a CSS class: <p class="star-rating Three">.

JSON mode never sees that HTML. It works from the Markdown version of the page, and Firecrawl's JSON mode docs say so plainly.

Firecrawl JSON mode docs: HTML attributes are not available in JSON extraction because it works on the markdown conversion of the page, plus tips including adding Return null if not found to each field description
The note that explains the missing rating, above the tip we followed. Source: docs.firecrawl.dev/features/llm-extract, captured 11 October 2026.

Null is the honest answer, but not the only one you'll get. Twelve minutes earlier, with an almost identical description, the same field came back as 0.

Zero is worse. It's a valid integer, so it slides straight through validation.

There are two fixes.

Give numeric fields a range, such as Field(None, ge=1, le=5), so a made-up zero fails Pydantic instead of reaching your database. For values stored in attributes, the docs suggest two routes. Ask for the rawHtml format and parse it yourself, or add an executeJavascript action that copies the value into visible text.

Prompt-only extraction: quicker to write, harder to build on

You can skip the schema and describe what you want in a sentence. Firecrawl calls this extraction without a schema, and "the llm chooses the structure of the data."

doc = app.scrape(URL, formats=[{
    "type": "json",
    "prompt": "Extract the book's title, price and how many copies are in stock.",
}])
print(doc.json)

It printed:

{'title': 'A Light in the Attic', 'price': 51.77, 'stock': 22}

The answers are right. The shape isn't yours: the price lost its pound sign and became a number, and the stock count arrived under a key called stock.

No code of ours expected that key. Ask again next week and it might be copies.

You're probably thinking a stricter prompt would fix that. Firecrawl's docs say the opposite: long prompts with many rules "increase variability", so move constraints into the schema.

Use prompt-only to explore a page you've never seen. It costs the same 5 credits as a schema, so don't build a pipeline on it.

A list page needs an array, and it's the cheap way in

Most sites show the fields you want twice: once on each product page and once on the listing. A schema with a list in it pulls the whole listing in one call.

class BookCard(BaseModel):
    title: str
    price: str

class Listing(BaseModel):
    books: List[BookCard]

doc = app.scrape("https://books.toscrape.com/", formats=[{"type": "json", "schema": Listing}])
listing = Listing.model_validate(doc.json)
print(len(listing.books), listing.books[0])

It printed:

20 title='A Light in the Attic' price='£51.77'

We checked all 20 records against the page's HTML: every title and price matched exactly. Wrap items in a list like this, because Firecrawl's docs warn that an object where a list belongs returns a single item. Skip minItems too, which the docs say can make the model invent entries to hit the count.

Now the arithmetic. Twenty books for 5 credits is a quarter of a credit per record, while the 20 detail pages would cost 100 credits.

If the listing already has the fields you need, extract the listing.

The check that catches a 404 dressed up as data

Here's the record from the opening. We asked for a book that doesn't exist, using the same schema.

MISSING = "https://books.toscrape.com/catalogue/this-book-does-not-exist_9999/index.html"

doc = app.scrape(MISSING, formats=[{"type": "json", "schema": Book}])

print(doc.metadata.status_code, doc.metadata.credits_used)
print(doc.json)

It printed:

404 5
{'title': '404 Not Found', 'price': 'N/A', 'in_stock': False, 'stock_count': None, 'rating': None}

The site said 404, Firecrawl charged 5 credits, and the model still filled in a book called "404 Not Found". Every field has the right type, so Pydantic accepts it. Our earlier run of the same URL priced it at "/", which is just as valid as a string.

Firecrawl is open about this. A page that answers with an error status still comes back as a document, and it still costs credits.

The pricing page says 1 credit, which is the plain-scrape figure. The billing docs add that format costs stack on top, which is how our 404 cost 5.

Firecrawl pricing at a glance: 1 credit per page, JSON, Question and Highlight formats add 4 credits per page, free tier 1,000 credits a month, Hobby 19 dollars monthly or 16 annually, and error pages such as 403 or 404 still cost 1 credit
The credit rules, including the line on failed requests. Source: firecrawl.dev/pricing, captured 11 October 2026.

The fix is one line: check doc.metadata.status_code before you trust anything in doc.json. Here it is inside the function we promised at the start, with validation and rate limits handled.

import time
from firecrawl import FirecrawlError, RateLimitError
from pydantic import ValidationError

def get_book(url: str, attempts: int = 4) -> Optional[Book]:
    for attempt in range(attempts):
        try:
            doc = app.scrape(url, formats=[{"type": "json", "schema": Book}])
            break
        except RateLimitError as e:  # 429: the SDK doesn't retry these
            wait = e.response.headers.get("Retry-After") if e.response is not None else None
            time.sleep(float(wait) if wait else 2 ** attempt)
        except FirecrawlError as e:  # 401, 402, 403, 422...
            print(f"skip {url}: {e}")
            return None
    else:
        print(f"skip {url}: still rate-limited")
        return None

    if doc.metadata.status_code != 200:
        print(f"skip {url}: page returned {doc.metadata.status_code}")
        return None
    try:
        return Book.model_validate(doc.json)
    except ValidationError as e:
        print(f"skip {url}: {e.error_count()} field(s) failed validation")
        return None

print(get_book("https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html"))
print(get_book(MISSING))

It printed:

title='Tipping the Velvet' price='£53.74' in_stock=True stock_count=20 rating=None
skip https://books.toscrape.com/catalogue/this-book-does-not-exist_9999/index.html: page returned 404
None

Why handle 429s yourself? The client's max_retries=3 sounds like it covers them. It doesn't.

In the SDK source, the client retries only 502 errors and dropped connections. A rate-limited request raises RateLimitError at once, and Firecrawl's errors page says to wait for the Retry-After header. A 402 means you're out of credits, so the function skips it rather than retrying.

Scaling up with batch scrape (watch the timeout)

For more than a handful of URLs, batch_scrape takes a list and the same JSON format. It needs a real API key, so this is the one example we couldn't run.

urls = [
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
    "https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html",
    "https://books.toscrape.com/catalogue/soumission_998/index.html",
]

job = app.batch_scrape(
    urls,
    formats=[{"type": "json", "schema": Book}],
    poll_interval=2,
    wait_timeout=300,  # seconds to wait for the whole job
)

books = [Book.model_validate(d.json) for d in job.data if d.metadata.status_code == 200]
print(job.status, job.completed, job.total, job.credits_used, len(books))

Two parameters look alike and aren't. wait_timeout is how many seconds the SDK waits for the whole job. timeout is the per-page scrape limit in milliseconds, so timeout=300 would give each page three-tenths of a second.

Your plan sets the pace, per the rate-limits page. Batch jobs share the crawl limit: 2 batch requests a minute on Free, 20 on Hobby. Pages then render on your concurrent browsers, 2 on Free, 5 on Hobby and 25 on Standard.

Credits are charged as each page finishes, and results stay on the API for 24 hours.

One more default to know about. The Python SDK asks for cached pages up to four hours old unless you pass max_age, and cached pages still cost credits. For prices that change hourly, pass max_age=0 and accept slower, slightly flakier scrapes.

Scrape JSON mode, /extract or Agent: which one?

Firecrawl has three ways to get structured data, and its own docs line them up.

Firecrawl docs comparison of /agent, /extract and /scrape JSON mode: /extract marked use /agent instead with token-based pricing at 1 credit per 15 tokens, /agent dynamic pricing with 5 free runs a day, and /scrape JSON mode active at 5 credits per page for known single-page extraction
Firecrawl's own comparison. Source: docs.firecrawl.dev, Choosing the Data Extractor, captured 11 October 2026.

If you know the URLs, use scrape with JSON format, singly or in a batch. You can predict the bill to the credit.

/extract takes many URLs and wildcards, but Firecrawl now marks it "Use /agent instead". It's priced in tokens (1 credit per 15), so you can't know a page's cost in advance. Agent finds pages for you, with five free runs a day and dynamic pricing after that; cap it with max_credits.

There's a fourth option for shops. The product format reads a product from the page's structured data with no language model, for 1 credit.

It refuses rather than guesses. Our test site has no such markup, so we got no product, a warning and a 1-credit charge. On real stores with product markup, try it before paying for JSON.

What 1,000 JSON pages cost on each plan

At 5 credits a page, each plan covers a fifth as many pages as its headline credit count suggests. Here's what that does to the price per 1,000 extracted pages, using the pricing page figures:

PlanPrice (monthly / annual)JSON pages a monthPer 1,000 JSON pages
Free$0200$0
Hobby$19 / $161,000$19.00 / $16.00
Standard$99 / $8320,000$4.95 / $4.15
Growth$399 / $333100,000$3.99 / $3.33
Scale$749 / $599200,000$3.75 / $3.00

Run past your plan and pay-as-you-go tops you up in $5 steps (paid plans only). On Hobby, $5 buys 1,000 credits, so overage JSON pages cost $25 per 1,000; on Standard, $12.50.

Unused credits don't roll over, except for one month on annual Scale.

Say you track 100 product pages a day, about 3,000 a month. As JSON that's 15,000 credits: Hobby plus $50 of top-ups ($69), or Standard at $99 ($83 billed yearly).

Pull the same data from listing pages, or with the product format where it works, and it fits inside Hobby's 5,000 credits for $19. The format you pick can decide your plan.

Our Firecrawl vs Playwright comparison prices the do-it-yourself route, if these numbers make you want to run your own browser. Our Apify vs Firecrawl breakdown covers ready-made scrapers for big sites.

So which setup should you build?

If you have a few known pages and the fields are visible text, copy get_book: a Pydantic schema, a status check, then validation. Start there.

If you have thousands of pages, check the listing first. On a 20-item listing, one JSON call does the work of 20. Then batch the rest with wait_timeout set.

If you're scraping shops, try product before JSON. If the data hides in attributes, take rawHtml and parse it yourself, because no prompt will make the model see it. And if you don't have URLs at all, that's an Agent job with a credit cap, not this one.

Frequently asked questions

How many credits does Firecrawl's JSON mode cost?

5 credits per page: 1 for the scrape and 4 for the JSON format. The free plan's 1,000 monthly credits cover 200 JSON pages, and Hobby's 5,000 credits ($19 a month, or $16 billed yearly) cover 1,000.

Can I use Firecrawl's Python SDK without an API key?

Not with the main client. The API has a keyless free tier for scrape, search and interact, but in firecrawl-py 4.50.0 (checked 11 October 2026) Firecrawl() raises ValueError: No API key provided. A free key gives you 1,000 credits a month with no card.

Should I use /extract or scrape with the JSON format?

Use scrape with JSON format when you know the URLs: it costs a predictable 5 credits a page. Firecrawl marks /extract as "Use /agent instead" and prices it in tokens (1 credit per 15), so a page's cost isn't known in advance.

Why does Firecrawl return null or 0 for a field I can see on the page?

JSON mode reads the page's Markdown, which keeps only visible text. Values held in HTML attributes, like a star rating stored in a CSS class, never reach the model. Request rawHtml and parse those fields yourself.

Does Firecrawl charge for pages that return 404?

Yes. A page that answers with an error status still comes back as a document, so it costs 1 credit plus any format costs. Our 404 with JSON format was billed 5 credits. A scrape that returns no document at all costs nothing.

Does the Firecrawl Python SDK retry rate-limited requests?

No. Its max_retries setting (3 by default) covers 502 errors and dropped connections only. A 429 raises RateLimitError straight away; wait for the Retry-After header, then try again.

Turn this into your next generation

Put these techniques to work across 7+ image and video models — free to start.

Start creating free
Firecrawl Review: a cluttered browser window with an orange stream flowing out of it into clean documents and a data card, surrounded by search, sitemap, crawl and click icons
Review8.6/10

Firecrawl Review (2026): Features, API & Pricing

Your AI agent needs to read the web. Firecrawl turns any page into clean Markdown in one call, but one setting decides if a page costs 0.1 cents or 1.6.

Ryan Mitchell7 min read