Runs onWeb Scraping API

HTML Extraction API

Get the fully rendered, post-JavaScript HTML of any page, the DOM a browser would see and feed it straight into the parser you already have.

4.9average across
G2 — 4.8 out of 5Google — 4.9 out of 5Product Hunt — 5.0 out of 5
REQUEST PAYLOAD
{
  "url": "example.com",
  "format": "html"
}
Output
Rendered HTML
Cost
1 credit a page
Best for
Teams with their own parser
  • 500 free credits every month
  • No card to start
  • Failed requests cost nothing
Playground
format
Request
POST /webscraping
{
  "url": "stripe.com",
  "format": [
    "html"
  ]
}
loading example…
Want every endpoint and every parameter?
The Geekflare playground has every API and option, plus saved requests. Free with 500 credits a month.
Open the playground
Trusted by teams at
PfizerNBCUniversalTCSHostingerKissflowLookoutPlivoClearSaleSparkianCBSplitOmreon
PfizerNBCUniversalTCSHostingerKissflowLookoutPlivoClearSaleSparkianCBSplitOmreon

Plain requests miss the content

  • curl and fetch return pre-JavaScript HTML — modern sites render it nearly empty
  • Anti-bot walls and CAPTCHAs block naive scrapers
  • You run and scale headless browsers and proxy pools yourself

The fully rendered DOM

Post-render, not the first response
  • Headless Chrome runs the page first, so you get the HTML a browser sees
  • Anti-bot handling and CAPTCHA solving included on every request
  • Feed the response straight into your own parser — BeautifulSoup, Cheerio, whatever you run

There is no separate endpoint to learn. Set format: "html" — or "html-llm" — on the Web Scraping API.

Unblock your pipeline

Your selectors, our browser fleet.

Keep the extraction rules you have already written and tested. What changes is who runs the browser, the proxies and the retries.

Post-render DOM

The page is rendered before it is returned, so client-side content is present rather than an empty root div.

renderJS: true

Anti-bot handled

Bot-detection challenges and CAPTCHAs are dealt with as part of the request, not as your problem.

stealth

Proxy rotation built in

Requests can route through a managed pool, so you are not sourcing and rotating IPs yourself.

proxyMode: "auto"

Complete page HTML

The whole document comes back, so existing CSS or XPath selectors keep working unchanged.

data.html

Trimmed variant

html-llm strips scripts, trackers and boilerplate when you want the structure without the noise.

format: "html-llm"

Several formats at once

Ask for HTML and Markdown in one request when part of your pipeline wants each, for the same single credit.

format: ["html", "markdown"]
Use cases

How teams run rendered HTML in production.

Swap the fetch layer, keep every selector

Most scraping rewrites fail on the extraction rules, not the fetching. Here the rules stay: point the parser you already have at this response instead of at your own browser pool, and the selectors resolve against the same DOM they always did.

  • No extraction rules to port
  • Retire the browser fleet and the proxy bill in one step
  • Roll over one target at a time
Typical request
POST /webscraping
{
  "url": "https://example.com/listing",
  "format": ["html"],
  "renderJS": true
}
What comes back
data.htmlrendered DOM
your selectorsunchanged
Quickstart

One request, rendered HTML back.

quickstart.pyofficial SDK
# pip install geekflare-api
from geekflare_api.client import GeekflareClient
from geekflare_api.models import WebScrapeDto

with GeekflareClient(api_key="<api-key>") as client:
    result = client.web_scrape(
        WebScrapeDto(
            url="https://example.com",
            format=["markdown"]
        )
    )
    print(result)
Pricing

One credit a page, whichever format.

Rendered HTML costs the same as Markdown or text, and credits are shared with every other Geekflare API. Failed requests cost nothing.

5002M

Growth covers up to ~100,000 pages a month (100K credits), about $0.69 per 1,000 pages. Or $58/mo billed yearly.

pages
60,000
Plan
Growth · $69/mo
FreeNo card
$0/mo
~500
pages / month
  • 500 credits / month
  • 1 team member
  • 7 days log retention
  • 1 request per second
Starter
$19/mo
~10,000
pages / month
  • 10K credits / month
  • 3 team members
  • 30 days log retention
  • 5 requests per second
GrowthMost popular
$69/mo
~100,000
pages / month
  • 100K credits / month
  • 5 team members
  • 30 days log retention
  • 10 requests per second
Business
$349/mo
~1,000,000
pages / month
  • 1M credits / month
  • 25 team members
  • 90 days log retention
  • 25 requests per second
Start freeCompare plans and credit packsEvery account starts free. Upgrade from the dashboard when you need more, or buy a credit pack from $10.
In production

Teams that stopped maintaining scrapers.

Read all reviews
“Found Geekflare API to get markdown from URL for my AI agents. It is fast and cheaper and works on almost every website.”
Ram DasiArchitect, PA Consulting
“After trying many scraping services, I selected Geekflare to scrape public directories. Mainly for two reasons - it is cheaper and fast.”
Vishu SharmaCTO
“Documentation was easy to follow. We were scraping dynamic pages within hours. Very reliable service.”
Marco SilvaAnalytics Engineer
Questions

Rendered HTML, specifically.

General scraping questions are answered on the Web Scraping API page.

Open the Web Scraping API page

After. The page is loaded in headless Chrome and the response is the rendered DOM, which is what your selectors were written against.

html returns the complete rendered document. html-llm strips scripts, trackers and boilerplate, keeping the structure but cutting the noise.

Yes, that's the point of this one. You get the markup and nothing is imposed on how you extract from it — BeautifulSoup, Cheerio, lxml or your own code.

Bot-detection challenges are handled as part of the request. For the hardest targets, pair it with proxyMode: "auto" so a blocked attempt is retried through a proxy, and stealth for stricter fingerprinting.

A page costs 1 credit. Routing through a proxy adds credits only on the requests that use one. Failed requests cost nothing.

Use HTML when you have a parser and want control. Use Markdown when the destination is a model, because it carries the same structure at roughly 60% fewer tokens.

Rendered HTML, on your first request.

Create a free account to try any URL in the playground and grab your API key. 500 credits a month, no card.