HTML Extraction API
Get the fully rendered, post-JavaScript HTML of any page, the DOM a browser would see and feed it straight into the parser you already have.
{
"url": "example.com",
"format": "html"
}- Output
- Rendered HTML
- Cost
- 1 credit a page
- Best for
- Teams with their own parser
- 500 free credits every month
- No card to start
- Failed requests cost nothing
POST /webscraping
{
"url": "stripe.com",
"format": [
"html"
]
}





Plain requests miss the content
curlandfetchreturn pre-JavaScript HTML — modern sites render it nearly empty- Anti-bot walls and CAPTCHAs block naive scrapers
- You run and scale headless browsers and proxy pools yourself
The fully rendered DOM
Post-render, not the first response- Headless Chrome runs the page first, so you get the HTML a browser sees
- Anti-bot handling and CAPTCHA solving included on every request
- Feed the response straight into your own parser — BeautifulSoup, Cheerio, whatever you run
There is no separate endpoint to learn. Set format: "html" — or "html-llm" — on the Web Scraping API.
Your selectors, our browser fleet.
Keep the extraction rules you have already written and tested. What changes is who runs the browser, the proxies and the retries.
Post-render DOM
The page is rendered before it is returned, so client-side content is present rather than an empty root div.
renderJS: trueAnti-bot handled
Bot-detection challenges and CAPTCHAs are dealt with as part of the request, not as your problem.
stealthProxy rotation built in
Requests can route through a managed pool, so you are not sourcing and rotating IPs yourself.
proxyMode: "auto"Complete page HTML
The whole document comes back, so existing CSS or XPath selectors keep working unchanged.
data.htmlTrimmed variant
html-llm strips scripts, trackers and boilerplate when you want the structure without the noise.
format: "html-llm"Several formats at once
Ask for HTML and Markdown in one request when part of your pipeline wants each, for the same single credit.
format: ["html", "markdown"]How teams run rendered HTML in production.
Swap the fetch layer, keep every selector
Most scraping rewrites fail on the extraction rules, not the fetching. Here the rules stay: point the parser you already have at this response instead of at your own browser pool, and the selectors resolve against the same DOM they always did.
- No extraction rules to port
- Retire the browser fleet and the proxy bill in one step
- Roll over one target at a time
POST /webscraping
{
"url": "https://example.com/listing",
"format": ["html"],
"renderJS": true
}One request, rendered HTML back.
# pip install geekflare-api
from geekflare_api.client import GeekflareClient
from geekflare_api.models import WebScrapeDto
with GeekflareClient(api_key="<api-key>") as client:
result = client.web_scrape(
WebScrapeDto(
url="https://example.com",
format=["markdown"]
)
)
print(result)One credit a page, whichever format.
Rendered HTML costs the same as Markdown or text, and credits are shared with every other Geekflare API. Failed requests cost nothing.
Growth covers up to ~100,000 pages a month (100K credits), about $0.69 per 1,000 pages. Or $58/mo billed yearly.
- 500 credits / month
- 1 team member
- 7 days log retention
- 1 request per second
- 10K credits / month
- 3 team members
- 30 days log retention
- 5 requests per second
- 100K credits / month
- 5 team members
- 30 days log retention
- 10 requests per second
- 1M credits / month
- 25 team members
- 90 days log retention
- 25 requests per second
Teams that stopped maintaining scrapers.
“Found Geekflare API to get markdown from URL for my AI agents. It is fast and cheaper and works on almost every website.”
Ram DasiArchitect, PA Consulting“After trying many scraping services, I selected Geekflare to scrape public directories. Mainly for two reasons - it is cheaper and fast.”
“Documentation was easy to follow. We were scraping dynamic pages within hours. Very reliable service.”
Rendered HTML, specifically.
General scraping questions are answered on the Web Scraping API page.
Open the Web Scraping API pagehtml returns the complete rendered document. html-llm strips scripts, trackers and boilerplate, keeping the structure but cutting the noise.proxyMode: "auto" so a blocked attempt is retried through a proxy, and stealth for stricter fingerprinting.Rendered HTML, on your first request.
Create a free account to try any URL in the playground and grab your API key. 500 credits a month, no card.