XPath Tester
Test XPath expressions on HTML, XML or a live page and generate scraping code.
Click an example to use it
Result
from lxml import html
with open("page.html", "rb") as f:
tree = html.fromstring(f.read())
for result in tree.xpath("//a[@class='card']//h3"):
# Elements have text; attribute values and text() come back as strings.
print(result if isinstance(result, str) else "".join(result.itertext()).strip())What Is an XPath Tester?
XPath is a query language for picking nodes out of an HTML or XML document: elements, attributes, text, or values computed from them. It's what lxml, Scrapy, Selenium and Playwright use to locate data on a page. An XPath tester lets you try an expression and see exactly what it selects before you put it in a scraper, so you find a wrong path in seconds rather than after a run comes back empty.
How to Use It
- Type an XPath expression. Results update as you type.
- Choose what to test it against:
- Sample: a product page, blog article, data table or RSS feed, ready to experiment with.
- Paste HTML / XML: your own markup. It's parsed in your browser.
- Load a URL: test in a browser through the Web Scraping API, so content added by JavaScript is there to match.
- Check the result: every matching node with its text or value and its absolute path, or a single value for expressions such as
count(). - Copy the generated code in Python, JavaScript, or as a ready-to-run Geekflare request.
XPath Cheat Sheet for Web Scraping
| Expression | Selects |
|---|---|
//h1 | Every <h1> in the document |
//a/@href | The href of every link |
//h1/text() | The text directly inside each <h1> |
//div[@id='price'] | The <div> whose id is exactly price |
//*[contains(@class, 'price')] | Any element with price in its class |
//li[1] | The first <li> in each list |
//tbody/tr[last()] | The last row of a table body |
//tr[td[1]='Growth']/td[2] | The second cell of the row whose first cell is "Growth" |
//h2[1]/following-sibling::p[1] | The paragraph right after the first <h2> |
//a[starts-with(@href, 'http')] | Links to absolute URLs |
//p[normalize-space()='In stock'] | A paragraph by its exact text, ignoring stray whitespace |
count(//a) | The number of links, as a number |
XPath vs CSS Selectors
CSS selectors are shorter and cover most scraping jobs. Reach for XPath when you need to:
- Select by text, such as the row whose first cell says "Growth". CSS can't match on text content.
- Move up or sideways, with
..,ancestor::orfollowing-sibling::. CSS only moves down and forward. - Return attributes or values directly, such as
//a/@hreforcount(//item). - Query XML, such as feeds, sitemaps and API responses, where XPath is the native tool.
If CSS is enough, try it in the CSS Selector Tester.
Writing XPaths That Don't Break
/html/body/div[3]/div[2]/spanbreaks as soon as a banner or wrapper is added. Anchor on something stable instead.@id,@data-testidand@itempropchange far less often than layout classes.- Match classes with
contains(),@class='price'fails when the element also hasprice sale. Usecontains(@class, 'price'), and check it doesn't also catchprice-old. - Use
normalize-space()for text. Markup is full of line breaks and indentation that make exact text comparisons fail. - Anchor on labels.
//th[.='SKU']/following-sibling::tdkeeps working when the table gains or loses rows.
Test Against the Rendered Page
Many sites build their content with JavaScript, so the HTML your HTTP client downloads isn't what you see in the browser. An XPath copied from DevTools can match in the browser and return nothing in your scraper because that node never existed in the raw response. Loading a URL here renders the page first, so you're testing against the same DOM a browser builds.
From XPath to a Working Scraper
Once the expression returns what you need, you can run it:
- In your own code, with the Python (
lxml) or JavaScript (document.evaluate) snippet the tester generates. - Through the Web Scraping API, which fetches and renders the page for you and returns each match as JSON. We manage JavaScript rendering and anti-bot handling.
{
"url": "https://example.com/product/1",
"format": ["json"],
"extractionMode": "xpathSchema",
"extractionSchema": {
"name": "Product",
"baseSelector": "//main",
"fields": [
{ "name": "title", "selector": ".//h1", "type": "text" },
{ "name": "price", "selector": ".//*[contains(@class, 'price')]", "type": "text" }
]
}
}baseSelector picks the repeating block and each field is an XPath relative to it, so a listing page comes back as an array of clean objects. See the Web Scraping API reference for every option.
Frequently Asked Questions
XPath 1.0, evaluated by your browser's own engine. That's the same version lxml, Scrapy, Selenium, Playwright and document.evaluate use, so an expression that works here works in your scraper.
Yes. Choose Load a URL and the page is rendered in a browser through the Geekflare Web Scraping API, so content added by JavaScript is there to match.
The usual cause is a default namespace, such as xmlns="http://www.w3.org/2005/Atom" on the root. In XPath 1.0 an unprefixed name never matches a namespaced element, so //entry finds nothing. Use //*[local-name()='entry'] instead.
End the expression with /@name for an attribute, for example //a/@href, or /text() for an element's own text, for example //h1/text(). Without either, you get the element itself and the tester shows its full text.
No. Pasted HTML and XML are parsed and evaluated in your browser. Only Load a URL sends a request, and that request carries just the URL to fetch.
Yes. The Geekflare API tab turns your expression into a Web Scraping API request with extractionMode: "xpathSchema". Each matching node comes back as an object in JSON, and JavaScript-heavy pages are rendered before your XPath runs.
Need this at scale? Try the Web Scraping API.
Run your XPath on any page via API. Pass an xpathSchema and get every match back as JSON, after the page is rendered.
curl -X POST https://api.geekflare.com/webscraping \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","format":["json"],"extractionMode":"xpathSchema","extractionSchema":{"name":"Extracted","baseSelector":"//h1","fields":[{"name":"text","selector":".","type":"text"}]}}'