XPath Tester

Test XPath expressions on HTML, XML or a live page and generate scraping code.

Click an example to use it

Result

python
from lxml import html

with open("page.html", "rb") as f:
    tree = html.fromstring(f.read())

for result in tree.xpath("//a[@class='card']//h3"):
    # Elements have text; attribute values and text() come back as strings.
    print(result if isinstance(result, str) else "".join(result.itertext()).strip())

What Is an XPath Tester?

XPath is a query language for picking nodes out of an HTML or XML document: elements, attributes, text, or values computed from them. It's what lxml, Scrapy, Selenium and Playwright use to locate data on a page. An XPath tester lets you try an expression and see exactly what it selects before you put it in a scraper, so you find a wrong path in seconds rather than after a run comes back empty.

How to Use It

  1. Type an XPath expression. Results update as you type.
  2. Choose what to test it against:
    • Sample: a product page, blog article, data table or RSS feed, ready to experiment with.
    • Paste HTML / XML: your own markup. It's parsed in your browser.
    • Load a URL: test in a browser through the Web Scraping API, so content added by JavaScript is there to match.
  3. Check the result: every matching node with its text or value and its absolute path, or a single value for expressions such as count().
  4. Copy the generated code in Python, JavaScript, or as a ready-to-run Geekflare request.

XPath Cheat Sheet for Web Scraping

ExpressionSelects
//h1Every <h1> in the document
//a/@hrefThe href of every link
//h1/text()The text directly inside each <h1>
//div[@id='price']The <div> whose id is exactly price
//*[contains(@class, 'price')]Any element with price in its class
//li[1]The first <li> in each list
//tbody/tr[last()]The last row of a table body
//tr[td[1]='Growth']/td[2]The second cell of the row whose first cell is "Growth"
//h2[1]/following-sibling::p[1]The paragraph right after the first <h2>
//a[starts-with(@href, 'http')]Links to absolute URLs
//p[normalize-space()='In stock']A paragraph by its exact text, ignoring stray whitespace
count(//a)The number of links, as a number

XPath vs CSS Selectors

CSS selectors are shorter and cover most scraping jobs. Reach for XPath when you need to:

  • Select by text, such as the row whose first cell says "Growth". CSS can't match on text content.
  • Move up or sideways, with .., ancestor:: or following-sibling::. CSS only moves down and forward.
  • Return attributes or values directly, such as //a/@href or count(//item).
  • Query XML, such as feeds, sitemaps and API responses, where XPath is the native tool.

If CSS is enough, try it in the CSS Selector Tester.

Writing XPaths That Don't Break

  • /html/body/div[3]/div[2]/span breaks as soon as a banner or wrapper is added. Anchor on something stable instead.
  • @id, @data-testid and @itemprop change far less often than layout classes.
  • Match classes with contains(), @class='price' fails when the element also has price sale. Use contains(@class, 'price'), and check it doesn't also catch price-old.
  • Use normalize-space() for text. Markup is full of line breaks and indentation that make exact text comparisons fail.
  • Anchor on labels. //th[.='SKU']/following-sibling::td keeps working when the table gains or loses rows.

Test Against the Rendered Page

Many sites build their content with JavaScript, so the HTML your HTTP client downloads isn't what you see in the browser. An XPath copied from DevTools can match in the browser and return nothing in your scraper because that node never existed in the raw response. Loading a URL here renders the page first, so you're testing against the same DOM a browser builds.

From XPath to a Working Scraper

Once the expression returns what you need, you can run it:

  • In your own code, with the Python (lxml) or JavaScript (document.evaluate) snippet the tester generates.
  • Through the Web Scraping API, which fetches and renders the page for you and returns each match as JSON. We manage JavaScript rendering and anti-bot handling.
json
{
  "url": "https://example.com/product/1",
  "format": ["json"],
  "extractionMode": "xpathSchema",
  "extractionSchema": {
    "name": "Product",
    "baseSelector": "//main",
    "fields": [
      { "name": "title", "selector": ".//h1", "type": "text" },
      { "name": "price", "selector": ".//*[contains(@class, 'price')]", "type": "text" }
    ]
  }
}

baseSelector picks the repeating block and each field is an XPath relative to it, so a listing page comes back as an array of clean objects. See the Web Scraping API reference for every option.

Frequently Asked Questions

XPath 1.0, evaluated by your browser's own engine. That's the same version lxml, Scrapy, Selenium, Playwright and document.evaluate use, so an expression that works here works in your scraper.

Yes. Choose Load a URL and the page is rendered in a browser through the Geekflare Web Scraping API, so content added by JavaScript is there to match.

The usual cause is a default namespace, such as xmlns="http://www.w3.org/2005/Atom" on the root. In XPath 1.0 an unprefixed name never matches a namespaced element, so //entry finds nothing. Use //*[local-name()='entry'] instead.

End the expression with /@name for an attribute, for example //a/@href, or /text() for an element's own text, for example //h1/text(). Without either, you get the element itself and the tester shows its full text.

No. Pasted HTML and XML are parsed and evaluated in your browser. Only Load a URL sends a request, and that request carries just the URL to fetch.

Yes. The Geekflare API tab turns your expression into a Web Scraping API request with extractionMode: "xpathSchema". Each matching node comes back as an object in JSON, and JavaScript-heavy pages are rendered before your XPath runs.

More Developers tools
Free to start · No card

Need this at scale? Try the Web Scraping API.

Run your XPath on any page via API. Pass an xpathSchema and get every match back as JSON, after the page is rendered.

bash
curl -X POST https://api.geekflare.com/webscraping \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","format":["json"],"extractionMode":"xpathSchema","extractionSchema":{"name":"Extracted","baseSelector":"//h1","fields":[{"name":"text","selector":".","type":"text"}]}}'