Ctrl + K
XML21 min read

XPath Guide

Learn how XPath works, how to select XML and HTML elements, and how to write practical XPath expressions.

Published: 2026-10-05

XPath is a language for selecting and navigating nodes in XML and HTML documents. It is commonly used in XML processing, web scraping, browser automation, testing and applications that need to locate specific elements inside a structured document.

Unlike a simple text search, XPath understands the hierarchical structure of a document. An expression can select an element by its name, attribute, position, text content or relationship to other nodes. XPath also provides functions and predicates that allow increasingly precise selections.

What Is XPath?

XPath stands for XML Path Language. It provides expressions for navigating through the nodes of an XML document and selecting nodes that match specific conditions. Although XPath was designed for XML, it is also widely used with HTML documents because HTML can be represented as a tree of elements.

An XPath expression describes where a desired node is located or what properties it must have. The result can be a single element, multiple elements, an attribute, text content or another type of node depending on the expression and XPath implementation.

//book
//book/title
//book/@id

The first expression selects book elements, the second selects title elements that are descendants of book elements, and the third selects id attributes belonging to book elements.

XPath and the Document Tree

XPath works with the hierarchical structure of a document. An XML document can be viewed as a tree where elements contain other elements, attributes belong to elements, and text appears inside element nodes.

<catalog>
  <book id="101">
    <title>XPath Guide</title>
    <author>Alice</author>
  </book>
  <book id="102">
    <title>XML Basics</title>
    <author>Bob</author>
  </book>
</catalog>

In this document, catalog is the root element. It contains two book elements, and each book contains title and author elements. The id value is an attribute of each book element.

Basic XPath Syntax

The slash character is one of the most important parts of XPath syntax. It describes movement through the document hierarchy. A single slash creates a path from the current context, while a double slash searches for matching descendants at any depth.

ExpressionMeaning
/Select from the root or move to a direct child.
//Select matching descendants at any depth.
.Current context node.
..Parent of the current node.
@Attribute selection.
*Wildcard matching any element or node of the relevant type.

Absolute XPath

An absolute XPath starts from the root of the document and describes the complete path to the target node. It can be useful when the document structure is fixed and well known, but it can become fragile when elements are added, removed or rearranged.

/catalog/book/title

This expression starts at the document root, selects catalog, then book, and finally title. If the target title is nested differently, the expression will no longer match it.

Relative XPath

A relative XPath does not require the complete path from the root. It can search from the current context or use descendant selection to find matching nodes deeper in the document.

book/title
//title
.//title

Relative expressions are often more flexible because they depend less on the exact position of the target node in the entire document.

Selecting Elements by Name

The simplest XPath expression selects elements by their element name. If an expression contains a node name without additional conditions, XPath returns nodes with that name that match the selected path.

//book
//title
//author

The double slash makes these expressions search for matching descendants regardless of their depth.

Selecting a Specific Child

A single slash can be used to select direct children. This is useful when the document structure is known and the relationship between elements matters.

/catalog/book
/catalog/book/title

The second expression selects title elements that are direct children of book elements, which are themselves direct children of catalog.

The Wildcard Selector

The asterisk wildcard selects elements without requiring a specific element name. This can be useful when the element name is unknown or when several different element types need to be selected.

/*
/catalog/*
/book/*

For example, /catalog/* selects all direct child elements of catalog, regardless of whether they are book, magazine or another element type.

Selecting Attributes

XPath uses the @ symbol to access attributes. An attribute can be selected directly or used as part of a condition.

//book/@id
//@id
//book/@*

The first expression selects id attributes of book elements. The second searches for id attributes anywhere in the document. The third selects all attributes belonging to book elements.

Selecting Elements by Attribute

Attributes are frequently used to identify specific elements. XPath predicates allow an expression to select only nodes whose attributes satisfy a condition.

//book[@id="101"]
//book[@id]
//*[@class="product"]

The first expression selects the book whose id is 101. The second selects book elements that have an id attribute regardless of its value. The third selects any element whose class attribute is exactly product.

Predicates in XPath

A predicate is a condition placed inside square brackets. Predicates filter the nodes selected by the part of the expression before the brackets.

//book[author="Alice"]
//book[@id="102"]
//book[title="XML Basics"]

Predicates can check element values, attributes, positions and more complex conditions. They are one of the most important features for writing precise XPath expressions.

Selecting by Text

The text() function can be used to match text nodes. This is especially common when selecting HTML elements based on visible text or XML elements containing a known value.

//title[text()="XPath Guide"]
//button[text()="Submit"]
//p[text()="Hello"]

These expressions look for elements whose direct text node exactly matches the specified value. In real HTML, text can sometimes be split across nested elements, so exact text matching may not always behave as expected.

The contains() Function

The contains() function checks whether one string contains another string. It is useful when the exact text or attribute value is not known or when a partial match is sufficient.

//title[contains(text(),"XPath")]
//div[contains(@class,"product")]
//a[contains(@href,"/products/")]

The expressions above select titles containing XPath, elements whose class attribute contains product, and links whose href contains /products/.

The starts-with() Function

The starts-with() function checks whether a string begins with a particular value. It can be useful for predictable prefixes in attributes or text.

//a[starts-with(@href,"/products/")]
//div[starts-with(@id,"item-")]

This can be useful when identifiers are generated dynamically but consistently begin with a known prefix.

Combining Conditions

XPath supports logical operators such as and and or. They allow several conditions to be combined inside a predicate.

//book[@id="101" and author="Alice"]
//book[@id="101" or @id="102"]
//product[@available="true" and price>10]

Use and when all conditions must match. Use or when at least one condition should match.

Selecting by Position

XPath can select nodes according to their position among matching siblings. Position-based selection is useful when a document contains repeated elements and the desired node is identified by its order.

//book[1]
//book[2]
//book[last()]
//book[position()=3]

These expressions select the first book, second book, last book and third book respectively.

⚠️ Position-based XPath can become fragile when the order of elements changes. If an element has a stable attribute or another unique property, selecting it by that property is often more robust.

The last() Function

The last() function returns the size of the current node set when used in a positional predicate. It is commonly used to select the final matching element.

//book[last()]
/catalog/book[last()]

Both expressions select the last matching book within their respective context.

XPath Axes

XPath axes describe relationships between the current node and other nodes in the document tree. They provide more control than simply moving from parent to child.

AxisDescription
childSelects direct child nodes.
parentSelects the parent node.
ancestorSelects all ancestor nodes.
ancestor-or-selfSelects ancestors and the current node.
descendantSelects descendant nodes.
descendant-or-selfSelects descendants and the current node.
following-siblingSelects later siblings.
preceding-siblingSelects earlier siblings.
followingSelects nodes appearing later in the document.
precedingSelects nodes appearing earlier in the document.
selfSelects the current node.
attributeSelects attributes of the current element.

Child Axis

The child axis selects direct children of the current node. It is equivalent to the normal slash-based child navigation used in many XPath expressions.

child::book
/catalog/child::book
child::title

The explicit child:: syntax is more verbose than the abbreviated form, so most everyday XPath expressions use the shorter syntax.

Parent Axis

The parent axis moves from the current node to its parent. The abbreviated form is two periods.

//title/..
/catalog/book/title/..

For example, //title/.. selects the parent elements of matching title nodes.

Ancestor Axis

The ancestor axis selects all parents above the current node. This is useful when an element needs to be related to a container higher in the document hierarchy.

//title/ancestor::book
//span/ancestor::div

These expressions find book ancestors of title elements and div ancestors of span elements.

Descendant Axis

The descendant axis selects nodes at any depth below the current node. The // abbreviation is commonly used for this type of search.

//catalog/descendant::title
/catalog//title

Both expressions can select title elements anywhere below catalog.

Sibling Axes

Sibling axes are useful when the target element is identified by its relationship to another element at the same level. following-sibling selects later siblings, while preceding-sibling selects earlier siblings.

//title/following-sibling::author
//author/preceding-sibling::title

These expressions select an author appearing after a title and a title appearing before an author within the same parent.

Using XPath with HTML

XPath is frequently used to locate elements in HTML documents. This makes it useful for browser automation, web scraping and automated testing.

<div class="product">
  <h2>Keyboard</h2>
  <span class="price">$49</span>
</div>

For this document, XPath can identify the product, its heading or its price using element names and attributes.

//div[@class="product"]
//div[@class="product"]/h2
//div[@class="product"]//span[@class="price"]

XPath for Web Automation

Automation frameworks can use XPath expressions to locate buttons, links, inputs and other elements. XPath is especially useful when an element does not have a convenient unique identifier.

//button[@type="submit"]
//input[@name="email"]
//a[contains(@href,"/account")]
//button[contains(text(),"Continue")]

However, an XPath should generally identify an element using stable characteristics rather than depending on a long chain of nested containers. Stable IDs, names, data attributes or distinctive combinations of attributes are usually easier to maintain.

XPath in Automated Testing

Testing frameworks can use XPath to locate elements before interacting with them or checking their state. This is common in browser-based end-to-end testing.

const button = page.locator('xpath=//button[@type="submit"]');
await button.click();

The exact API depends on the testing framework. XPath itself only defines the expression used to locate nodes; the framework determines how that expression is evaluated and how the resulting elements are manipulated.

XPath for Web Scraping

Web scraping libraries can use XPath to extract structured information from HTML. XPath is useful when data follows a predictable document structure and selectors need to express relationships between elements.

//article//h2
//article//a/@href
//article[contains(@class,"featured")]//h2

For example, these expressions can select article headings, article links and headings inside featured articles.

XPath Functions

XPath includes functions for working with strings, numbers, node sets and other values. Functions become particularly useful when simple element and attribute matching is not enough.

FunctionPurpose
text()Selects text nodes.
contains()Checks whether a string contains another string.
starts-with()Checks whether a string begins with a value.
last()Returns the size of the current node set.
position()Returns the position of the current node.
count()Counts nodes in a node set.
normalize-space()Removes surrounding whitespace and normalizes internal whitespace.
string-length()Returns the length of a string.
substring()Extracts part of a string.

normalize-space()

The normalize-space() function is useful when text contains inconsistent whitespace. It trims leading and trailing whitespace and normalizes sequences of whitespace characters.

//p[normalize-space()="Hello world"]
//button[normalize-space(text())="Submit"]

This can make text-based matching more tolerant of formatting differences in the document.

count() and string Functions

XPath can also calculate values rather than simply selecting nodes. For example, count() returns the number of matching nodes, while string-related functions can transform or inspect text.

count(//book)
string-length(//title[1])
substring(//title[1],1,5)

Whether these expressions return useful results depends on the XPath version and the API evaluating them. XPath 1.0, XPath 2.0 and XPath 3.x provide different capabilities, so compatibility should be checked when using advanced expressions.

XPath 1.0, 2.0 and 3.x

XPath has evolved through several versions. XPath 1.0 provides the core navigation, selection and function features used by many browser automation and XML libraries. Later versions add more powerful data types, functions, operators and expression capabilities.

VersionGeneral Characteristics
XPath 1.0Widely implemented and commonly encountered in browsers and automation tools.
XPath 2.0Adds richer data types, sequences and substantially more functions.
XPath 3.0Extends the expression language with additional capabilities.
XPath 3.1Adds features such as maps and arrays.

When writing XPath for a specific application, use the version supported by the parser, browser, automation framework or XML processing library. An expression valid in a newer XPath version may not work in an XPath 1.0 environment.

XPath Operators

XPath supports comparison and logical operators that can be used inside predicates and expressions. Common operators include =, !=, <, >, <=, >=, and, or and not().

//product[price>50]
//product[@status!="inactive"]
//book[author="Alice" and @id="101"]
//book[not(@archived)]

These operators make it possible to filter nodes using several properties instead of relying only on element names.

Using Multiple XPath Conditions

Complex documents often require several conditions at the same time. Combining attributes, text and structural relationships can produce a precise selector without relying on a long absolute path.

//div[
  @class="product" and
  .//h2[contains(text(),"Keyboard")] and
  .//span[@class="price"]
]

This expression searches for a product container with the expected class and requires it to contain both a matching heading and a price element.

XPath and CSS Selectors

XPath and CSS selectors can both locate elements in HTML, but they provide different capabilities. CSS selectors are often concise and familiar to frontend developers, while XPath provides powerful navigation through parent, ancestor and sibling relationships.

FeatureXPathCSS Selector
Element selectionYesYes
Attribute matchingYesYes
Text-based conditionsYesLimited depending on environment
Parent navigationYesLimited
Ancestor navigationYesLimited
Sibling relationshipsYesYes
Complex document relationshipsVery flexibleGenerally simpler
Browser familiarityCommonVery common

For straightforward element selection, CSS selectors are often easier to read. XPath becomes particularly useful when the target is defined by text or by a relationship to another node.

Choosing Stable XPath Expressions

A technically valid XPath is not necessarily a good XPath. In automation and testing, selectors should remain useful when the page changes slightly. Expressions that depend on stable attributes are generally easier to maintain than expressions that depend on exact element positions.

  • Prefer stable IDs, names or data attributes when they are available.
  • Use meaningful attributes instead of long absolute paths.
  • Use text matching when the visible text is stable and distinctive.
  • Avoid unnecessary positional predicates.
  • Avoid selectors that depend on automatically generated class names.
  • Keep complex expressions understandable and testable.
  • Check the selector against realistic document variations.
💡 When creating an XPath for automation, start with the most stable identifier available and add conditions only when they are necessary to make the selection unique.

Common XPath Mistakes

XPath expressions often fail because of small differences in document structure, text content or XPath version support. Understanding the most common mistakes makes debugging much easier.

  • Using an absolute path that breaks after a small DOM change.
  • Assuming text() matches all visible text when nested elements are present.
  • Using exact class matching when an element contains multiple classes.
  • Relying on element position when a stable attribute is available.
  • Forgetting the difference between direct children and descendants.
  • Using functions that are unavailable in the XPath version being used.
  • Ignoring namespaces when working with XML documents.
  • Assuming an XPath that works in one tool is automatically supported by every XPath implementation.

Namespaces in XPath

XML namespaces can make XPath more complicated because element names may belong to a namespace even when the visible XML name looks familiar. XPath implementations usually require the namespace to be registered or referenced correctly before namespaced elements can be selected.

<book xmlns="https://example.com/books">
  <title>XPath Guide</title>
</book>

An expression such as //title may not match the title element when the document uses a default namespace. The exact solution depends on the XML library and its namespace handling API.

⚠️ When XPath unexpectedly returns no results for valid XML, check whether the document uses namespaces. Namespace handling is a common source of confusion in XML queries.

How to Test an XPath Expression

Testing an XPath expression against the actual document is often faster than trying to reason about it from the expression alone. XPath testers and browser developer tools can help verify which nodes are selected.

  • Start with a simple expression that identifies the general element.
  • Add an attribute or text condition.
  • Check how many nodes are returned.
  • Add positional or structural conditions only when necessary.
  • Test the expression against realistic variations of the document.
  • Verify that the XPath version supported by the target environment can evaluate the expression.

Practical XPath Examples

The following examples demonstrate common patterns that can be adapted for XML and HTML documents.

GoalXPath
Find all links//a
Find links with a specific href//a[@href="/login"]
Find inputs by name//input[@name="email"]
Find buttons containing text//button[contains(text(),"Submit")]
Find elements with a class//*[contains(@class,"product")]
Find the first book//book[1]
Find the last book//book[last()]
Find a parent element//title/..
Find a specific descendant//article//h2
Find elements without an attribute//book[not(@id)]

XPath Best Practices

  • Prefer relative expressions when they make the selector less fragile.
  • Use stable attributes whenever possible.
  • Keep expressions as simple as the task allows.
  • Use predicates to express meaningful conditions.
  • Use axes when relationships between nodes matter.
  • Avoid unnecessary dependence on element positions.
  • Be careful with exact text matching when content can change.
  • Consider CSS selectors when the task does not require XPath-specific capabilities.
  • Test expressions in the same environment where they will be used.
  • Account for XML namespaces when querying namespaced documents.

XPath vs Regular Expressions

XPath and regular expressions solve different problems. XPath understands the structure of an XML or HTML document and selects nodes based on that structure. Regular expressions operate on strings and are better suited to finding or transforming text patterns.

For structured XML or HTML data, XPath is generally more appropriate for selecting elements and relationships. Regular expressions can still be useful for processing text values after the relevant node has been selected.

XPath vs Parsing the Document Manually

XPath provides a structured way to query a document without manually traversing every element in application code. A parser builds or exposes the document structure, while XPath provides expressions for locating the nodes that matter.

This separation can make extraction and automation logic easier to maintain. Instead of writing custom traversal code for every query, an application can use a concise XPath expression against the parsed document.

When Should You Use XPath?

  • You need to select specific nodes from an XML document.
  • You are extracting structured data from HTML.
  • You are writing browser automation or end-to-end tests.
  • You need to locate elements based on text or relationships.
  • You need to navigate from an element to its parent or ancestors.
  • You need precise filtering using attributes and predicates.
  • Your XML processing library supports XPath queries.

When XPath May Not Be the Best Choice

XPath is not automatically the best selector for every HTML task. CSS selectors can be shorter and easier to maintain for straightforward element and attribute selection. In application code, a dedicated parser API may also be clearer when only a small amount of traversal is required.

The important consideration is the problem being solved. XPath is particularly valuable when document relationships, text conditions or XML-specific navigation are central to the query.

Frequently Asked Questions

What is XPath used for?

XPath is used to navigate and select nodes in XML and HTML documents. Common applications include XML processing, web scraping, browser automation and automated testing.

What does // mean in XPath?

The // syntax searches for matching descendant nodes at any depth below the current context. For example, //title can find title elements throughout a document.

What does @ mean in XPath?

The @ symbol is used to select or reference attributes. For example, //book/@id selects id attributes from book elements.

What is an XPath predicate?

A predicate is a condition inside square brackets that filters selected nodes. Examples include [@id="101"], [1] and [contains(text(),"Guide")].

How do I select an element by text in XPath?

The text() function can be used to match element text, such as //button[text()="Submit"]. The contains() function can be used when only part of the text needs to match.

Is XPath better than CSS selectors?

Neither is universally better. CSS selectors are often simpler for straightforward HTML selection, while XPath provides powerful navigation through parents, ancestors, siblings and text-based conditions.

Why is my XPath not finding an XML element?

Common causes include an incorrect document path, namespace handling, text differences, unsupported XPath functions or an expression that assumes a different document structure.

Helpful XPath and XML Tools

An XPath Tester can help verify expressions against XML or HTML and inspect which nodes are selected. An XPath Generator can produce candidate expressions from a document structure, while an XML Tree Viewer makes nested elements, attributes and relationships easier to inspect. XML Formatters improve readability when the source document is compressed or poorly formatted, and XML Validators can identify structural or syntax problems that may otherwise make XPath debugging confusing.

Conclusion

XPath provides a powerful way to navigate and query XML and HTML documents using paths, predicates, attributes, functions and node relationships. Basic expressions can select elements and attributes, while more advanced expressions can filter nodes by text, position, structure and multiple conditions. XPath is widely useful in XML processing, web scraping, browser automation and automated testing. For reliable selectors, prefer stable attributes and meaningful relationships instead of fragile absolute paths or unnecessary positional conditions. Once the basic syntax and axes become familiar, XPath can express complex document queries in a relatively compact form.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.