XPath Nodes
XPath works on a tree model. Every piece of an XML document — every element, every attribute, every piece of text — is a node in this tree. Understanding node types is essential because XPath expressions select, filter, and navigate specific types of nodes.
The Seven XPath Node Types
1. Root Node (Document Node)
The root node is the invisible top of the tree. It is the parent of the root element and everything else in the document. Select it with a single forward slash.
/ ← selects the root node
The root node is different from the root element. The root node contains the root element, processing instructions, and comments that appear outside the root element.
2. Element Nodes
Element nodes are the XML tags. Every element — opening and closing tag pair — is one element node in the tree. An element node can have child nodes (other elements, text, attributes, comments).
XML:
<catalog>
<product>
<name>USB Hub</name>
</product>
</catalog>
Element nodes: catalog, product, name
3. Attribute Nodes
Attribute nodes belong to element nodes but are not considered children. They are a separate node type. Select them using the @ symbol.
XML: <product id="P001" category="electronics"> Attribute nodes of <product>: id = "P001" category = "electronics" XPath to select: /catalog/product/@id XPath to select all attributes: /catalog/product/@*
4. Text Nodes
Text nodes hold the actual text content inside elements. A text node is a child of the element that contains the text. Whitespace between elements also creates text nodes.
XML: <name>USB Hub</name> Text node: "USB Hub" XPath to select: /catalog/product/name/text() Without text(), name selects the element node. With text(), name/text() selects only the string "USB Hub".
5. Comment Nodes
XML comments are nodes in the tree. Select them using the comment() function in XPath.
XML: <catalog> <!-- Products updated March 2024 --> <product>...</product> </catalog> XPath to select comment: /catalog/comment()
6. Processing Instruction Nodes
Processing instructions that appear in the XML document are also nodes. Select them using the processing-instruction() function.
XML: <?xml-stylesheet type="text/xsl" href="style.xsl"?> <catalog>...</catalog> XPath: /processing-instruction()
7. Namespace Nodes
Each element that has a namespace declaration carries namespace nodes. These are rarely selected directly in XPath but exist in the tree model.
XML: <catalog xmlns:prod="http://example.com/products"> Namespace node on catalog: prod → http://example.com/products
Node Type Summary Diagram
Document Root Node (/)
|
+-- Processing Instruction Node (<?xml-stylesheet ...?>)
|
+-- Element Node: <catalog>
|
+-- Comment Node: <!-- updated -->
|
+-- Element Node: <product>
|
+-- Attribute Node: id="P001"
|
+-- Attribute Node: category="electronics"
|
+-- Element Node: <name>
| |
| +-- Text Node: "USB Hub"
|
+-- Element Node: <price>
|
+-- Text Node: "1299"
Node Relationships
XPath uses family-style terms to describe relationships between nodes.
| Relationship | Meaning |
|---|---|
| Parent | The node directly above this node |
| Child | A node directly below this node |
| Sibling | A node at the same level with the same parent |
| Ancestor | Parent, grandparent, and all above |
| Descendant | Child, grandchild, and all below |
Example with the library XML
<library>
<book>
<title>XML Basics</title>
<author>Meena</author>
</book>
</library>
library → parent of book
book → child of library; parent of title and author
title → child of book; sibling of author
author → child of book; sibling of title
"XML Basics" → text node; child of title
library → ancestor of title
title → descendant of library
Selecting Nodes by Type
| XPath Function | Selects |
|---|---|
| node() | All node types — elements, text, comments, PIs |
| text() | Text nodes only |
| comment() | Comment nodes only |
| processing-instruction() | Processing instruction nodes only |
| element() | Element nodes only (XPath 2.0+) |
Examples
/library/book/node() → All child nodes of book (elements + text + comments) /library/book/text() → Text nodes directly inside book /library/comment() → Comments inside library /library//text() → All text nodes anywhere inside library
Node Values
Each node type has a string value — the text that XPath extracts when you call string() or use xsl:value-of.
| Node Type | String Value |
|---|---|
| Element node | Concatenation of all descendant text nodes |
| Attribute node | The attribute value |
| Text node | The text string itself |
| Comment node | The comment text (without <!-- -->) |
| Root node | Concatenation of all text in the document |
Atomic Values vs Nodes
XPath 2.0 distinguishes between nodes (parts of the XML tree) and atomic values (plain data like strings and numbers). An element node contains structure; an atomic value is just a piece of data extracted from a node.
Node: <price>1299</price> ← has structure, children, position in tree Atomic value: "1299" ← just the string, no tree context
Key Points to Remember
- XPath treats an XML document as a tree of seven types of nodes.
- The root node (/) is the invisible top of the tree, parent of the root element.
- Element nodes are XML tags; attribute nodes belong to elements but are not children.
- Text nodes hold the actual text content inside elements.
- Use text() to select text nodes, comment() for comments, and node() for all types.
- Node relationships use family terms: parent, child, sibling, ancestor, descendant.
- The string value of an element node is the concatenation of all its descendant text.
