XML Tree Structure
Every XML document has a tree structure. This means the document has one starting point at the top, and all elements branch downward from it in a parent-child hierarchy — exactly like an upside-down tree. Understanding this structure is essential because every tool that reads, searches, or transforms XML — XPath, XSLT, the DOM, XQuery — navigates this tree.
Think of an XML tree like a company organisation chart. The CEO is at the top. Each manager below the CEO has their own team. Each team member reports to exactly one manager. No employee can report to two managers at the same time. XML elements follow the same rule — every element has exactly one parent.
Tree Terminology
| Term | Meaning |
|---|---|
| Root | The single top-level element — every document has exactly one |
| Parent | An element that contains other elements |
| Child | An element directly inside another element |
| Sibling | Elements at the same level with the same parent |
| Ancestor | A parent, grandparent, or any element above this one |
| Descendant | A child, grandchild, or any element below this one |
| Leaf node | An element with no children — it holds only text |
A Tree Built from XML
XML Document
<?xml version="1.0" encoding="UTF-8"?>
<school>
<class grade="10">
<student roll="S01">
<name>Aditya Sharma</name>
<marks>92</marks>
</student>
<student roll="S02">
<name>Bhavna Patel</name>
<marks>85</marks>
</student>
</class>
<class grade="11">
<student roll="S03">
<name>Chirag Nair</name>
<marks>78</marks>
</student>
</class>
</school>
The Same Document as a Tree
Document Root
|
[school] ←── Root element
|
├── [class grade="10"] ←── Child of school
│ |
│ ├── ←── Child of class, sibling of S02
│ │ ├── [name] → "Aditya Sharma" ←── Leaf node
│ │ └── [marks] → "92" ←── Leaf node
│ │
│ └── ←── Sibling of S01
│ ├── [name] → "Bhavna Patel"
│ └── [marks] → "85"
│
└── [class grade="11"] ←── Sibling of the first class
|
└──
├── [name] → "Chirag Nair"
└── [marks] → "78"
Relationships in the Tree
school → parent of both class elements class[10] → child of school; parent of student S01 and S02 student[S01] → child of class[10]; sibling of student[S02] name → child of student; leaf node (no children) school → ancestor of name (grandparent) name → descendant of school
The Document Node vs The Root Element
The tree has one invisible node at the very top called the document node. It is the parent of the root element, any processing instructions before the root, and any comments before the root. The root element is the first child of the document node.
Document Node (invisible, at the very top)
|
├── Processing Instruction: <?xml version="1.0"?>
├── Comment: <!-- School data -->
└── Element: <school> ← this is the root element
XPath uses a single forward slash / to refer to the document node, not the root element. The root element is /school.
Node Types in the Tree
The XML tree contains several types of nodes, not just elements:
| Node Type | Example |
|---|---|
| Document node | The invisible root of the entire tree |
| Element node | <school>, <student>, <name> |
| Attribute node | grade="10", roll="S01" |
| Text node | "Aditya Sharma", "92" |
| Comment node | <!-- comment text --> |
| Processing instruction node | <?xml-stylesheet ...?> |
Why the Tree Structure Matters
Every XML technology uses the tree to do its work:
- XPath: Navigates the tree using path expressions like /school/class/student/name
- XSLT: Matches elements in the tree using templates and outputs a transformed tree
- DOM: Loads the entire tree into memory so your code can read and modify any node
- XQuery: Queries the tree using FLWOR expressions to extract and reshape data
- SAX: Reads the tree node by node, firing events as it goes
Nesting Rules the Tree Enforces
Because XML is a tree, elements cannot overlap. Every child element must open and close completely inside its parent. This is called proper nesting.
WRONG — overlapping (not a valid tree): <bold><italic>text</bold></italic> CORRECT — properly nested (valid tree): <bold><italic>text</italic></bold>
Key Points to Remember
- An XML document is a tree with one root element at the top.
- Every element except the root has exactly one parent.
- Parent, child, sibling, ancestor, and descendant describe element relationships.
- Leaf nodes are elements that contain only text and have no children.
- The document node is the invisible top of the tree — the parent of the root element.
- All XML tools — XPath, XSLT, DOM, XQuery — navigate and operate on this tree.
- Proper nesting is required because overlapping elements break the tree structure.
