XML DOM Access
The DOM API provides methods to find and read any node in an XML document. You access the document, then navigate to specific elements, attributes, or text nodes using these methods. All major programming languages — JavaScript, Python, Java, C# — expose the DOM through similar interfaces.
Loading XML into the DOM
In JavaScript (Browser)
var parser = new DOMParser(); var xmlDoc = parser.parseFromString(xmlString, "text/xml");
Loading from a File in JavaScript
var xhr = new XMLHttpRequest();
xhr.open("GET", "catalog.xml", false);
xhr.send();
var xmlDoc = xhr.responseXML;
In Python (lxml)
from lxml import etree
tree = etree.parse("catalog.xml")
root = tree.getroot()
Accessing the Root Element
var root = xmlDoc.documentElement; console.log(root.nodeName); // prints the root element name
getElementsByTagName()
Returns a live NodeList of all elements with the given tag name, searched throughout the entire document (or within a specific element).
// All book elements in the document
var books = xmlDoc.getElementsByTagName("book");
console.log(books.length); // number of books
// Access by index (0-based)
var firstBook = books[0];
var secondBook = books[1];
// Search within a specific element
var prices = firstBook.getElementsByTagName("price");
console.log(prices[0].textContent); // text inside first price
getElementById()
Returns the single element with the specified id attribute. The attribute must be declared as type ID in the DTD or Schema for this to work in XML (unlike HTML where id always works).
var book = xmlDoc.getElementById("B001");
getAttribute()
Reads the value of a named attribute on an element.
var book = xmlDoc.getElementsByTagName("book")[0];
var id = book.getAttribute("id"); // "B001"
var genre = book.getAttribute("genre"); // "fiction"
// Returns null if attribute does not exist
var missing = book.getAttribute("xyz"); // null
textContent and textNode
var title = xmlDoc.getElementsByTagName("title")[0];
// textContent: all text inside the element (including child element text)
console.log(title.textContent); // "XML Basics"
// childNodes[0].nodeValue: the first text node's value
console.log(title.childNodes[0].nodeValue); // "XML Basics"
Practical Access Example
XML: catalog.xml
<catalog>
<product id="P001" category="electronics">
<name>Wireless Mouse</name>
<price>799</price>
<stock>45</stock>
</product>
<product id="P002" category="accessories">
<name>USB Hub</name>
<price>1299</price>
<stock>20</stock>
</product>
</catalog>
JavaScript to Read All Products
var products = xmlDoc.getElementsByTagName("product");
for (var i = 0; i < products.length; i++) {
var p = products[i];
var id = p.getAttribute("id");
var name = p.getElementsByTagName("name")[0].textContent;
var price = p.getElementsByTagName("price")[0].textContent;
var stock = p.getElementsByTagName("stock")[0].textContent;
console.log(id + ": " + name + " — INR " + price + " (" + stock + " units)");
}
Output in Console
P001: Wireless Mouse — INR 799 (45 units) P002: USB Hub — INR 1299 (20 units)
hasAttribute() and hasChildNodes()
var p = products[0];
p.hasAttribute("category"); // true
p.hasAttribute("discount"); // false
p.hasChildNodes(); // true (has name, price, stock children)
Key Points to Remember
- getElementsByTagName() returns all matching elements as a NodeList.
- getAttribute() reads an attribute value by name — returns null if missing.
- textContent returns all text content inside an element including nested elements.
- childNodes[0].nodeValue accesses the raw text node value directly.
- hasAttribute() and hasChildNodes() safely check before reading.
- Access is read-only — modifying requires the DOM modification methods covered next.
XML DOM Modify
Reading the DOM gives you data. Modifying the DOM changes the document in memory. You can add new elements, change text, update attributes, and delete nodes — all without touching the original XML file. The changes exist in memory and can be serialized back to XML when needed.
Changing Element Text Content
// Get the price element
var price = xmlDoc.getElementsByTagName("price")[0];
// Change the text content
price.childNodes[0].nodeValue = "999";
// Or using textContent (simpler, modern browsers)
price.textContent = "999";
Setting Attribute Values
var product = xmlDoc.getElementsByTagName("product")[0];
// Change existing attribute
product.setAttribute("category", "peripherals");
// Add a new attribute
product.setAttribute("onSale", "true");
// Remove an attribute
product.removeAttribute("onSale");
Creating New Elements
// Create a new element
var newProduct = xmlDoc.createElement("product");
// Create child elements for it
var nameEl = xmlDoc.createElement("name");
var priceEl = xmlDoc.createElement("price");
// Create text nodes for the children
var nameText = xmlDoc.createTextNode("Webcam HD");
var priceText = xmlDoc.createTextNode("2499");
// Build the tree
nameEl.appendChild(nameText);
priceEl.appendChild(priceText);
newProduct.appendChild(nameEl);
newProduct.appendChild(priceEl);
// Add the new product to the catalog
var catalog = xmlDoc.documentElement;
catalog.appendChild(newProduct);
Inserting Before an Existing Node
// Insert newProduct before the second existing product
var secondProduct = xmlDoc.getElementsByTagName("product")[1];
catalog.insertBefore(newProduct, secondProduct);
Removing Nodes
// Remove a specific element
var products = xmlDoc.getElementsByTagName("product");
var toRemove = products[1]; // second product
toRemove.parentNode.removeChild(toRemove);
Replacing Nodes
// Create a replacement node
var replacement = xmlDoc.createElement("product");
replacement.setAttribute("id", "P999");
// Replace the first product
var old = xmlDoc.getElementsByTagName("product")[0];
old.parentNode.replaceChild(replacement, old);
Cloning Nodes
// Shallow clone — copies the element but not its children var shallowCopy = product.cloneNode(false); // Deep clone — copies the element AND all its children var deepCopy = product.cloneNode(true); catalog.appendChild(deepCopy);
Complete Modification Example
// Add a discount attribute to all products over 1000
var products = xmlDoc.getElementsByTagName("product");
for (var i = 0; i < products.length; i++) {
var price = parseInt(products[i].getElementsByTagName("price")[0].textContent);
if (price > 1000) {
products[i].setAttribute("discount", "10%");
}
}
DOM Modification Methods Summary
| Method | Purpose |
|---|---|
| createElement(tag) | Creates a new element node |
| createTextNode(text) | Creates a new text node |
| appendChild(node) | Adds node as the last child |
| insertBefore(new, ref) | Inserts new node before ref node |
| removeChild(node) | Removes a child node |
| replaceChild(new, old) | Replaces old child with new child |
| cloneNode(deep) | Copies a node (true = deep copy) |
| setAttribute(name, val) | Sets or creates an attribute |
| removeAttribute(name) | Removes an attribute |
Key Points to Remember
- DOM modifications happen in memory — the original XML file is not changed.
- createElement() and createTextNode() create nodes; appendChild() attaches them.
- setAttribute() adds or updates an attribute; removeAttribute() deletes it.
- removeChild() requires calling it on the parent node.
- cloneNode(true) deep-copies an element including all its children.
- After modifications, serialize the DOM to a string to save the updated XML.
XML DOM Navigation
DOM navigation means moving through the node tree using parent-child and sibling relationships. The DOM provides properties on every node that point to neighboring nodes. Navigation is essential when you process complex XML structures where you need to move up, down, or sideways through the tree.
Navigation Properties
| Property | Returns |
|---|---|
| parentNode | The parent node |
| childNodes | NodeList of all child nodes |
| firstChild | The first child node |
| lastChild | The last child node |
| nextSibling | The next node at the same level |
| previousSibling | The previous node at the same level |
| firstElementChild | First child that is an element (skips text nodes) |
| lastElementChild | Last child that is an element |
| nextElementSibling | Next sibling that is an element |
| previousElementSibling | Previous sibling that is an element |
Navigation Example
var catalog = xmlDoc.documentElement;
// First child (may be a text/whitespace node)
var first = catalog.firstChild;
// First element child (skips whitespace text nodes)
var firstProduct = catalog.firstElementChild;
console.log(firstProduct.nodeName); // "product"
// Move to the next product sibling
var secondProduct = firstProduct.nextElementSibling;
console.log(secondProduct.getAttribute("id")); // "P002"
// Move back
var backToFirst = secondProduct.previousElementSibling;
// Go up to parent
var parent = firstProduct.parentNode;
console.log(parent.nodeName); // "catalog"
Navigating Through All Children
var book = xmlDoc.getElementsByTagName("book")[0];
var children = book.childNodes;
for (var i = 0; i < children.length; i++) {
var child = children[i];
// Skip whitespace text nodes
if (child.nodeType === 1) {
console.log(child.nodeName + ": " + child.textContent);
}
}
Key Points to Remember
- parentNode, childNodes, firstChild, lastChild, nextSibling, previousSibling navigate the full node tree including text nodes.
- firstElementChild, nextElementSibling, and similar properties skip text nodes — preferred for element-only navigation.
- Always check nodeType === 1 before reading nodeName or accessing child elements when using childNodes.
- Navigation properties return null when no matching node exists.
XML DOM NodeList
When DOM methods like getElementsByTagName() return multiple nodes, they return a NodeList — a collection of nodes. Understanding how to work with NodeLists is essential for processing XML with multiple repeating elements.
NodeList Basics
var books = xmlDoc.getElementsByTagName("book");
// Length — number of nodes in the list
console.log(books.length); // e.g., 3
// Access by index (0-based)
var first = books[0];
var second = books[1];
var last = books[books.length - 1];
// Using item() method — same as bracket notation
var first2 = books.item(0);
Looping Over a NodeList
var products = xmlDoc.getElementsByTagName("product");
// Traditional for loop — most compatible
for (var i = 0; i < products.length; i++) {
console.log(products[i].getAttribute("id"));
}
// Modern Array.from (converts NodeList to a true Array)
Array.from(products).forEach(function(product) {
console.log(product.textContent);
});
Live vs Static NodeLists
NodeLists returned by getElementsByTagName() are live — they update automatically when the document changes. If you add a book element, the books NodeList immediately contains the new node. NodeLists returned by querySelectorAll() are static — they do not update after creation.
// Live NodeList
var books = xmlDoc.getElementsByTagName("book");
console.log(books.length); // 2
// Add a new book
var newBook = xmlDoc.createElement("book");
xmlDoc.documentElement.appendChild(newBook);
console.log(books.length); // 3 — automatically updated!
Converting NodeList to Array
// Method 1: Array.from
var booksArray = Array.from(xmlDoc.getElementsByTagName("book"));
// Method 2: spread operator
var booksArray2 = [...xmlDoc.getElementsByTagName("book")];
// Now you can use array methods
var titles = booksArray.map(book =>
book.getElementsByTagName("title")[0].textContent
);
Key Points to Remember
- A NodeList is a collection of nodes — not a JavaScript array, but array-like.
- Access nodes by index (books[0]) or using the item() method (books.item(0)).
- NodeLists from getElementsByTagName() are live — they reflect document changes automatically.
- Use Array.from() or the spread operator [...] to convert a NodeList to a true array.
- Always check .length before accessing nodes to avoid null reference errors.
