XML DOM Access

The DOM API provides methods to find and read any node in an XML document. You access the document, then navigate to specific elements, attributes, or text nodes using these methods. All major programming languages — JavaScript, Python, Java, C# — expose the DOM through similar interfaces.

Loading XML into the DOM

In JavaScript (Browser)

var parser = new DOMParser();
var xmlDoc = parser.parseFromString(xmlString, "text/xml");

Loading from a File in JavaScript

var xhr = new XMLHttpRequest();
xhr.open("GET", "catalog.xml", false);
xhr.send();
var xmlDoc = xhr.responseXML;

In Python (lxml)

from lxml import etree
tree = etree.parse("catalog.xml")
root = tree.getroot()

Accessing the Root Element

var root = xmlDoc.documentElement;
console.log(root.nodeName);   // prints the root element name

getElementsByTagName()

Returns a live NodeList of all elements with the given tag name, searched throughout the entire document (or within a specific element).

// All book elements in the document
var books = xmlDoc.getElementsByTagName("book");
console.log(books.length);          // number of books

// Access by index (0-based)
var firstBook = books[0];
var secondBook = books[1];

// Search within a specific element
var prices = firstBook.getElementsByTagName("price");
console.log(prices[0].textContent);  // text inside first price

getElementById()

Returns the single element with the specified id attribute. The attribute must be declared as type ID in the DTD or Schema for this to work in XML (unlike HTML where id always works).

var book = xmlDoc.getElementById("B001");

getAttribute()

Reads the value of a named attribute on an element.

var book = xmlDoc.getElementsByTagName("book")[0];
var id = book.getAttribute("id");          // "B001"
var genre = book.getAttribute("genre");    // "fiction"

// Returns null if attribute does not exist
var missing = book.getAttribute("xyz");    // null

textContent and textNode

var title = xmlDoc.getElementsByTagName("title")[0];

// textContent: all text inside the element (including child element text)
console.log(title.textContent);      // "XML Basics"

// childNodes[0].nodeValue: the first text node's value
console.log(title.childNodes[0].nodeValue);  // "XML Basics"

Practical Access Example

XML: catalog.xml

<catalog>
  <product id="P001" category="electronics">
    <name>Wireless Mouse</name>
    <price>799</price>
    <stock>45</stock>
  </product>
  <product id="P002" category="accessories">
    <name>USB Hub</name>
    <price>1299</price>
    <stock>20</stock>
  </product>
</catalog>

JavaScript to Read All Products

var products = xmlDoc.getElementsByTagName("product");

for (var i = 0; i < products.length; i++) {
  var p = products[i];
  var id = p.getAttribute("id");
  var name = p.getElementsByTagName("name")[0].textContent;
  var price = p.getElementsByTagName("price")[0].textContent;
  var stock = p.getElementsByTagName("stock")[0].textContent;

  console.log(id + ": " + name + " — INR " + price + " (" + stock + " units)");
}

Output in Console

P001: Wireless Mouse — INR 799 (45 units)
P002: USB Hub — INR 1299 (20 units)

hasAttribute() and hasChildNodes()

var p = products[0];

p.hasAttribute("category");   // true
p.hasAttribute("discount");   // false

p.hasChildNodes();            // true (has name, price, stock children)

Key Points to Remember

  • getElementsByTagName() returns all matching elements as a NodeList.
  • getAttribute() reads an attribute value by name — returns null if missing.
  • textContent returns all text content inside an element including nested elements.
  • childNodes[0].nodeValue accesses the raw text node value directly.
  • hasAttribute() and hasChildNodes() safely check before reading.
  • Access is read-only — modifying requires the DOM modification methods covered next.

XML DOM Modify

Reading the DOM gives you data. Modifying the DOM changes the document in memory. You can add new elements, change text, update attributes, and delete nodes — all without touching the original XML file. The changes exist in memory and can be serialized back to XML when needed.

Changing Element Text Content

// Get the price element
var price = xmlDoc.getElementsByTagName("price")[0];

// Change the text content
price.childNodes[0].nodeValue = "999";

// Or using textContent (simpler, modern browsers)
price.textContent = "999";

Setting Attribute Values

var product = xmlDoc.getElementsByTagName("product")[0];

// Change existing attribute
product.setAttribute("category", "peripherals");

// Add a new attribute
product.setAttribute("onSale", "true");

// Remove an attribute
product.removeAttribute("onSale");

Creating New Elements

// Create a new element
var newProduct = xmlDoc.createElement("product");

// Create child elements for it
var nameEl = xmlDoc.createElement("name");
var priceEl = xmlDoc.createElement("price");

// Create text nodes for the children
var nameText = xmlDoc.createTextNode("Webcam HD");
var priceText = xmlDoc.createTextNode("2499");

// Build the tree
nameEl.appendChild(nameText);
priceEl.appendChild(priceText);
newProduct.appendChild(nameEl);
newProduct.appendChild(priceEl);

// Add the new product to the catalog
var catalog = xmlDoc.documentElement;
catalog.appendChild(newProduct);

Inserting Before an Existing Node

// Insert newProduct before the second existing product
var secondProduct = xmlDoc.getElementsByTagName("product")[1];
catalog.insertBefore(newProduct, secondProduct);

Removing Nodes

// Remove a specific element
var products = xmlDoc.getElementsByTagName("product");
var toRemove = products[1];         // second product
toRemove.parentNode.removeChild(toRemove);

Replacing Nodes

// Create a replacement node
var replacement = xmlDoc.createElement("product");
replacement.setAttribute("id", "P999");

// Replace the first product
var old = xmlDoc.getElementsByTagName("product")[0];
old.parentNode.replaceChild(replacement, old);

Cloning Nodes

// Shallow clone — copies the element but not its children
var shallowCopy = product.cloneNode(false);

// Deep clone — copies the element AND all its children
var deepCopy = product.cloneNode(true);

catalog.appendChild(deepCopy);

Complete Modification Example

// Add a discount attribute to all products over 1000
var products = xmlDoc.getElementsByTagName("product");
for (var i = 0; i < products.length; i++) {
  var price = parseInt(products[i].getElementsByTagName("price")[0].textContent);
  if (price > 1000) {
    products[i].setAttribute("discount", "10%");
  }
}

DOM Modification Methods Summary

MethodPurpose
createElement(tag)Creates a new element node
createTextNode(text)Creates a new text node
appendChild(node)Adds node as the last child
insertBefore(new, ref)Inserts new node before ref node
removeChild(node)Removes a child node
replaceChild(new, old)Replaces old child with new child
cloneNode(deep)Copies a node (true = deep copy)
setAttribute(name, val)Sets or creates an attribute
removeAttribute(name)Removes an attribute

Key Points to Remember

  • DOM modifications happen in memory — the original XML file is not changed.
  • createElement() and createTextNode() create nodes; appendChild() attaches them.
  • setAttribute() adds or updates an attribute; removeAttribute() deletes it.
  • removeChild() requires calling it on the parent node.
  • cloneNode(true) deep-copies an element including all its children.
  • After modifications, serialize the DOM to a string to save the updated XML.

XML DOM Navigation

DOM navigation means moving through the node tree using parent-child and sibling relationships. The DOM provides properties on every node that point to neighboring nodes. Navigation is essential when you process complex XML structures where you need to move up, down, or sideways through the tree.

Navigation Properties

PropertyReturns
parentNodeThe parent node
childNodesNodeList of all child nodes
firstChildThe first child node
lastChildThe last child node
nextSiblingThe next node at the same level
previousSiblingThe previous node at the same level
firstElementChildFirst child that is an element (skips text nodes)
lastElementChildLast child that is an element
nextElementSiblingNext sibling that is an element
previousElementSiblingPrevious sibling that is an element

Navigation Example

var catalog = xmlDoc.documentElement;

// First child (may be a text/whitespace node)
var first = catalog.firstChild;

// First element child (skips whitespace text nodes)
var firstProduct = catalog.firstElementChild;
console.log(firstProduct.nodeName);       // "product"

// Move to the next product sibling
var secondProduct = firstProduct.nextElementSibling;
console.log(secondProduct.getAttribute("id"));  // "P002"

// Move back
var backToFirst = secondProduct.previousElementSibling;

// Go up to parent
var parent = firstProduct.parentNode;
console.log(parent.nodeName);     // "catalog"

Navigating Through All Children

var book = xmlDoc.getElementsByTagName("book")[0];
var children = book.childNodes;

for (var i = 0; i < children.length; i++) {
  var child = children[i];
  // Skip whitespace text nodes
  if (child.nodeType === 1) {
    console.log(child.nodeName + ": " + child.textContent);
  }
}

Key Points to Remember

  • parentNode, childNodes, firstChild, lastChild, nextSibling, previousSibling navigate the full node tree including text nodes.
  • firstElementChild, nextElementSibling, and similar properties skip text nodes — preferred for element-only navigation.
  • Always check nodeType === 1 before reading nodeName or accessing child elements when using childNodes.
  • Navigation properties return null when no matching node exists.

XML DOM NodeList

When DOM methods like getElementsByTagName() return multiple nodes, they return a NodeList — a collection of nodes. Understanding how to work with NodeLists is essential for processing XML with multiple repeating elements.

NodeList Basics

var books = xmlDoc.getElementsByTagName("book");

// Length — number of nodes in the list
console.log(books.length);     // e.g., 3

// Access by index (0-based)
var first = books[0];
var second = books[1];
var last = books[books.length - 1];

// Using item() method — same as bracket notation
var first2 = books.item(0);

Looping Over a NodeList

var products = xmlDoc.getElementsByTagName("product");

// Traditional for loop — most compatible
for (var i = 0; i < products.length; i++) {
  console.log(products[i].getAttribute("id"));
}

// Modern Array.from (converts NodeList to a true Array)
Array.from(products).forEach(function(product) {
  console.log(product.textContent);
});

Live vs Static NodeLists

NodeLists returned by getElementsByTagName() are live — they update automatically when the document changes. If you add a book element, the books NodeList immediately contains the new node. NodeLists returned by querySelectorAll() are static — they do not update after creation.

// Live NodeList
var books = xmlDoc.getElementsByTagName("book");
console.log(books.length);     // 2

// Add a new book
var newBook = xmlDoc.createElement("book");
xmlDoc.documentElement.appendChild(newBook);

console.log(books.length);     // 3 — automatically updated!

Converting NodeList to Array

// Method 1: Array.from
var booksArray = Array.from(xmlDoc.getElementsByTagName("book"));

// Method 2: spread operator
var booksArray2 = [...xmlDoc.getElementsByTagName("book")];

// Now you can use array methods
var titles = booksArray.map(book =>
  book.getElementsByTagName("title")[0].textContent
);

Key Points to Remember

  • A NodeList is a collection of nodes — not a JavaScript array, but array-like.
  • Access nodes by index (books[0]) or using the item() method (books.item(0)).
  • NodeLists from getElementsByTagName() are live — they reflect document changes automatically.
  • Use Array.from() or the spread operator [...] to convert a NodeList to a true array.
  • Always check .length before accessing nodes to avoid null reference errors.

Leave a Comment

Your email address will not be published. Required fields are marked *