XML DOM Introduction

The Document Object Model, commonly called the DOM, is a programming interface that represents an XML (or HTML) document as a tree of objects in memory. Once a parser loads an XML file into the DOM, every element, attribute, text node, and comment becomes a live object your program can read, search, modify, add, or remove — without touching the original file on disk.

Think of the DOM like a physical model of a building. The architect's blueprint (the XML file) describes the structure. The physical model (the DOM) is the actual 3D object you can rearrange — move rooms, add floors, knock down walls — all while the original blueprint stays unchanged.

Why the DOM Exists

Before the DOM, programs had to manually scan XML text character by character to extract data — an enormous amount of repetitive, error-prone code. The DOM standardises this: any language that implements the DOM API (JavaScript, Java, Python, C#, PHP) provides the same set of methods to navigate and manipulate any XML document.

DOM vs SAX — Two Ways to Process XML

FeatureDOMSAX
Loads entire documentYes — full tree in memoryNo — reads one event at a time
Random accessYes — jump to any nodeNo — forward-only
Modify the documentYesNo
Memory usageHigherVery low
Best forSmall–medium files, editing, searchingVery large files, one-pass reading

The DOM Tree

When a DOM parser loads an XML file, it builds a tree of node objects. Every part of the XML document becomes a node.

XML File

<?xml version="1.0"?>
<library>
  <book id="B01">
    <title>XML Basics</title>
    <author>Meena Iyer</author>
  </book>
</library>

DOM Tree in Memory

Document Node
  └── Element: library
        └── Element: book
              ├── Attribute: id = "B01"
              ├── Element: title
              │     └── Text: "XML Basics"
              └── Element: author
                    └── Text: "Meena Iyer"

Core DOM Node Types

Node TypenodeType NumberExample
Document9The top of the entire tree
Element1<library>, <book>
Attribute2id="B01"
Text3"XML Basics"
Comment8<!-- comment -->
Processing Instruction7<?xml-stylesheet ...?>

Essential DOM Properties

PropertyReturns
nodeNameThe tag name for elements, "#text" for text nodes
nodeValueThe text for text nodes; null for elements
nodeTypeNumeric type (1 = element, 3 = text, etc.)
parentNodeThe parent node
childNodesAll child nodes as a NodeList
firstChildThe first child node
lastChildThe last child node
nextSiblingThe next node at the same level
attributesAll attribute nodes of an element
textContentAll text inside an element including descendants

Essential DOM Methods

MethodPurpose
getElementsByTagName(name)Returns all elements with the given tag name
getElementById(id)Returns the element with the given id
getAttribute(name)Returns the value of a named attribute
setAttribute(name, value)Sets or adds an attribute
createElement(name)Creates a new element node
createTextNode(text)Creates a new text node
appendChild(node)Adds a node as the last child
removeChild(node)Removes a child node

A Quick DOM Example in JavaScript

var parser = new DOMParser();
var xmlDoc = parser.parseFromString(xmlString, "text/xml");

// Get all book elements
var books = xmlDoc.getElementsByTagName("book");

// Read the first book's title
var title = books[0].getElementsByTagName("title")[0].textContent;
console.log(title);   // "XML Basics"

// Read an attribute
var id = books[0].getAttribute("id");
console.log(id);      // "B01"

DOM in Different Languages

LanguageDOM API
JavaScript (Browser)Built-in DOMParser and document object
Pythonxml.dom.minidom, xml.etree.ElementTree, lxml
Javajavax.xml.parsers.DocumentBuilder
C# / .NETSystem.Xml.XmlDocument
PHPDOMDocument class

When to Use the DOM

  • You need to read and modify multiple parts of the document in any order.
  • Your XML file is small enough to fit comfortably in memory.
  • You want to build new XML documents from scratch in code.
  • You need to search for specific nodes by tag name, id, or XPath expression.

Key Points to Remember

  • The DOM represents an XML document as a navigable, modifiable tree of objects in memory.
  • Every element, attribute, text, and comment is a node with properties and a nodeType number.
  • getElementsByTagName() and getAttribute() are the two most-used DOM methods for reading XML.
  • The DOM API is standardised — the same concepts apply in JavaScript, Java, Python, and C#.
  • DOM loads the entire document into memory — use SAX for very large XML files instead.
  • Changes made to the DOM are in memory only — write the tree back to a file to save them.

Leave a Comment

Your email address will not be published. Required fields are marked *