XML SAX Parser
SAX stands for Simple API for XML. A SAX parser reads the XML document as a stream and fires event callbacks each time it encounters a start tag, end tag, or text content. It never builds a complete tree in memory. This makes SAX ideal for processing gigabyte-scale XML files where a DOM tree would exhaust available memory.
SAX Event Model
XML Input:
<order id="O1">
<item>USB Cable</item>
<price>99</price>
</order>
SAX fires these events in sequence:
startDocument()
startElement("order", {id: "O1"})
startElement("item", {})
characters("USB Cable")
endElement("item")
startElement("price", {})
characters("99")
endElement("price")
endElement("order")
endDocument()
SAX Parser in Python
import xml.sax
class ProductHandler(xml.sax.ContentHandler):
def __init__(self):
self.current_element = ""
self.current_product = {}
def startElement(self, name, attrs):
self.current_element = name
if name == "product":
self.current_product = {"id": attrs["id"]}
def characters(self, content):
content = content.strip()
if content:
if self.current_element == "name":
self.current_product["name"] = content
elif self.current_element == "price":
self.current_product["price"] = content
def endElement(self, name):
if name == "product":
print(f"{self.current_product['id']}: "
f"{self.current_product['name']} — "
f"INR {self.current_product['price']}")
self.current_element = ""
parser = xml.sax.make_parser()
handler = ProductHandler()
parser.setContentHandler(handler)
parser.parse("catalog.xml")
Key Points to Remember
- SAX parses XML as a stream, firing events — never builds a tree in memory.
- Write handler methods: startElement, endElement, characters.
- SAX is faster and uses far less memory than DOM for large files.
- SAX is forward-only — you cannot go back or access nodes out of order.
- Use SAX when processing very large XML files or when you only need a subset of the data.
