XML with Python
Python ships with multiple XML libraries in its standard library. For most tasks, xml.etree.ElementTree is the simplest choice. For advanced processing with XPath and XSLT, the third-party lxml library is the standard tool.
Reading XML with ElementTree
import xml.etree.ElementTree as ET
tree = ET.parse("products.xml")
root = tree.getroot()
print(root.tag) # catalog
for product in root.findall("product"):
pid = product.get("id")
name = product.find("name").text
price = product.find("price").text
print(f"{pid}: {name} — INR {price}")
Parsing XML from a String
xml_data = """<catalog>
<product id="P001">
<name>Notebook</name>
<price>40</price>
</product>
</catalog>"""
root = ET.fromstring(xml_data)
for p in root:
print(p.find("name").text)
Searching with find and findall
# First matching element
first = root.find("product")
# All matching elements
all_products = root.findall("product")
# XPath-like selectors in ElementTree
expensive = root.findall("product[price]") # products that have a price child
# Iterate all with a condition
for p in root.findall("product"):
if int(p.find("price").text) > 500:
print(p.find("name").text)
Building XML with ElementTree
root = ET.Element("catalog")
product = ET.SubElement(root, "product")
product.set("id", "P001")
name = ET.SubElement(product, "name")
name.text = "Webcam"
price = ET.SubElement(product, "price")
price.text = "2499"
tree = ET.ElementTree(root)
tree.write("new_catalog.xml", xml_declaration=True, encoding="utf-8")
Using lxml for XPath
from lxml import etree
tree = etree.parse("products.xml")
# XPath query
results = tree.xpath("//product[price > 1000]/name/text()")
for name in results:
print(name)
Key Points to Remember
- xml.etree.ElementTree is the standard, batteries-included XML library in Python.
- ET.parse() opens a file; ET.fromstring() parses an XML string.
- find() returns the first match; findall() returns a list of all matches.
- element.get("attr") reads attributes; element.text reads the text content.
- lxml adds full XPath and XSLT support and better performance for large files.
