XML White Space
White space refers to spaces, tabs, line breaks, and carriage returns in an XML file. How XML parsers handle this white space is a topic that confuses many beginners. Understanding white space helps you write cleaner XML and avoid unexpected behavior in applications.
Think of white space like the empty space on a printed page. It makes text readable for humans. But the printer does not print blank space — it prints only characters. XML parsers face the same question: should they keep or discard the spaces your eyes see?
Two Types of White Space in XML
1. Insignificant White Space
White space that appears between elements — in the indentation and line breaks you add to make your XML readable — is called insignificant white space. Parsers typically ignore it because it carries no data meaning.
<bookstore>
<book>
<title>XML Basics</title>
</book>
</bookstore>
The two spaces before <book> and the four spaces before <title> are insignificant white space. They help you read the structure but carry no data.
2. Significant White Space
White space inside element content — the text between an opening and closing tag — is significant. Parsers preserve it because it is part of the data.
<address>42, MG Road Bangalore</address>
The three spaces between "Road" and "Bangalore" are significant. They are part of the address text and must not be removed.
White Space Diagram
<library> ← line break after this tag = insignificant
<book> ← 2 spaces indent = insignificant
<title> ← 4 spaces indent = insignificant
Learn XML ← "Learn XML" text is content
</title> ← spaces here = significant (inside element)
</book>
</library>
The xml:space Attribute
XML provides the xml:space attribute to tell parsers how to handle white space in a specific element. It accepts two values.
xml:space="preserve"
The parser keeps all white space exactly as written — spaces, tabs, and line breaks included.
<poem xml:space="preserve"> Roses are red, Violets are blue, XML is structured, And so are you. </poem>
Every line break and leading space in the poem is preserved. Remove xml:space="preserve" and the application may collapse all the white space into a single block of text.
xml:space="default"
The parser handles white space using its default behavior, which usually means collapsing multiple spaces into one and trimming leading and trailing spaces.
<description xml:space="default"> This is a product description. </description>
With the default setting, an application might collapse multiple spaces into single spaces.
White Space in Attribute Values
Attribute values follow their own white space rules. XML parsers normalize attribute white space automatically:
- Tab characters become single spaces.
- Line breaks become single spaces.
- Multiple consecutive spaces collapse into one space.
- Leading and trailing spaces are removed.
<!-- Written like this --> <item code=" A 100 " /> <!-- Parser normalizes the attribute value to --> <item code="A 100" />
White Space in Different Contexts
Between Elements
<catalog> <item>Pen</item> <item>Book</item> </catalog>
The blank lines between elements are insignificant white space. Most parsers ignore them, but a strict parser may create text nodes for them in the DOM.
Inside a Mixed-Content Element
A mixed-content element holds both text and child elements. White space inside such elements is always significant because removing it changes the meaning of the text.
<paragraph> Click <link>here</link> to continue. </paragraph>
The space before "here" and after "here" are both significant. Removing them changes "Click here to continue" into "Clickherecontinue".
White Space in XML vs HTML
| Situation | XML Behavior | HTML Behavior |
|---|---|---|
| Multiple spaces in content | Preserved | Collapsed to one |
| Line breaks in content | Preserved | Treated as space |
| Leading/trailing spaces in attributes | Normalized | Normalized |
| Spaces between elements | May create text nodes | Usually ignored |
Practical Tips for White Space
- Use xml:space="preserve" for poetry, code snippets, and pre-formatted text.
- Avoid unnecessary white space between elements if your application processes the DOM, as it creates extra text nodes.
- Never rely on white space normalization for attribute values in business logic — trim values in your application code instead.
- Test your XML output in the actual parser your application uses, as different parsers handle edge cases differently.
Key Points to Remember
- Insignificant white space appears between elements and is used for readability.
- Significant white space appears inside element content and carries data meaning.
- The xml:space="preserve" attribute tells parsers to keep all white space unchanged.
- Attribute white space is always normalized — tabs and line breaks become spaces.
- Mixed-content elements treat all white space as significant.
- Different parsers handle insignificant white space differently — test your specific parser.
