XML CDATA

When you write an XML document, certain characters cause serious problems. For example, the less-than sign < tells the XML parser that a new tag is starting. If you write < inside your text content, the parser gets confused and throws an error. This is where CDATA saves you.

CDATA stands for Character Data. It tells the XML parser: "Ignore everything inside this block. Treat it as plain text, not as XML code."

Why CDATA Exists

Think of a recipe card inside a box. The box has a label on the outside that says "Do Not Read Instructions Inside." Anyone who picks up the box reads the outer label and skips the inner content. CDATA works the same way — it wraps content so the XML parser skips over it without trying to interpret it as XML tags.

Without CDATA, you must replace special characters with escape codes:

  • < becomes &lt;
  • > becomes &gt;
  • & becomes &amp;

Replacing every special character manually is tedious and error-prone. CDATA removes this problem entirely.

CDATA Syntax

A CDATA section starts with <![CDATA[ and ends with ]]>. Everything between these two markers is treated as raw text.

Basic Structure

<![CDATA[ your text goes here ]]>

Real Example Without CDATA

Suppose you store a math formula inside XML. Without CDATA, you must escape the less-than sign:

<formula>5 &lt; 10 and 20 &gt; 15</formula>

Same Example With CDATA

<formula><![CDATA[5 < 10 and 20 > 15]]></formula>

The CDATA version is cleaner and much easier to read.

A Diagram-Based Example

Imagine a bookstore XML file. You want to store a book description that contains HTML code for formatting.

<book>
  <title>Learning XML</title>
  <description>
    <![CDATA[
      <b>This book</b> covers <i>XML basics</i>.
      Price is always 5 < 10 dollars for members.
    ]]>
  </description>
</book>

The XML parser reads the <book> and <description> tags normally. When it reaches <![CDATA[, it stops parsing and reads everything until ]]> as plain text. The HTML tags <b> and <i> inside the CDATA block do not confuse the XML parser at all.

Visual Flow Diagram

XML Parser reads file
        |
        v
  Finds <description>
        |
        v
  Finds <![CDATA[  ----> STOP PARSING
        |
        v
  Reads everything as plain text
        |
        v
  Finds ]]>  ----------> RESUME PARSING
        |
        v
  Finds </description> — closes tag normally

Where CDATA Is Most Useful

CDATA appears most often in these situations:

  • Storing HTML inside XML: Web feeds like RSS use XML to store HTML-formatted articles. CDATA wraps the HTML content so the XML parser ignores the HTML tags.
  • Storing code snippets: Programming tutorials often put code examples inside XML. Code frequently contains <, >, and & characters.
  • Storing mathematical expressions: Formulas use symbols like <, >, and & that conflict with XML syntax.
  • Storing SQL queries: Database queries often use the ampersand and comparison operators.

CDATA Limitations

CDATA sections have one restriction you must remember. The closing marker ]]> cannot appear anywhere inside the CDATA block. If your text contains ]]>, you must split your CDATA into two separate sections.

Example of the Problem

<!-- This causes an error -->
<data><![CDATA[ Hello ]]> World ]]></data>

Correct Fix Using Two CDATA Sections

<data><![CDATA[ Hello ]]>&gt;<![CDATA[ World ]]></data>

CDATA vs Escape Characters — Which to Use?

Both methods solve the same problem. Use CDATA when your content has many special characters and escaping them all would make the text unreadable. Use escape characters when only one or two special characters appear in the text.

Quick Comparison

SituationBest Choice
One or two special charactersEscape codes (&lt;, &gt;)
Large block of HTML or codeCDATA section
SQL queries or formulasCDATA section
Simple product descriptionsEscape codes

CDATA in RSS Feeds

RSS feeds are XML files that news websites and blogs use to publish updates. Almost every RSS feed uses CDATA to wrap article descriptions. This lets the description contain full HTML formatting without breaking the XML structure.

<item>
  <title>New Product Launch</title>
  <description>
    <![CDATA[
      <p>Our <b>latest product</b> is now available.</p>
      <ul>
        <li>Feature 1</li>
        <li>Feature 2</li>
      </ul>
    ]]>
  </description>
</item>

Key Points to Remember

  • CDATA tells the XML parser to skip the content inside and treat it as plain text.
  • Use <![CDATA[ to start and ]]> to end a CDATA section.
  • The string ]]> must not appear inside a CDATA block.
  • CDATA does not change the data — it only changes how the parser handles it.
  • Multiple CDATA sections can appear inside one XML element.
  • CDATA works only inside element content, not inside attribute values.

Leave a Comment

Your email address will not be published. Required fields are marked *