XML Encoding

Every XML file stores text as a sequence of numbers in memory. The encoding tells the computer which number corresponds to which character. Without encoding information, a computer cannot correctly read your file — it may display garbled text or throw an error.

Think of encoding like a translation dictionary. The file contains numbers like 65, 66, 67. The encoding tells the system: "65 means A, 66 means B, 67 means C." Change the encoding and the same numbers give different characters.

The XML Declaration and Encoding

XML files declare their encoding in the very first line called the XML prolog. The prolog looks like this:

<?xml version="1.0" encoding="UTF-8"?>

The encoding attribute tells the XML parser which character set to use when reading the file. If you skip the encoding declaration, most parsers assume UTF-8 by default.

Why Encoding Matters

Imagine a French restaurant menu stored in XML. The menu has words like résumé and naïve. If you save the file with one encoding but declare a different encoding in the prolog, the parser reads the wrong numbers for those accented characters. The result is broken text that shows strange symbols instead of proper letters.

Correct Encoding Declaration

<?xml version="1.0" encoding="UTF-8"?>
<menu>
  <item>Café au Lait</item>
  <item>Crêpe Suzette</item>
</menu>

The encoding declaration and the actual file encoding match — so the parser reads accented characters correctly.

Common XML Encodings

UTF-8

UTF-8 is the most popular encoding for XML files worldwide. It covers every character from every language — English, Arabic, Chinese, Hindi, and thousands more. It uses between 1 and 4 bytes per character. Simple English letters use only 1 byte, which keeps the file size small.

UTF-16

UTF-16 uses 2 bytes minimum for each character. It covers the same huge range of characters as UTF-8. UTF-16 files are sometimes larger than UTF-8 for English text, but they work better for languages like Chinese and Japanese where most characters need 2 bytes anyway.

ISO-8859-1

ISO-8859-1 is an older encoding also called Latin-1. It covers Western European languages like English, French, German, Spanish, and Portuguese. It uses exactly 1 byte per character and supports 256 different characters. Avoid this encoding for modern projects because it cannot handle Asian characters, Arabic, or many other scripts.

Encoding Comparison Diagram

Character: "A"
  UTF-8      → stores as: 0x41 (1 byte)
  UTF-16     → stores as: 0x00 0x41 (2 bytes)
  ISO-8859-1 → stores as: 0x41 (1 byte)

Character: "é" (e with accent)
  UTF-8      → stores as: 0xC3 0xA9 (2 bytes)
  UTF-16     → stores as: 0x00 0xE9 (2 bytes)
  ISO-8859-1 → stores as: 0xE9 (1 byte)

Character: "中" (Chinese character)
  UTF-8      → stores as: 0xE4 0xB8 0xAD (3 bytes)
  UTF-16     → stores as: 0x4E 0x2D (2 bytes)
  ISO-8859-1 → CANNOT represent this character

What Happens When Encoding Goes Wrong

Suppose a file is saved as UTF-8 but the XML prolog says ISO-8859-1. The parser reads the file using ISO-8859-1 rules. Multi-byte UTF-8 characters get split into separate ISO-8859-1 bytes. The result is garbage text like Crêpe instead of Crêpe.

Wrong Declaration

<?xml version="1.0" encoding="ISO-8859-1"?>
<!-- File is actually saved as UTF-8 -->
<item>Crêpe</item>
<!-- Parser reads: Crêpe -->

Correct Declaration

<?xml version="1.0" encoding="UTF-8"?>
<!-- File is saved as UTF-8 -->
<item>Crêpe</item>
<!-- Parser reads: Crêpe -->

Rules for XML Encoding

  • The encoding declaration must appear in the XML prolog on the very first line.
  • No characters can appear before the XML prolog — not even a blank line.
  • The encoding value is case-insensitive. UTF-8, utf-8, and Utf-8 all mean the same thing.
  • The actual file encoding on disk must match the encoding declared in the prolog.
  • If no encoding is declared, XML parsers default to UTF-8 or UTF-16.

Byte Order Mark (BOM)

Some text editors add a special hidden marker at the very start of UTF-16 files called the Byte Order Mark or BOM. The BOM is two bytes that tell the parser whether the file is big-endian or little-endian — two different ways computers store 2-byte numbers.

Big-endian stores the most important byte first. Little-endian stores the least important byte first. Think of writing a phone number: big-endian puts the country code first; little-endian puts it last.

UTF-16 Big-endian BOM:    0xFE 0xFF
UTF-16 Little-endian BOM: 0xFF 0xFE

UTF-8 does not need a BOM, though some editors add one anyway. XML parsers that encounter a UTF-8 BOM should handle it gracefully, but some older parsers reject it.

Choosing the Right Encoding for Your XML

Follow these guidelines when creating XML files:

Use UTF-8 when:

  • Your content mixes English with any other language.
  • You share the file across different operating systems.
  • You work with web applications, REST APIs, or RSS feeds.

Use UTF-16 when:

  • Your content is primarily East Asian text (Chinese, Japanese, Korean).
  • The application you work with requires UTF-16.

Use ISO-8859-1 only when:

  • You maintain a legacy system that already uses this encoding.
  • Your content uses only Western European characters and no other scripts.

Checking and Setting Encoding in Text Editors

Most text editors let you set the encoding before saving a file. In Notepad++ on Windows, go to Encoding in the menu and choose Encode in UTF-8. In VS Code, click the encoding label at the bottom right of the screen and select the encoding you want.

Always match the editor encoding to the encoding you declare in the XML prolog. This one step prevents the most common encoding errors in XML development.

Key Points to Remember

  • Encoding tells the parser how to convert bytes into readable characters.
  • Declare encoding in the XML prolog using the encoding attribute.
  • UTF-8 is the recommended encoding for nearly all XML files.
  • The declared encoding must match the actual file encoding on disk.
  • Mismatched encoding causes garbled text — often called mojibake.
  • XML parsers default to UTF-8 when no encoding declaration appears.

Leave a Comment

Your email address will not be published. Required fields are marked *