UTF-16 Explained
A practical guide to UTF-16 encoding, covering code units, surrogate pairs, the Basic Multilingual Plane, BOMs, endianness, JavaScript strings, Unicode escapes, and common encoding problems.
UTF-16 is a Unicode encoding that represents Unicode text using 16-bit code units. It was designed to provide a practical way to represent a large portion of Unicode using one 16-bit unit while still supporting characters outside that range through pairs of code units called surrogate pairs.
UTF-16 is especially important for JavaScript, Java, .NET, Windows APIs, and other environments that historically use 16-bit code units for strings. It is also the reason that a Unicode character such as an emoji can appear to have a string length of two in JavaScript even though it is perceived as one character.
This guide explains how UTF-16 works, what code units and code points are, how surrogate pairs are calculated, why endianness matters, how the BOM works, and how UTF-16 compares with UTF-8 and UTF-32.
What Is UTF-16?
UTF-16 stands for Unicode Transformation Format - 16-bit. It is a variable-width Unicode encoding based on 16-bit code units.
A UTF-16 encoded text stream uses one or two 16-bit code units for each Unicode code point. Characters in the Basic Multilingual Plane can generally be represented using one code unit, while supplementary characters require two code units in the form of a surrogate pair.
| Unicode Range | UTF-16 Representation |
|---|---|
| U+0000–U+D7FF | 1 code unit |
| U+E000–U+FFFF | 1 code unit |
| U+10000–U+10FFFF | 2 code units |
The surrogate range U+D800 through U+DFFF is reserved for surrogate pairs and is not used as an independent Unicode scalar value.
UTF-16 Is Not the Same as Unicode
Unicode and UTF-16 describe different layers of text representation. Unicode defines a universal character repertoire and assigns code points to characters. UTF-16 defines how those code points are represented using 16-bit code units.
- Unicode defines code points such as U+0041, U+0416, and U+1F600.
- A Unicode code point identifies a value in the Unicode code space.
- UTF-16 converts that code point into one or two 16-bit code units.
- The resulting code units are ultimately stored or transmitted as bytes.
Character:
😀
Unicode code point:
U+1F600
UTF-16 code units:
D83D DE00The important distinction is that D83D and DE00 are UTF-16 code units, not two separate Unicode characters. Together they represent the single Unicode code point U+1F600.
What Is a UTF-16 Code Unit?
A code unit is the basic storage unit of an encoding. UTF-16 uses 16-bit code units, so each code unit contains exactly 16 bits or two bytes.
1 UTF-16 code unit:
16 bits
2 bytes
Example:
0041
Hexadecimal:
00 41 (big-endian)
41 00 (little-endian)A code unit should not automatically be treated as a Unicode character. Some Unicode code points require two UTF-16 code units, and a surrogate pair must be interpreted together.
UTF-16 and the Basic Multilingual Plane
Unicode is divided into several planes. The Basic Multilingual Plane, or BMP, contains code points from U+0000 through U+FFFF.
Most commonly used scripts and symbols are located in the BMP. A large number of BMP characters can therefore be represented directly by a single UTF-16 code unit.
A
U+0041
UTF-16 code unit:
0041
Я
U+042F
UTF-16 code unit:
042F
€
U+20AC
UTF-16 code unit:
20ACThese examples fit directly into a 16-bit value and therefore do not need a surrogate pair.
Why UTF-16 Needs Surrogate Pairs
A single 16-bit value can represent 65,536 possible values, but Unicode has more code points than can fit into one 16-bit unit. UTF-16 solves this by reserving two ranges of 16-bit values for surrogate pairs.
| Range | Purpose |
|---|---|
| U+D800–U+DBFF | High surrogates |
| U+DC00–U+DFFF | Low surrogates |
| U+D800–U+DFFF | Reserved surrogate range |
A supplementary Unicode code point from U+10000 through U+10FFFF is represented by one high surrogate followed by one low surrogate.
What Is a Surrogate Pair?
A surrogate pair is two UTF-16 code units that together encode one Unicode code point above U+FFFF.
😀
Unicode code point:
U+1F600
UTF-16:
D83D DE00The first code unit is the high surrogate and the second is the low surrogate. They must be interpreted together to recover the original supplementary code point.
How Surrogate Pairs Are Calculated
To encode a supplementary code point using UTF-16, subtract U+10000 from the code point. The resulting 20-bit value is split into a high 10-bit portion and a low 10-bit portion.
Code point:
U+10000–U+10FFFF
1. Subtract U+10000
2. Split the remaining 20 bits
3. Add the upper 10 bits to D800
4. Add the lower 10 bits to DC00For a supplementary code point P, the mathematical formulas can be expressed as follows:
T = P - 0x10000
High surrogate = 0xD800 + (T >> 10)
Low surrogate = 0xDC00 + (T & 0x3FF)Surrogate Pair Example: U+1F600
The grinning face emoji is U+1F600. It is outside the BMP, so it requires a surrogate pair.
Code point:
U+1F600
Subtract U+10000:
0x1F600 - 0x10000 = 0xF600
High surrogate:
D800 + upper 10 bits = D83D
Low surrogate:
DC00 + lower 10 bits = DE00
Result:
D83D DE00Therefore, U+1F600 is represented by the UTF-16 code-unit sequence D83D DE00.
UTF-16 Byte Representation
A UTF-16 code unit is 16 bits, but computers store bytes. This introduces a byte-order question: which byte comes first?
| Encoding | Example for U+0041 |
|---|---|
| UTF-16BE | 00 41 |
| UTF-16LE | 41 00 |
UTF-16 therefore has two commonly encountered byte orders: big-endian and little-endian.
UTF-16 Big-Endian
In big-endian UTF-16, the most significant byte of each 16-bit code unit is stored first.
Character:
A
UTF-16 code unit:
0041
UTF-16BE bytes:
00 41The high byte appears before the low byte. This ordering is often written as UTF-16BE.
UTF-16 Little-Endian
In little-endian UTF-16, the least significant byte is stored first.
Character:
A
UTF-16 code unit:
0041
UTF-16LE bytes:
41 00This ordering is written as UTF-16LE. It is common on systems based on little-endian processor architectures.
What Is a BOM?
A byte-order mark, or BOM, is a special Unicode character placed at the beginning of a text stream to indicate encoding information. In UTF-16, the BOM is particularly useful because it can identify byte order.
| Encoding | BOM Bytes |
|---|---|
| UTF-16BE | FE FF |
| UTF-16LE | FF FE |
The Unicode code point represented by the BOM is U+FEFF. Its byte representation changes depending on the byte order.
How a UTF-16 BOM Identifies Endianness
Suppose a file starts with the bytes FF FE. A decoder can interpret these bytes as the UTF-16LE representation of U+FEFF and conclude that the remaining data should be interpreted as little-endian UTF-16.
UTF-16LE:
FF FE
UTF-16BE:
FE FFThe BOM therefore acts as an encoding signature and byte-order indicator. Applications should still follow the rules of the relevant file or protocol rather than blindly assuming that every UTF-16 stream must contain a BOM.
UTF-16 With and Without a BOM
A UTF-16 encoded stream does not inherently require a BOM. If the byte order is specified externally, a BOM may be unnecessary. For example, a protocol can explicitly state that all text is UTF-16LE.
| Situation | BOM Role |
|---|---|
| Byte order known from protocol | May not be needed |
| Byte order must be detected from file | Can be useful |
| UTF-16LE with BOM | FF FE at beginning |
| UTF-16BE with BOM | FE FF at beginning |
UTF-16 and ASCII
Unlike UTF-8, UTF-16 is not byte-compatible with ASCII. ASCII uses one byte per character, while UTF-16 uses a 16-bit code unit for the same character.
Character:
A
ASCII:
41
UTF-16BE:
00 41
UTF-16LE:
41 00The Unicode code point for A is U+0041, but UTF-16 still stores it in a 16-bit code unit. This makes UTF-16 less compact than UTF-8 for ASCII-heavy text.
UTF-16 vs UTF-8
UTF-8 and UTF-16 can represent the same Unicode code points, but they use different encoding strategies. UTF-8 uses one to four bytes, while UTF-16 uses one or two 16-bit code units, equivalent to two or four bytes.
| Character | UTF-8 | UTF-16 Code Units | Typical UTF-16 Bytes |
|---|---|---|---|
| A | 1 byte | 1 | 2 |
| Я | 2 bytes | 1 | 2 |
| € | 3 bytes | 1 | 2 |
| 😀 | 4 bytes | 2 | 4 |
For ASCII text, UTF-8 is more compact. For many characters in the BMP, UTF-16 uses a predictable two bytes per code unit. Supplementary characters require four bytes in both UTF-8 and UTF-16.
UTF-16 vs UTF-32
UTF-32 uses one fixed 32-bit code unit for every Unicode code point. UTF-16 uses one 16-bit code unit for most BMP characters and two for supplementary characters.
| Character | UTF-8 | UTF-16 | UTF-32 |
|---|---|---|---|
| A | 1 byte | 2 bytes | 4 bytes |
| Я | 2 bytes | 2 bytes | 4 bytes |
| € | 3 bytes | 2 bytes | 4 bytes |
| 😀 | 4 bytes | 4 bytes | 4 bytes |
UTF-32 makes every Unicode code point the same storage size, but that predictability comes at the cost of increased memory usage for most text. UTF-16 occupies a middle position between UTF-8 and UTF-32.
UTF-16 in JavaScript
JavaScript strings are specified in terms of UTF-16 code units. This does not mean every JavaScript string is stored as a literal UTF-16 byte sequence in memory in every JavaScript engine, but the observable string model is based on UTF-16 code units.
const text = "A";
console.log(text.length); // 1For characters in the BMP, one Unicode code point commonly corresponds to one JavaScript code unit.
const text = "😀";
console.log(text.length); // 2The emoji U+1F600 is represented by the surrogate pair D83D DE00, so JavaScript's `length` reports two UTF-16 code units rather than one Unicode code point.
JavaScript Code Units vs Code Points
The difference becomes especially important when iterating over strings or calculating their length. Indexing a JavaScript string operates on UTF-16 code units.
const text = "😀";
console.log(text[0]); // high surrogate
console.log(text[1]); // low surrogateThe individual values are not useful as independent characters. They are two halves of the surrogate pair.
const text = "😀";
console.log([...text].length); // 1The string iterator understands surrogate pairs and yields the supplementary code point as one iteration result.
Getting Unicode Code Points in JavaScript
The `codePointAt()` method can read a Unicode code point rather than returning only one UTF-16 code unit.
const text = "😀";
console.log(text.charCodeAt(0).toString(16));
// d83d
console.log(text.codePointAt(0).toString(16));
// 1f600The difference illustrates the distinction between a UTF-16 code unit and the Unicode code point represented by one or more code units.
JavaScript UTF-16 Escape Sequences
JavaScript supports `\uXXXX` escapes for specifying UTF-16 code units. A supplementary character can therefore be represented using two escape sequences corresponding to its surrogate pair.
const text = "\uD83D\uDE00";
console.log(text); // 😀Modern JavaScript also supports code point escapes using braces, which are often easier to read for supplementary characters.
const text = "\u{1F600}";
console.log(text); // 😀The first form explicitly specifies the two UTF-16 surrogate code units, while the second specifies the Unicode code point directly.
UTF-16 and Java
Java's `char` type represents a 16-bit UTF-16 code unit rather than an arbitrary Unicode code point. As a result, a supplementary Unicode character can require two `char` values.
String text = "😀";
System.out.println(text.length());
// 2Java also provides APIs such as `codePointAt()` and code-point-aware iteration for applications that need to work with Unicode code points rather than raw UTF-16 code units.
UTF-16 and .NET
The .NET `System.String` type uses UTF-16 code units as its conceptual representation. The `char` type represents a 16-bit UTF-16 code unit.
string text = "😀";
Console.WriteLine(text.Length);
// 2As with JavaScript and Java, applications that need to reason about Unicode code points or user-perceived characters must distinguish those concepts from UTF-16 code-unit count.
UTF-16 and Supplementary Characters
Supplementary characters occupy Unicode code points from U+10000 through U+10FFFF. They include many emoji and characters from several historical and modern writing systems.
😀 U+1F600 → D83D DE00
🚀 U+1F680 → D83D DE80The exact surrogate pair depends on the Unicode code point. The pair must be interpreted as a unit when converting back to the original code point.
How to Decode a Surrogate Pair
To reconstruct a supplementary code point from a high surrogate H and low surrogate L, subtract the surrogate range bases and combine the resulting 10-bit values.
Code point =
0x10000
+ ((H - 0xD800) << 10)
+ (L - 0xDC00)For D83D DE00, this calculation produces U+1F600.
Unpaired Surrogates
A valid surrogate pair consists of a high surrogate followed by a low surrogate. An isolated high surrogate or isolated low surrogate does not represent a valid Unicode scalar value by itself.
Valid:
D83D DE00
Invalid as a scalar value:
D83D
Invalid as a scalar value:
DE00Programming languages and APIs differ in how they handle strings containing unpaired surrogates. Some allow them internally, while strict Unicode processing may reject them or replace them during encoding.
UTF-16 and Invalid Byte Sequences
UTF-16 bytes cannot be interpreted correctly without knowing their byte order. Reading UTF-16LE data as UTF-16BE, or vice versa, changes the numerical value of every 16-bit code unit.
Correct UTF-16LE:
41 00 → U+0041 → A
Incorrectly interpreted as UTF-16BE:
41 00 → U+4100This is one reason byte order must be explicitly defined or reliably detected when working with raw UTF-16 files or network data.
Common UTF-16 Encoding Problems
| Problem | Likely Cause | What to Check |
|---|---|---|
| Text looks corrupted | Wrong byte order | UTF-16LE vs UTF-16BE |
| Unexpected characters at the beginning | BOM interpreted as text | FF FE or FE FF |
| Emoji breaks during processing | Surrogate pair split | High and low surrogate handling |
| JavaScript length seems too large | UTF-16 code-unit counting | Surrogate pairs |
| Character indexing breaks emoji | Code-unit indexing | Use code-point-aware iteration |
| File opens incorrectly | Encoding mismatch | Declared encoding and actual bytes |
| Conversion produces invalid text | Unpaired surrogate | Source string validity |
UTF-16 and the BOM Problem
A UTF-16 BOM is useful for identifying byte order, but software must know whether the BOM is metadata or actual content. A decoder that correctly handles the encoding should generally consume the BOM rather than expose it as an ordinary character.
UTF-16LE file:
FF FE 48 00 69 00
BOM:
FF FE
Content:
48 00 69 00If the BOM is accidentally preserved as U+FEFF inside application data, it can behave like an invisible character and cause unexpected comparisons or parsing behavior.
UTF-16 and Invisible Characters
The BOM character U+FEFF is historically associated with the byte-order mark, but Unicode also defines its use as a zero-width no-break space in older contexts. Modern text processing should generally treat U+FEFF at the beginning of a stream as a BOM rather than ordinary content.
Other invisible Unicode characters can also appear in UTF-16 text. These characters can make strings that look identical behave differently in comparisons, searches, or validation.
UTF-16 in Files
UTF-16 can be used for text files, but a decoder must know whether the file is UTF-16LE or UTF-16BE. Some formats specify the encoding explicitly, while others rely on a BOM or external metadata.
UTF-16LE:
FF FE
48 00
65 00
6C 00
6C 00
6F 00The repeated zero bytes in ASCII-heavy UTF-16LE text make its representation visibly different from UTF-8. This can also make UTF-16 files less compact for English text.
UTF-16 in Web Development
Developers encounter UTF-16 frequently in web applications because JavaScript strings use the UTF-16 code-unit model. However, web protocols and files commonly use UTF-8 for serialized text and network transport.
This means a web application can manipulate text using a UTF-16-oriented string model while receiving and sending data encoded as UTF-8. The runtime converts between these representations at API boundaries.
const text = "Привет 😀";
const utf8 = new TextEncoder().encode(text);
console.log(text.length);
console.log(utf8.length);The string's JavaScript length and its UTF-8 byte length answer different questions. Confusing the two can lead to incorrect limits, truncation, or buffer calculations.
UTF-16 and Unicode Escapes
Unicode escape notation often appears in programming languages and data formats. A four-digit `\uXXXX` escape represents a 16-bit value, which is particularly relevant to UTF-16-based string models.
const letter = "\u0041";
const emoji = "\uD83D\uDE00";
console.log(letter); // A
console.log(emoji); // 😀The emoji example contains two four-digit escapes because U+1F600 cannot fit into one 16-bit UTF-16 code unit.
Unicode Code Point Escapes vs UTF-16 Escapes
Modern JavaScript provides two useful forms of Unicode escape notation. `\uXXXX` specifies a 16-bit value, while `\u{...}` specifies a Unicode code point.
const a = "\uD83D\uDE00";
const b = "\u{1F600}";
console.log(a === b); // trueThe first form explicitly describes the surrogate pair, while the second directly identifies the Unicode code point. The JavaScript engine produces the same logical string.
UTF-16 and String Truncation
A common bug occurs when code truncates strings by UTF-16 code-unit index. Cutting between the high and low surrogates can create an incomplete surrogate pair.
const text = "A😀B";
console.log(text.slice(0, 2));
// May contain "A" plus only part of the emojiThe exact result depends on the selected range, but the underlying issue is that `slice()` indexes UTF-16 code units rather than Unicode grapheme clusters.
UTF-16 and User-Perceived Characters
Even counting Unicode code points is not always equivalent to counting characters as users perceive them. A visible symbol can consist of multiple code points, and each code point can contain one or two UTF-16 code units.
const text = "👨💻";
console.log(text.length);
console.log([...text].length);The displayed emoji can be a sequence involving multiple Unicode code points connected by zero-width joiners. A grapheme-aware algorithm is required when the application needs to count user-perceived characters accurately.
UTF-16 and Memory Usage
UTF-16 requires at least two bytes per code unit. This makes it relatively predictable for BMP-heavy text, but ASCII text generally takes twice as much space as its UTF-8 representation.
| Text | UTF-8 | UTF-16 |
|---|---|---|
| A | 1 byte | 2 bytes |
| Я | 2 bytes | 2 bytes |
| € | 3 bytes | 2 bytes |
| 😀 | 4 bytes | 4 bytes |
UTF-16 can therefore be smaller than UTF-8 for some BMP characters, while UTF-8 is generally smaller for ASCII-heavy content.
When UTF-16 Can Be Useful
UTF-16 can be convenient in systems whose internal string representation is already based on 16-bit code units. It provides direct representation for every BMP code point and a standardized surrogate-pair mechanism for supplementary characters.
- Environments with UTF-16-based string APIs
- Legacy Windows APIs and file formats
- Java and .NET string representations
- JavaScript string processing
- Applications interoperating with systems that explicitly require UTF-16
For new network protocols and web-oriented text interchange, UTF-8 is often preferred, but compatibility requirements may make UTF-16 the appropriate choice in specific systems.
When UTF-16 Can Be Problematic
UTF-16 introduces several concepts that developers must handle correctly: two-byte code units, byte order, BOMs, surrogate pairs, and the difference between code-unit count and code-point count.
- ASCII-heavy text generally requires more bytes than UTF-8.
- Raw data requires explicit handling of endianness.
- Supplementary characters require surrogate pairs.
- Code-unit indexing can split a surrogate pair.
- A BOM can be mishandled as content.
- String length can differ from Unicode code-point count.
- Different APIs may handle unpaired surrogates differently.
UTF-16 Encoding Workflow
A simplified UTF-16 encoding process starts with Unicode code points and determines whether each value belongs to the BMP or supplementary range.
Unicode code point
↓
Is it in the BMP?
↓
Yes → one 16-bit code unit
No
↓
Subtract U+10000
↓
Split into two 10-bit values
↓
High surrogate + low surrogate
↓
Two 16-bit code unitsAfter code units are produced, they can be serialized into bytes according to either big-endian or little-endian byte order.
UTF-16 Decoding Workflow
Decoding reverses the process. Raw bytes are first interpreted according to the correct byte order. The resulting 16-bit code units are then examined for surrogate pairs.
Bytes
↓
Determine byte order
↓
Read 16-bit code units
↓
High surrogate followed by low surrogate?
↓
Yes → combine into one supplementary code point
No → interpret as a BMP code pointCommon UTF-16 Mistakes
- Confusing UTF-16 code units with Unicode code points
- Assuming every Unicode character uses one 16-bit unit
- Ignoring surrogate pairs
- Treating a high surrogate as an independent character
- Splitting a string in the middle of a surrogate pair
- Reading UTF-16LE as UTF-16BE
- Reading UTF-16BE as UTF-16LE
- Treating the BOM as ordinary text
- Assuming `string.length` always means number of Unicode characters
- Assuming UTF-16 is byte-compatible with ASCII
- Confusing `\uXXXX` escapes with UTF-8 bytes
- Assuming every visible character corresponds to exactly one code point
- Ignoring normalization and combining characters
How to Debug UTF-16 Problems
The most useful debugging technique is to inspect the representation at each layer. Determine whether the problem occurs in the raw bytes, byte order, code units, surrogate-pair processing, or final rendering.
- Check whether the source is actually UTF-16.
- Determine whether the data is UTF-16LE or UTF-16BE.
- Inspect the first bytes for FF FE or FE FF.
- Group the data into 16-bit code units.
- Look for valid high-surrogate and low-surrogate pairs.
- Check whether an operation splits a surrogate pair.
- Compare UTF-16 code-unit length with Unicode code-point count.
- Check for invisible characters such as U+FEFF.
- Verify that the receiving application expects the same encoding.
A UTF-16 inspector is particularly useful for examining code units and byte order. A BOM detector can identify the presence and type of byte-order mark, while a Unicode escape converter can help compare escape sequences with their corresponding code points.
UTF-16 Validation
Validating UTF-16 requires more than checking whether the byte count is even. The bytes must first be interpreted in the correct byte order, and the resulting code units must follow the rules for surrogate pairs.
An isolated surrogate can indicate malformed data, depending on the context and the representation being validated. Strict Unicode processing should distinguish valid Unicode scalar values from unpaired surrogate code units.
UTF-16 and Security
Encoding inconsistencies can become security problems when different components interpret the same data differently. For example, one component may process UTF-16 code units while another assumes that each code unit is a complete character.
Security-sensitive applications should establish clear boundaries between byte decoding, Unicode processing, validation, normalization, and application-level parsing.
UTF-16 and Normalization
UTF-16 defines how Unicode code points are represented using 16-bit code units. It does not determine whether canonically equivalent sequences should be normalized.
const a = "é";
const b = "e\u0301";
console.log(a === b); // false
console.log(
a.normalize("NFC") === b.normalize("NFC")
); // trueThe example contains the same visible accented character but different underlying code-point sequences. UTF-16 can encode both sequences; normalization is a separate Unicode operation.
UTF-16 vs UTF-8 vs UTF-32 at a Glance
| Property | UTF-8 | UTF-16 | UTF-32 |
|---|---|---|---|
| Basic unit | 8 bits | 16 bits | 32 bits |
| Bytes per unit | 1 | 2 | 4 |
| Variable width | 1–4 bytes | 1–2 code units | Fixed 4 bytes |
| ASCII compatibility | Yes | No | No |
| Supplementary characters | 4 bytes | 2 code units | 4 bytes |
| Byte order issue | No | Yes | Yes |
| BOM commonly relevant | Optional | Useful for byte order | Useful for byte order |
| Common web transport | Very common | Less common | Rare |
UTF-16 and UTF-8 Conversion
Converting between UTF-8 and UTF-16 should conceptually pass through Unicode text or code points. One encoding should not be interpreted directly as the other.
UTF-8 bytes
↓
UTF-8 decoding
↓
Unicode text
↓
UTF-16 encoding
↓
UTF-16 code units
↓
Byte serializationThe reverse conversion follows the same principle in the opposite direction. Decode the source encoding correctly, then encode the resulting Unicode text using the destination encoding.
UTF-16 and Byte Length
Because every UTF-16 code unit is 16 bits, the serialized representation normally uses two bytes per code unit. BMP characters generally require two bytes, while supplementary characters require four bytes.
A
1 code unit
2 UTF-16 bytes
Я
1 code unit
2 UTF-16 bytes
😀
2 code units
4 UTF-16 bytesThis makes UTF-16 byte length easier to estimate than UTF-8 byte length once the number of code units is known, but determining code units correctly still requires handling surrogate pairs.
UTF-16 and Character Counting
There are several different quantities that developers may call a character count: UTF-16 code units, Unicode code points, grapheme clusters, and encoded bytes. These counts can all differ.
| Measurement | Example: 😀 |
|---|---|
| UTF-16 code units | 2 |
| Unicode code points | 1 |
| UTF-8 bytes | 4 |
| UTF-16 bytes | 4 |
| Visible grapheme clusters | 1 |
Choosing the correct measurement depends on the problem. A network protocol may care about bytes, a JavaScript API may expose code units, and a user-interface character limit may need grapheme clusters.
Testing UTF-16 Correctly
Testing only ASCII characters is not enough to verify correct UTF-16 handling. A useful test set should contain BMP characters, accented characters, characters near the surrogate boundary, and supplementary characters.
ASCII:
Hello
BMP:
Привет
€
漢字
Supplementary:
😀
🚀
Combining sequence:
e + combining acute accentTests should also cover both UTF-16LE and UTF-16BE when an application reads raw UTF-16 data, as well as files with and without a BOM when both forms are permitted.
A Practical UTF-16 Checklist
- Determine whether the data is UTF-16LE or UTF-16BE.
- Check for a BOM when applicable.
- Remember that each UTF-16 code unit is 16 bits.
- Do not confuse code units with Unicode code points.
- Handle surrogate pairs as a pair.
- Do not split supplementary characters by code-unit index.
- Use code-point-aware APIs when code points are required.
- Use grapheme-aware processing when user-perceived characters are required.
- Inspect raw bytes when debugging encoding issues.
- Check for unpaired surrogates in malformed input.
- Do not assume UTF-16 is ASCII-compatible.
- Keep UTF-8, UTF-16, and UTF-32 conversions explicit.
Frequently Asked Questions
What is UTF-16 in simple terms?
UTF-16 is a Unicode encoding that represents text using 16-bit code units. Most BMP characters use one code unit, while supplementary characters use two code units in a surrogate pair.
How many bytes does UTF-16 use per character?
A BMP code point generally requires two bytes, while a supplementary code point requires four bytes because it uses two 16-bit code units. The exact concept is code units rather than characters.
What is a UTF-16 surrogate pair?
A surrogate pair consists of one high surrogate and one low surrogate. Together they represent a Unicode code point from U+10000 through U+10FFFF.
Why does JavaScript report emoji as length 2?
JavaScript strings use a UTF-16 code-unit model. An emoji such as U+1F600 is represented by two UTF-16 code units, so `length` reports 2 even though the emoji is one Unicode code point.
What is the UTF-16 BOM?
A UTF-16 BOM is the byte sequence FF FE for little-endian data or FE FF for big-endian data. It can identify both UTF-16 encoding and byte order.
What is the difference between UTF-16LE and UTF-16BE?
They use different byte orders for each 16-bit code unit. UTF-16LE stores the low byte first, while UTF-16BE stores the high byte first.
Is UTF-16 better than UTF-8?
They solve the same general problem using different representations. UTF-8 is usually more compact for ASCII-heavy text and is widely used for web data, while UTF-16 is convenient in systems whose string model is based on 16-bit code units.
Can UTF-16 represent every Unicode character?
Yes. Valid Unicode scalar values through U+10FFFF can be represented using either one UTF-16 code unit or a surrogate pair of two code units.
Helpful UTF-16 Tools
A UTF-16 inspector is useful for examining 16-bit code units, surrogate pairs, and byte order. A UTF-8 inspector can help compare the same Unicode text across different encodings, while a UTF-32 inspector shows the corresponding fixed-width code-point representation.
A Unicode escape converter is useful when working with `\uXXXX` sequences and surrogate pairs. A BOM detector can quickly reveal whether a file begins with the UTF-16LE or UTF-16BE byte-order mark.
Conclusion
UTF-16 is a variable-width Unicode encoding based on 16-bit code units. Unicode code points in the BMP can generally be represented by one code unit, while supplementary code points require two code units forming a surrogate pair.
The most important concepts to remember are the difference between code points and code units, the role of surrogate pairs, and the distinction between UTF-16LE and UTF-16BE. A BOM can help identify the byte order, but UTF-16 data can also be defined as little-endian or big-endian by external metadata.
UTF-16 remains particularly relevant because several major programming environments use a UTF-16-oriented string model. Understanding it is therefore essential when working with JavaScript, Java, .NET, Windows APIs, Unicode escapes, emoji, and systems that exchange UTF-16 encoded data.