Byte Order Mark (BOM) Explained
A practical guide to Byte Order Marks: what BOM is, how it works in UTF-8, UTF-16 and UTF-32, why byte order matters, and how to detect and handle BOM correctly.
A Byte Order Mark (BOM) is a special sequence of bytes placed at the beginning of a text file. It can help a program identify the encoding of the file and, for encodings such as UTF-16 and UTF-32, determine the byte order used to represent multi-byte values.
BOM is most commonly encountered when working with UTF-8, UTF-16, or UTF-32 files. In UTF-8, the BOM does not indicate byte order because UTF-8 has only one byte order. Instead, it can act as an encoding signature. In UTF-16 and UTF-32, the BOM can distinguish between little-endian and big-endian representations.
Although a BOM is invisible when text is displayed normally, it is part of the underlying byte sequence. This is why a file can look perfectly normal in an editor while scripts, parsers, command-line tools, or APIs still behave unexpectedly.
What Is a Byte Order Mark?
A Byte Order Mark is a sequence of bytes at the beginning of a Unicode text stream. Its original purpose is to provide information about the byte order of encoded data. The Unicode standard defines BOM sequences for UTF-16 and UTF-32, while UTF-8 has a defined byte sequence that can be used as an encoding signature.
The name comes from byte order, also called endianness. Some Unicode encodings represent one logical value using multiple bytes. Those bytes can be stored in different orders. A decoder needs to know which order is being used to reconstruct the original value correctly.
For example, a UTF-16 code unit can be represented using two bytes. In little-endian order, the least significant byte comes first. In big-endian order, the most significant byte comes first. A BOM at the beginning of a stream can tell a decoder which interpretation to use.
| Encoding | BOM bytes | Purpose |
|---|---|---|
| UTF-8 | EF BB BF | Encoding signature; not needed for byte order |
| UTF-16 LE | FF FE | Indicates little-endian byte order |
| UTF-16 BE | FE FF | Indicates big-endian byte order |
| UTF-32 LE | FF FE 00 00 | Indicates little-endian byte order |
| UTF-32 BE | 00 00 FE FF | Indicates big-endian byte order |
Why Does Byte Order Matter?
Byte order matters whenever a single logical value is represented by multiple bytes. Consider the hexadecimal value 0x1234 stored using two bytes. Big-endian representation stores 12 first and 34 second. Little-endian representation stores 34 first and 12 second.
Big-endian:
12 34
Little-endian:
34 12If a decoder expects the wrong byte order, it can reconstruct a completely different value. This is particularly important for UTF-16 and UTF-32 because their basic encoded units are larger than one byte.
UTF-8 does not have this problem. Each encoded unit is an individual byte, and the order of bytes has a defined meaning within the UTF-8 sequence. There is therefore no UTF-8 little-endian or big-endian variant.
BOM and Endianness
Endianness describes the order in which the bytes of a multi-byte value are stored. The two common forms are little-endian and big-endian.
- Little-endian stores the least significant byte first.
- Big-endian stores the most significant byte first.
- UTF-16 can be encoded as UTF-16LE or UTF-16BE.
- UTF-32 can be encoded as UTF-32LE or UTF-32BE.
- UTF-8 does not have an endianness distinction.
For UTF-16 and UTF-32, a BOM provides an unambiguous signal when the encoding itself does not otherwise specify the byte order. A decoder can inspect the first bytes before processing the rest of the stream.
The UTF-8 BOM
The UTF-8 BOM consists of three bytes: EF BB BF. When interpreted as Unicode, these bytes correspond to U+FEFF.
EF BB BFUnlike UTF-16 and UTF-32, these bytes do not identify a byte order. UTF-8 has no alternative byte ordering. Instead, the UTF-8 BOM can serve as an encoding signature that helps software recognize that a file is UTF-8.
Some Windows tools and applications have historically used the UTF-8 BOM to distinguish UTF-8 text from legacy encodings. Other tools prefer UTF-8 without a BOM because UTF-8 can already be identified through metadata, file format rules, or context.
As a result, UTF-8 with BOM and UTF-8 without BOM are both common in real-world files. The correct choice depends on the format and the software consuming the file.
UTF-8 With BOM vs UTF-8 Without BOM
| Property | UTF-8 without BOM | UTF-8 with BOM |
|---|---|---|
| Encoding | UTF-8 | UTF-8 |
| Byte order | Not applicable | Not applicable |
| Initial bytes | None | EF BB BF |
| Encoding signature | No BOM signature | Yes |
| Compatibility | Widely supported | Widely supported, but not universally expected |
| Extra bytes | None | 3 bytes at the beginning |
For many modern text formats, UTF-8 without a BOM is the simplest representation. However, this does not mean that a UTF-8 BOM is automatically invalid. Compatibility depends on the specific format and parser.
The UTF-16 BOM
UTF-16 uses 16-bit code units. Because each code unit contains two bytes, byte order can matter. UTF-16 therefore has little-endian and big-endian forms.
| Encoding | BOM | Meaning |
|---|---|---|
| UTF-16LE | FF FE | Little-endian |
| UTF-16BE | FE FF | Big-endian |
The two-byte sequences FF FE and FE FF are deliberately distinct. A decoder can examine the beginning of a file and determine whether the UTF-16 code units should be interpreted in little-endian or big-endian order.
For example, the Unicode code point U+0041 represents the Latin capital letter A. A UTF-16 representation of that value consists of the 16-bit value 0041. The bytes are 41 00 in UTF-16LE and 00 41 in UTF-16BE.
UTF-16LE:
41 00
UTF-16BE:
00 41If a decoder reads UTF-16LE data as UTF-16BE, the two bytes are reversed and the resulting value is no longer U+0041. The BOM prevents this ambiguity when it is present.
The UTF-32 BOM
UTF-32 represents Unicode code points using 32-bit values. Because four bytes are used for each encoded value, byte order can also be little-endian or big-endian.
| Encoding | BOM | Meaning |
|---|---|---|
| UTF-32LE | FF FE 00 00 | Little-endian |
| UTF-32BE | 00 00 FE FF | Big-endian |
UTF-32 is straightforward to reason about because each Unicode scalar value occupies one 32-bit code unit. However, UTF-32 uses significantly more storage than UTF-8 for most text and is therefore uncommon for general web and interchange formats.
What Is U+FEFF?
The BOM is closely associated with the Unicode character U+FEFF. Historically, U+FEFF was known as ZERO WIDTH NO-BREAK SPACE and could be used as a zero-width, non-breaking character. Its use for that purpose has been deprecated in favor of WORD JOINER, U+2060.
Today, U+FEFF at the beginning of a Unicode stream is primarily interpreted as a signature indicating the encoding and, where relevant, byte order. When the same character appears elsewhere in text, it should not generally be treated as a BOM.
This distinction is important because the exact same Unicode value can be represented as the three UTF-8 bytes EF BB BF. At the beginning of a stream, those bytes may be recognized as a BOM. Inside ordinary text, an occurrence of U+FEFF has a different semantic interpretation.
BOM Is Part of the Byte Stream
One of the easiest ways to understand BOM-related problems is to remember that the BOM exists at the byte level. A text editor may hide it completely, but a hexadecimal viewer or byte inspector can reveal it immediately.
UTF-8 file without BOM:
48 65 6C 6C 6F
UTF-8 file with BOM:
EF BB BF 48 65 6C 6C 6FBoth files display the same visible text: Hello. However, the underlying byte sequences are different. A program that reads raw bytes and does not handle the BOM correctly can see the initial EF BB BF sequence as data rather than metadata.
How BOM Causes Invisible Character Problems
A BOM is usually invisible when text is rendered. This can make debugging confusing. A string can appear to start with an ordinary letter while actually containing an invisible character before it.
For example, a value copied from a UTF-8 file with a BOM might internally behave like a string beginning with U+FEFF followed by the expected text. A comparison against a normal string can then fail even though the values look identical in a UI.
const value = "\uFEFFHello";
console.log(value === "Hello");
// false
console.log(value.charCodeAt(0).toString(16));
// "feff"This is one reason invisible Unicode characters should be considered when debugging unexpected string comparisons, malformed identifiers, strange CSV headers, or values that appear to have unexplained leading whitespace.
BOM in JavaScript and TypeScript Files
JavaScript and TypeScript tooling generally handles UTF-8 source files well, including files that contain a UTF-8 BOM. Modern parsers and build tools can usually recognize and ignore a leading BOM.
However, the situation can become more complicated when source files are passed through custom scripts, generated files, shell commands, text-processing utilities, or tools that treat the file as a generic sequence of bytes.
A BOM can also matter when a program reads source or configuration files manually rather than relying on a mature parser. In those cases, the first key or first line can contain an unexpected U+FEFF character if the reader does not strip the BOM.
Detecting BOM in JavaScript
When working with a JavaScript string, U+FEFF can be detected explicitly. A simple check is useful when you need to diagnose input that may have originated from a BOM-containing file.
function hasLeadingBom(value) {
return value.startsWith("\uFEFF");
}
console.log(hasLeadingBom("\uFEFFHello"));
// true
console.log(hasLeadingBom("Hello"));
// falseWhen working with raw bytes, the UTF-8 BOM can be checked as the byte sequence EF BB BF.
function hasUtf8Bom(bytes) {
return (
bytes.length >= 3 &&
bytes[0] === 0xef &&
bytes[1] === 0xbb &&
bytes[2] === 0xbf
);
}
console.log(hasUtf8Bom(new Uint8Array([0xef, 0xbb, 0xbf, 0x48])));
// trueRemoving a UTF-8 BOM in JavaScript
If an application receives text where a leading U+FEFF should not be part of the value, it can remove the character explicitly. This should be done only when the input contract expects ordinary text rather than a meaningful U+FEFF character.
function removeLeadingBom(value) {
return value.replace(/^\uFEFF/, "");
}
const input = "\uFEFFHello";
const cleaned = removeLeadingBom(input);
console.log(cleaned);
// "Hello"Using a leading-only replacement is preferable to blindly removing every occurrence of U+FEFF from the string. A BOM is specifically associated with the beginning of a Unicode stream.
BOM and TextDecoder
The Web API TextDecoder provides a convenient way to decode byte sequences. Its behavior regarding a UTF-8 BOM is important when processing binary input in browser or server-side JavaScript environments.
const bytes = new Uint8Array([
0xef, 0xbb, 0xbf,
0x48, 0x65, 0x6c, 0x6c, 0x6f
]);
const decoder = new TextDecoder("utf-8");
const text = decoder.decode(bytes);
console.log(text);
// "Hello"The decoder can recognize the UTF-8 BOM as an encoding signature instead of returning it as ordinary text. This is one reason using a standards-aware decoder is preferable to manually converting individual bytes whenever possible.
BOM in Python
Python provides explicit encoding names for dealing with UTF-8 files that may contain a BOM. The utf-8-sig codec is designed to consume a UTF-8 BOM when reading and write one when encoding.
with open("data.txt", "r", encoding="utf-8-sig") as file:
text = file.read()
print(text)Using utf-8-sig can be useful when consuming files produced by applications that commonly add a UTF-8 BOM. If your application controls the format and does not need a BOM, plain UTF-8 is often a simpler choice.
BOM in CSV Files
CSV files are a common place to encounter UTF-8 BOMs. Some spreadsheet applications use a UTF-8 BOM as a signal that a CSV file is encoded as UTF-8, which can help when opening files containing non-ASCII characters.
However, a CSV parser that does not recognize the BOM may treat it as part of the first field. This can lead to a subtle problem where the first column name contains an invisible U+FEFF character.
Expected header:
name,email
Possible internal value:
\uFEFFnameThe result may be especially confusing when code accesses columns by exact name. A lookup for name can fail because the actual key is technically different even though both strings look identical when printed.
BOM and JSON
JSON is normally exchanged as Unicode text and is commonly encoded as UTF-8. A UTF-8 BOM is not necessary for JSON to identify itself as UTF-8. Parsers and implementations can differ in how they handle an unexpected BOM at the beginning of a JSON document.
For maximum interoperability, generated JSON is commonly emitted as UTF-8 without a BOM. This avoids making the consumer responsible for handling an additional signature before the JSON text begins.
UTF-8 JSON without BOM:
7B 22 6E 61 6D 65 22 3A 22 41 6C 69 63 65 22 7D
UTF-8 JSON with BOM:
EF BB BF 7B 22 6E 61 6D 65 22 3A 22 41 6C 69 63 65 22 7DIf a JSON parser reports an unexpected character at the very beginning of an otherwise valid document, checking for a UTF-8 BOM is one useful diagnostic step.
BOM in HTML and Web Development
Web documents are commonly served as UTF-8, and HTML can contain a BOM at the beginning of the file. Browsers are generally designed to handle it correctly.
For web applications, the more important issue is usually consistency. The server should provide the correct Content-Type and charset information, source files should use a predictable encoding, and build tools should agree on how text files are encoded.
A BOM should not be treated as a replacement for correct HTTP metadata. When a server can explicitly declare the character encoding, that metadata is more appropriate than relying on a signature embedded in the document.
BOM and HTTP Content-Type
HTTP responses can declare the media type and character encoding using the Content-Type header. For example, an HTML response can be served with a UTF-8 charset declaration.
Content-Type: text/html; charset=utf-8When encoding information is already provided reliably through the protocol or file format, adding a BOM is often unnecessary. The BOM is most useful when it solves a real encoding-identification or byte-order problem.
BOM and File Editors
Text editors may expose BOM handling through settings such as UTF-8, UTF-8 with BOM, UTF-16 LE, or UTF-16 BE. The exact labels depend on the editor.
A file can therefore change its byte representation even if the visible text does not change. Saving a UTF-8 file as UTF-8 with BOM adds EF BB BF to the beginning. Saving it back as UTF-8 without BOM removes those three bytes.
This can matter in source repositories because changing the encoding of a file may create noisy diffs or cause scripts to behave differently. For team projects, it is useful to have a consistent encoding policy.
BOM and Git
Git tracks file content as bytes. Therefore, adding or removing a BOM changes the file content from Git's perspective even if an editor displays exactly the same text.
Before:
48 65 6C 6C 6F
After adding UTF-8 BOM:
EF BB BF 48 65 6C 6C 6FDepending on the diff viewer, the change may appear as an encoding-related modification rather than an obvious textual change. If a large number of files suddenly acquire or lose BOMs, a commit can become unnecessarily noisy.
How to Detect a BOM
The most reliable way to diagnose a BOM is to inspect the first few bytes of the file. A hexadecimal view immediately reveals whether the file begins with one of the recognized BOM sequences.
| Bytes at start | Likely interpretation |
|---|---|
| EF BB BF | UTF-8 BOM |
| FF FE | UTF-16LE BOM |
| FE FF | UTF-16BE BOM |
| FF FE 00 00 | UTF-32LE BOM |
| 00 00 FE FF | UTF-32BE BOM |
A BOM Detector can automate this check and show whether the file starts with a recognized marker. A UTF-8, UTF-16, or UTF-32 inspector can then be used to examine the rest of the byte sequence.
BOM Detection With a Hexadecimal Check
If you are debugging an unknown file format, looking at the first bytes is often enough to identify a BOM.
EF BB BF 3C 68 74 6D 6C 3E
The first three bytes are a UTF-8 BOM.
The remaining bytes begin with "<html>".The same method works for UTF-16 and UTF-32. The key is to inspect the bytes before attempting to interpret the entire file using a particular encoding.
BOM vs Encoding Detection
A BOM can help with encoding detection, but it is not a universal encoding detector. Many text files contain no BOM at all, and several encodings can represent the same initial bytes.
For example, a normal ASCII-compatible UTF-8 file beginning with English text may contain no special signature. A decoder cannot always determine the intended encoding with certainty from arbitrary bytes alone.
Reliable encoding identification should therefore use the strongest available source of information: an explicit protocol declaration, file format specification, metadata, known application convention, or BOM where applicable.
BOM and Unicode Compatibility
The BOM is an encoding-level mechanism. It should not be confused with Unicode normalization. Normalization changes the representation of Unicode text according to canonical or compatibility equivalence, while a BOM identifies properties of the encoded byte stream.
For example, NFC and NFD can produce different Unicode code point sequences that represent canonically equivalent text. A BOM does not normalize those characters and does not determine whether a string is NFC or NFD.
BOM vs Zero-Width Characters
BOM-related problems are sometimes grouped together with invisible Unicode characters, but the concepts are not identical. A BOM is specifically associated with the beginning of an encoded stream. Unicode also contains other invisible or formatting characters that can occur inside ordinary text.
This distinction matters when cleaning data. Removing all invisible Unicode characters can destroy legitimate text formatting or semantics. BOM removal should normally be limited to a leading BOM when the input format calls for it.
Should You Use a BOM?
There is no universal answer for every format. The right choice depends on the encoding, file format, platform, and consuming software.
| Situation | Typical approach |
|---|---|
| UTF-8 web application source | Usually UTF-8 without BOM |
| UTF-8 JSON interchange | Usually UTF-8 without BOM |
| UTF-8 CSV for spreadsheet compatibility | BOM may be useful |
| UTF-16 file requiring byte-order identification | BOM can be useful |
| UTF-32 file requiring byte-order identification | BOM can be useful |
| Unknown legacy pipeline | Follow the consumer's requirements |
The important principle is to choose deliberately rather than treating BOM as either universally required or universally harmful.
Common BOM Problems
Most BOM-related bugs come from software that assumes a text stream starts immediately with meaningful content. Several symptoms can point toward an unexpected BOM.
- The first CSV column name does not match the expected string.
- A JSON parser reports an unexpected character at the beginning of the document.
- A configuration key works everywhere except for the first key in a file.
- A string comparison fails even though both values look identical.
- A generated file behaves differently depending on which editor saved it.
- A script works with one text file but fails with another visually identical file.
- A Git diff shows changes that are not obvious in the displayed text.
How to Troubleshoot a BOM Problem
When a text file behaves unexpectedly, start by inspecting the beginning of the byte stream rather than immediately changing application logic.
- Check the first bytes of the file.
- Identify whether a BOM is present.
- Determine the intended character encoding.
- Check whether the consuming library supports that BOM.
- Inspect the first parsed string or key for U+FEFF.
- Compare the actual bytes with a known-good file.
- Remove or preserve the BOM according to the target format.
- Add an automated encoding rule if the same problem can recur.
This approach separates encoding problems from ordinary application bugs. If the first parsed key contains U+FEFF, for example, changing business logic or database queries will not fix the underlying input problem.
BOM and Line Endings Are Different
BOM and line endings are both file-level concerns, but they solve completely different problems. A BOM identifies encoding or byte order. Line endings determine how line breaks are represented.
| Feature | BOM | Line ending |
|---|---|---|
| Purpose | Encoding signature / byte order | Represents line breaks |
| Common examples | EF BB BF, FF FE | LF, CRLF, CR |
| Location | Usually at file beginning | Between lines |
| Encoding-specific | Yes | No |
| Typical tools | BOM detector / encoding inspector | Line ending converter |
A file can have a UTF-8 BOM and CRLF line endings, or no BOM and LF line endings. These properties are independent and should be handled separately.
BOM in UTF-8, UTF-16 and UTF-32 Compared
| Property | UTF-8 | UTF-16 | UTF-32 |
|---|---|---|---|
| Basic unit | 8 bits | 16 bits | 32 bits |
| Byte order | Not applicable | LE or BE | LE or BE |
| BOM size | 3 bytes | 2 bytes | 4 bytes |
| BOM required? | No | No, but useful for byte order | No, but useful for byte order |
| Typical web usage | Very common | Less common | Rare |
| ASCII-compatible | Yes | No | No |
UTF-8 is dominant for modern web content because it is compact for ASCII-heavy text, preserves ASCII compatibility, and does not require an endianness choice. UTF-16 remains important in programming environments and some file formats, while UTF-32 is mostly useful when fixed-width code point representation is desirable.
BOM and Character Count
A BOM can affect byte counts even when it is not displayed. A UTF-8 BOM adds three bytes to the beginning of a file. This means a file's byte length can differ from what you might expect based only on its visible characters.
const text = "Hello";
const withBom = "\uFEFFHello";
console.log(new TextEncoder().encode(text).length);
// 5
console.log(new TextEncoder().encode(withBom).length);
// 8The difference is especially important when working with file sizes, binary protocols, hashes, signatures, offsets, or systems that compare raw byte sequences.
BOM and Hashes
Adding or removing a BOM changes the underlying bytes and therefore changes a cryptographic hash of the file. Two files can display exactly the same visible text but produce different SHA-256 or other hashes.
Without BOM:
48 65 6C 6C 6F
With UTF-8 BOM:
EF BB BF 48 65 6C 6C 6FThis matters for checksums, content-addressed storage, cache keys, digital signatures, reproducible builds, and file integrity verification. If a system hashes raw file bytes, BOM handling must be consistent.
BOM and Security
Invisible characters can create security and reliability issues when different components interpret text differently. A leading BOM can be harmless when handled consistently but problematic when one component removes it and another preserves it.
Security-sensitive code should therefore normalize its input assumptions and avoid silently applying inconsistent transformations. When comparing identifiers, parsing structured input, or validating text, it is useful to know whether the input may contain encoding markers or other invisible characters.
BOM Best Practices
- Prefer UTF-8 for modern web and application text unless a format requires another encoding.
- Do not add a UTF-8 BOM merely because a file contains Unicode characters.
- Use a BOM when the target format or consumer benefits from it.
- Use explicit encoding metadata whenever the surrounding protocol supports it.
- Inspect raw bytes when a text file behaves unexpectedly.
- Handle UTF-16 and UTF-32 byte order explicitly when required.
- Avoid treating every U+FEFF occurrence as a BOM.
- Keep encoding conventions consistent across a project.
- Be careful when changing BOM settings because the visible text may not change while the bytes do.
- Test generated files with the actual software that will consume them.
Frequently Asked Questions
What does BOM stand for?
BOM stands for Byte Order Mark. It is a special sequence of bytes at the beginning of a Unicode text stream that can identify the encoding and, for UTF-16 and UTF-32, indicate byte order.
What is the UTF-8 BOM?
The UTF-8 BOM is the three-byte sequence EF BB BF. It does not indicate little-endian or big-endian order because UTF-8 has no byte-order variants. Instead, it can act as an encoding signature.
Is a UTF-8 BOM required?
No. UTF-8 works correctly without a BOM. Whether a BOM is useful depends on the file format and consuming software.
What is the difference between UTF-16LE and UTF-16BE?
UTF-16LE stores the bytes of each 16-bit code unit in little-endian order, while UTF-16BE stores them in big-endian order. Their BOMs are FF FE and FE FF respectively.
Can a BOM cause parsing errors?
Yes. A parser that does not recognize or remove a leading BOM can interpret it as part of the input. This can cause problems with JSON, CSV headers, configuration keys, or exact string comparisons.
How can I check whether a file has a BOM?
Inspect the first bytes of the file. Common signatures are EF BB BF for UTF-8, FF FE for UTF-16LE, FE FF for UTF-16BE, FF FE 00 00 for UTF-32LE, and 00 00 FE FF for UTF-32BE.
Is BOM the same as an invisible Unicode character?
Not exactly. A BOM is an encoding-level marker associated with the beginning of a Unicode stream. U+FEFF can also occur as a character value, but it should not automatically be treated as a BOM everywhere it appears.
Does BOM affect file hashes?
Yes. Adding or removing a BOM changes the underlying bytes, so the resulting file has a different cryptographic hash even if the visible text appears unchanged.
Helpful Encoding Tools
When diagnosing text encoding problems, several types of tools can make invisible differences easier to identify. BOM detectors can identify UTF-8, UTF-16, and UTF-32 byte-order markers, while UTF-8, UTF-16, and UTF-32 inspectors can expose the underlying byte representation and help compare different encoding forms. Hexadecimal viewers are also useful when the raw file bytes need to be inspected directly.
Line ending converters can help separate encoding issues from differences between LF and CRLF line endings. Text comparison tools are useful for finding invisible differences between strings when two pieces of text appear identical but behave differently during processing.
Conclusion
A Byte Order Mark is a small piece of metadata at the beginning of a Unicode byte stream. Its most important historical purpose is to identify byte order in encodings such as UTF-16 and UTF-32. In UTF-8, the BOM does not describe byte order because UTF-8 has no endianness; instead, EF BB BF can serve as an encoding signature.
Understanding BOM becomes especially useful when debugging files that look correct in an editor but behave differently in code. A leading U+FEFF can affect string comparisons, CSV headers, parsers, hashes, generated files, and other byte-sensitive operations.
The practical approach is to treat encoding, BOM, and line endings as separate properties of a text file. Inspect the actual bytes when something looks suspicious, follow the requirements of the target format, and keep encoding conventions consistent across the tools that produce and consume your files.