Ctrl + K
Encoding29 min read

UTF-8 Explained

A practical guide to UTF-8 encoding, covering Unicode code points, byte sequences, ASCII compatibility, multi-byte characters, BOMs, JavaScript, HTML, APIs, and common encoding problems.

Published: 2026-10-05

UTF-8 is one of the most widely used character encodings in modern software. It is used for web pages, APIs, source code, configuration files, databases, text files, logs, and many other forms of digital text. Its main advantage is that it can represent the entire Unicode character set while remaining fully compatible with standard ASCII for the first 128 characters.

Despite its widespread use, UTF-8 is often confused with Unicode itself. Unicode defines characters and assigns them code points, while UTF-8 defines how those code points are represented as bytes. Understanding this distinction makes it much easier to understand why different characters require different numbers of bytes and why encoding mistakes can produce corrupted text.

This guide explains UTF-8 from the basic concepts to the byte-level representation, including ASCII compatibility, multi-byte sequences, Unicode code points, BOMs, JavaScript, HTTP, JSON, databases, and common debugging techniques.

What Is UTF-8?

UTF-8 stands for Unicode Transformation Format - 8-bit. It is a variable-width encoding used to represent Unicode code points as sequences of bytes.

UTF-8 uses between one and four bytes to represent a Unicode code point. Characters in the ASCII range use one byte, while characters outside that range use two, three, or four bytes depending on their code point.

Unicode Code Point RangeUTF-8 Length
U+0000–U+007F1 byte
U+0080–U+07FF2 bytes
U+0800–U+FFFF3 bytes
U+10000–U+10FFFF4 bytes

This variable-width design allows UTF-8 to represent simple English text efficiently while still supporting scripts and symbols from the entire Unicode repertoire.

UTF-8, Unicode, and Characters Are Different Concepts

Before looking at byte sequences, it is important to separate three concepts: characters, Unicode code points, and encodings.

  • A character is the text element represented to the user.
  • A Unicode code point is a numeric value assigned within the Unicode standard.
  • UTF-8 is an encoding that converts Unicode code points into bytes.
Character:
€

Unicode code point:
U+20AC

UTF-8 bytes:
E2 82 AC

The euro sign is therefore not itself a UTF-8 byte sequence. Its Unicode code point is U+20AC, and UTF-8 defines how that code point is represented as the three bytes E2 82 AC.

Why UTF-8 Uses a Variable Number of Bytes

Unicode contains characters from many writing systems, so a fixed one-byte representation would not provide enough possible values. UTF-8 solves this by using additional bytes when a code point cannot fit into the ASCII-compatible one-byte range.

ASCII character:
A → 1 byte

Latin/Cyrillic examples:
é → 2 bytes
Я → 2 bytes

Many other Unicode characters:
€ → 3 bytes

Supplementary characters:
😀 → 4 bytes

The first byte of a multi-byte UTF-8 sequence also indicates how many bytes belong to the sequence. Continuation bytes have a different bit pattern, allowing a decoder to distinguish the parts of a character.

UTF-8 and ASCII Compatibility

One of UTF-8's most important properties is its compatibility with standard ASCII. Unicode code points from U+0000 through U+007F are encoded using exactly the same byte values used by ASCII.

Character   Unicode    ASCII    UTF-8
A           U+0041     41       41
B           U+0042     42       42
0           U+0030     30       30
!           U+0021     21       21
space       U+0020     20       20

This means an ASCII text file is also valid UTF-8. A file containing only standard ASCII characters does not need any special transformation to become valid UTF-8.

💡 If a text contains only ASCII characters, its UTF-8 byte representation is identical to its ASCII byte representation. The difference becomes visible when characters outside the ASCII range are introduced.

How UTF-8 Represents One-Byte Characters

A UTF-8 sequence containing one byte represents a Unicode code point from U+0000 through U+007F. The byte is simply the value of the code point.

A
Unicode: U+0041
Decimal:  65
Hex:      41
UTF-8:    41

Because there are only 128 possible values in this range, the highest possible byte is 0x7F. Bytes from 0x80 upward cannot be interpreted as standalone ASCII-compatible UTF-8 characters.

How UTF-8 Represents Two-Byte Characters

Unicode code points from U+0080 through U+07FF require two UTF-8 bytes. The first byte identifies the sequence as a two-byte character, while the second byte is a continuation byte.

Я
Unicode code point: U+042F
UTF-8 bytes:        D0 AF

Many characters from Latin-based alphabets with accents and many Cyrillic characters fall into this range. This is one reason UTF-8 remains relatively compact for European-language text.

How UTF-8 Represents Three-Byte Characters

Code points from U+0800 through U+FFFF generally require three UTF-8 bytes, excluding the surrogate range, which is not encoded directly as UTF-8.

€
Unicode code point: U+20AC
UTF-8 bytes:        E2 82 AC

Many characters used by writing systems around the world fall into the three-byte range. This includes a large portion of the Basic Multilingual Plane.

How UTF-8 Represents Four-Byte Characters

Unicode code points from U+10000 through U+10FFFF require four UTF-8 bytes. This range includes many supplementary characters, including a large number of emoji.

😀
Unicode code point: U+1F600
UTF-8 bytes:        F0 9F 98 80

The four-byte form is necessary because the Unicode code point does not fit into the smaller UTF-8 sequence sizes.

The Structure of UTF-8 Bytes

UTF-8 uses specific bit patterns to distinguish one-byte, two-byte, three-byte, and four-byte sequences. The exact structure is useful when inspecting raw bytes or implementing a decoder.

SequenceByte Pattern
1 byte0xxxxxxx
2 bytes110xxxxx 10xxxxxx
3 bytes1110xxxx 10xxxxxx 10xxxxxx
4 bytes11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The `x` positions contain bits from the Unicode code point. Continuation bytes always begin with `10`, while the first byte contains a prefix identifying the total sequence length.

UTF-8 Encoding Example: The Euro Sign

The euro sign has Unicode code point U+20AC. Since this value is greater than U+07FF but within the three-byte range, UTF-8 represents it using three bytes.

Code point:
U+20AC

Binary value:
0010 0000 1010 1100

UTF-8:
1110xxxx 10xxxxxx 10xxxxxx

Result:
11100010 10000010 10101100

Hex:
E2 82 AC

The example shows that UTF-8 does not simply store the hexadecimal code point value directly. It distributes the code point's bits across the appropriate UTF-8 byte structure.

UTF-8 Encoding Example: An ASCII Character

The letter A provides the simplest example because its Unicode code point is U+0041 and it belongs to the ASCII range.

U+0041
Binary: 1000001

UTF-8:
01000001

Hex:
41

No additional bytes are required because the code point fits entirely within the seven data bits available in a one-byte UTF-8 sequence.

UTF-8 Encoding Example: An Emoji

The grinning face emoji has code point U+1F600, which is outside the Basic Multilingual Plane. UTF-8 therefore uses four bytes.

Character:
😀

Code point:
U+1F600

UTF-8:
F0 9F 98 80

This is one reason byte-oriented operations can produce surprising results when processing Unicode text. One visible symbol may occupy four UTF-8 bytes even though it appears to the user as a single character.

UTF-8 Does Not Mean One Character Equals One Byte

A common misconception is that UTF-8 stores every character using one byte. It does not. Only the ASCII-compatible range uses one byte. Other Unicode code points require multiple bytes.

Text:
ABC

UTF-8 bytes:
41 42 43

Text:
Привет

UTF-8 uses multiple bytes per character.

This distinction matters when calculating byte lengths, allocating buffers, processing network data, or comparing a string's character count with its encoded size.

UTF-8 Byte Length vs Character Count

The number of characters in a string and the number of UTF-8 bytes required to encode it are different measurements.

const text = "Hello";

const bytes = new TextEncoder().encode(text);

console.log(text.length); // 5
console.log(bytes.length); // 5

For ASCII text, the values happen to be equal because every character requires one UTF-8 byte. With non-ASCII text, they can differ.

const text = "Привет";

const bytes = new TextEncoder().encode(text);

console.log(text.length); // 6
console.log(bytes.length); // 12

Each Cyrillic character in this example requires two UTF-8 bytes, so six characters require twelve bytes.

UTF-8 vs UTF-16

UTF-8 and UTF-16 are both Unicode encodings, but they represent code points differently. UTF-8 uses one to four bytes, while UTF-16 uses one or two 16-bit code units for a Unicode code point.

PropertyUTF-8UTF-16
Basic ASCII character1 byte1 code unit
Many European characters1–2 bytes1 code unit
Many BMP characters3 bytes1 code unit
Supplementary character4 bytes2 code units
ASCII compatibilityYesNo
Common web encodingYesLess common

The right encoding depends on the environment and requirements, but UTF-8 is particularly convenient for web data and interoperability because ASCII data remains byte-compatible.

UTF-8 vs UTF-32

UTF-32 represents every Unicode code point using a fixed 32-bit unit. This makes indexing by code point conceptually straightforward, but it uses considerably more memory for ordinary text.

CharacterUTF-8UTF-32
A1 byte4 bytes
Я2 bytes4 bytes
€3 bytes4 bytes
😀4 bytes4 bytes

UTF-32 can simplify certain internal processing tasks, but UTF-8 is generally much more space-efficient for storage and transmission.

Why UTF-8 Is Efficient for English

English text is mostly composed of ASCII characters, and every ASCII character requires exactly one UTF-8 byte. This gives UTF-8 the same compact representation for English text that older ASCII-oriented systems expect.

Hello, world!

UTF-8 bytes:
48 65 6C 6C 6F 2C 20
77 6F 72 6C 64 21

The fact that UTF-8 can efficiently represent English while also supporting international text is one of its major practical advantages.

UTF-8 and Cyrillic Text

Cyrillic characters are outside the ASCII range, so they require multiple UTF-8 bytes. For many Cyrillic characters, the UTF-8 representation consists of two bytes.

П → U+041F → D0 9F
р → U+0440 → D1 80
и → U+0438 → D0 B8
в → U+0432 → D0 B2
е → U+0435 → D0 B5
т → U+0442 → D1 82

This is why a Cyrillic string can require roughly twice as many UTF-8 bytes as characters even though the visible text appears to contain one character per letter.

UTF-8 and Emoji

Many emoji are represented by Unicode code points above U+FFFF and therefore require four UTF-8 bytes per code point. More complex emoji can contain multiple code points combined into one displayed sequence.

😀
U+1F600
F0 9F 98 80

This means that the number of visible emoji is not necessarily equal to the number of Unicode code points or UTF-8 bytes. A displayed emoji sequence can contain multiple code points, including variation selectors and zero-width joiners.

UTF-8 and Unicode Combining Characters

Unicode can represent some visible text using a base character followed by one or more combining characters. UTF-8 encodes each code point separately.

const text = "e\u0301";
const bytes = new TextEncoder().encode(text);

console.log([...bytes]);

The displayed result can look like a single accented character even though the string contains more than one Unicode code point. UTF-8 does not combine those code points into one conceptual character; it simply encodes each code point according to the UTF-8 rules.

UTF-8 and the Basic Multilingual Plane

Unicode divides its code space into ranges called planes. The Basic Multilingual Plane covers code points from U+0000 through U+FFFF. Many commonly used characters are located there.

UTF-8 represents most BMP characters using one, two, or three bytes. Characters outside the BMP use four-byte UTF-8 sequences.

Unicode RangeTypical UTF-8 Representation
U+0000–U+007F1 byte
U+0080–U+07FF2 bytes
U+0800–U+FFFF3 bytes
U+10000–U+10FFFF4 bytes

Invalid UTF-8 Sequences

UTF-8 has strict structural rules. Not every arbitrary sequence of bytes is valid UTF-8. A decoder must verify that continuation bytes appear in the correct positions and that the resulting code point is within the permitted Unicode range.

Valid ASCII:
41

Valid UTF-8:
E2 82 AC

Invalid example:
E2 28 AC

The final example contains a byte in a position where a UTF-8 continuation byte is required. A strict UTF-8 decoder should reject such a sequence rather than interpreting it as a valid character.

Overlong UTF-8 Encodings

An overlong encoding represents a code point using more bytes than necessary. Modern UTF-8 processing must reject overlong forms because allowing multiple byte representations of the same character creates ambiguity and security problems.

For example, an ASCII character must use its one-byte representation. It must not be encoded using a multi-byte sequence that could otherwise represent the same value.

⚠️ A UTF-8 decoder should reject malformed and overlong sequences. Accepting non-standard byte representations can create inconsistencies between components and has historically contributed to security vulnerabilities.

UTF-8 and Surrogate Code Points

UTF-16 uses surrogate code units to represent supplementary Unicode code points. UTF-8 does not encode surrogate code points directly. The surrogate range U+D800 through U+DFFF is reserved for UTF-16 surrogate handling.

When working with UTF-8 at the byte level, a valid encoder produces byte sequences corresponding to valid Unicode scalar values rather than encoding UTF-16 surrogate code points as ordinary UTF-8 characters.

UTF-8 and the Byte Order Mark

A byte order mark, or BOM, is the byte sequence EF BB BF when it appears at the beginning of a UTF-8 file. It corresponds to the Unicode character U+FEFF.

UTF-8 BOM:
EF BB BF

Unlike UTF-16 and UTF-32, UTF-8 does not need a BOM to determine byte order because UTF-8 is byte-oriented and has no multi-byte byte-order ambiguity. A UTF-8 BOM may still be used as an encoding signature in some environments.

💡 If a UTF-8 file unexpectedly starts with an invisible character, check whether it contains the UTF-8 BOM sequence EF BB BF.

UTF-8 BOM vs UTF-8 Without BOM

VariantBeginning BytesTypical Meaning
UTF-8 without BOMActual content starts immediatelyCommon modern form
UTF-8 with BOMEF BB BF before contentEncoding signature or compatibility marker

Both forms contain UTF-8 encoded text. Whether a BOM should be present depends on the file format, platform, and compatibility requirements. Applications should not assume that every UTF-8 document begins with EF BB BF.

UTF-8 in HTML

UTF-8 is the standard practical choice for modern HTML documents. The encoding can be declared using a meta element near the beginning of the document.

<!doctype html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>UTF-8 Example</title>
</head>
<body>
  <p>Привет, мир!</p>
</body>
</html>

The document's actual bytes should agree with the declared encoding. Declaring UTF-8 while serving bytes encoded using another character set can still produce corrupted text.

UTF-8 in HTTP

HTTP responses can communicate media types and character encoding through the Content-Type header. For text-based responses, the server and client need a consistent understanding of how the bytes should be decoded.

Content-Type: text/html; charset=UTF-8

For JSON APIs, UTF-8 is the normal encoding used for interoperable text exchange. The important principle is that the producer and consumer must agree on the representation of the response bytes.

UTF-8 in JSON

JSON supports Unicode text and is commonly transported using UTF-8. Unicode characters can appear directly in a JSON document or certain characters can be represented using Unicode escape sequences.

{
  "message": "Привет, мир!",
  "currency": "€"
}

A JSON escape such as `\u20AC` is not itself a UTF-8 byte sequence. It is a textual escape representation of a Unicode value. If the JSON document is stored as UTF-8, those characters or escape characters are then encoded as UTF-8 bytes.

UTF-8 in JavaScript

JavaScript works with Unicode text, but JavaScript strings are based on UTF-16 code units. This is separate from how a string is encoded when it is transmitted or stored as UTF-8.

const text = "Привет";

const bytes = new TextEncoder().encode(text);

console.log(text.length);
console.log(bytes.length);

The first value describes the JavaScript string in terms of UTF-16 code units, while the second value is the number of UTF-8 bytes produced by `TextEncoder`.

const text = "😀";

console.log(text.length); // 2

const bytes = new TextEncoder().encode(text);

console.log(bytes.length); // 4

The emoji occupies two UTF-16 code units in JavaScript but four UTF-8 bytes when encoded for transmission or storage.

UTF-8 Encoding With TextEncoder

The Web API `TextEncoder` provides a convenient way to encode JavaScript strings into UTF-8 bytes.

const encoder = new TextEncoder();

const bytes = encoder.encode("Hello, 世界!");

console.log([...bytes]);

The resulting `Uint8Array` contains the actual UTF-8 byte values. This is useful when working with binary protocols, files, cryptographic operations, network payloads, or other APIs that operate on bytes.

Decoding UTF-8 in JavaScript

The corresponding `TextDecoder` API can decode UTF-8 bytes back into a JavaScript string.

const bytes = new Uint8Array([
  0xD0,
  0x9F,
  0xD1,
  0x80,
  0xD0,
  0xB8,
  0xD0,
  0xB2,
  0xD0,
  0xB5,
  0xD1,
  0x82,
]);

const decoder = new TextDecoder("utf-8");

console.log(decoder.decode(bytes));
// Привет

Encoding converts text into bytes, while decoding performs the reverse operation. Both sides need to agree on UTF-8 for the original text to be reconstructed correctly.

UTF-8 in Databases

Modern databases can store Unicode text, but correct behavior depends on the database engine, column type, connection settings, client library, and server configuration. An application can still experience corrupted text if one layer interprets the bytes differently from another.

When investigating database encoding problems, inspect the entire path rather than checking only the database column. The application input, connection, database storage, query result, and output encoding all need to be compatible.

💡 A character-encoding problem is often a pipeline problem. Check the encoding at the point where text enters the system, where it is stored, and where it is decoded or displayed.

UTF-8 in Source Code

Most modern development environments support UTF-8 source files. This allows source code and comments to contain international characters, although language-specific identifier rules may still impose restrictions on what can be used as an identifier.

const greeting = "Привет, мир!";

console.log(greeting);

Even when source files contain Unicode characters, the development tools ultimately read bytes and decode them using an encoding. Keeping the editor, compiler, build system, and repository configuration consistent prevents many avoidable problems.

UTF-8 in Configuration Files

Configuration files frequently contain URLs, names, comments, localized values, or other text. UTF-8 is a practical default for many modern formats, but the exact file specification should always be followed.

Problems can appear when one tool writes UTF-8 with a BOM and another tool expects UTF-8 without one, or when a legacy tool assumes a different code page. Inspecting the actual file bytes can reveal these differences quickly.

UTF-8 and File Extensions

A file extension such as `.txt`, `.json`, or `.html` does not by itself determine the encoding of the file's bytes. The file format, metadata, application configuration, and encoding declarations can all influence how the content is interpreted.

⚠️ Renaming a file does not convert its encoding. If a file needs to be converted from one character encoding to UTF-8, the bytes must be decoded using the original encoding and then encoded as UTF-8.

How to Convert Text to UTF-8 Correctly

Converting a file to UTF-8 is a decode-then-encode operation. First, the original bytes must be interpreted using their actual source encoding. The resulting characters are then encoded as UTF-8.

Original bytes
      ↓
Decode using source encoding
      ↓
Unicode text
      ↓
Encode as UTF-8
      ↓
UTF-8 bytes

If the original encoding is identified incorrectly, conversion can produce corrupted text. Therefore, determining the source encoding is an important part of a reliable conversion process.

Common UTF-8 Encoding Problems

ProblemLikely CauseWhat to Check
Garbled charactersWrong decodingActual source encoding and declared encoding
Unexpected invisible character at file startUTF-8 BOMFirst three bytes: EF BB BF
Emoji displayed incorrectlyIncorrect Unicode handlingCode points, decoder, and application support
Cyrillic appears corruptedWrong code page or encodingSource bytes and response encoding
Byte length is larger than character countMulti-byte UTF-8 representationUTF-8 bytes per code point
Works locally but fails elsewhereDifferent encoding configurationEditor, server, database, and runtime settings
JSON contains escaped charactersUnicode escape representationDistinguish escapes from UTF-8 bytes

Mojibake: What It Is and Why It Happens

Mojibake is text that appears corrupted because bytes were decoded using the wrong character encoding. The original bytes may still be intact; the problem is the interpretation applied to them.

A typical example occurs when UTF-8 bytes are mistakenly interpreted as a legacy single-byte encoding. The resulting characters can look completely unrelated to the original text.

The correct fix is usually not to manually replace every corrupted character. Instead, identify the original bytes and decode them using the correct encoding. If the data has already been re-encoded incorrectly and the original bytes are lost, recovery can be more complicated.

How to Debug a UTF-8 Problem

The fastest way to debug an encoding issue is to inspect the data at the byte level. Do not rely only on how the text looks on screen because the visual representation may hide the actual problem.

  • Identify the exact text that is displayed incorrectly.
  • Determine the expected Unicode characters.
  • Inspect the raw bytes if possible.
  • Check whether the bytes form valid UTF-8.
  • Check for a UTF-8 BOM at the beginning of a file.
  • Verify the declared encoding.
  • Verify the encoding used by the decoder.
  • Check HTTP headers when working with network responses.
  • Check database and connection settings for stored text.
  • Compare byte length with expected UTF-8 length.

A UTF-8 inspector can help reveal the exact byte sequence behind a string. A BOM detector can identify an unexpected byte-order mark, while an invisible-character detector can expose characters that are present but difficult to see.

UTF-8 Validation

Valid UTF-8 must follow the structural rules of the encoding. A validator can check whether byte sequences contain valid leading and continuation bytes, whether the encoded code point is legal, and whether the sequence uses the shortest valid representation.

Validation is particularly important when processing binary input from untrusted sources. A system should not assume that arbitrary bytes are valid UTF-8 merely because they are intended to contain text.

UTF-8 and Security

Encoding differences can become security issues when different components interpret the same input differently. If a proxy, web server, application, validator, and database disagree about how bytes should be decoded, an attacker may be able to exploit the discrepancy.

Strict UTF-8 validation, consistent decoding, and clear boundaries between bytes and text help reduce these risks. Applications should decode input once using a well-defined encoding and then perform validation on the resulting text.

⚠️ Do not validate raw bytes using assumptions about their textual meaning and then decode them using a different interpretation. Security-sensitive systems should establish a consistent decoding and normalization strategy before validation.

UTF-8 and Invisible Characters

Unicode contains characters that may have little or no visible appearance. Examples include zero-width characters, non-breaking spaces, variation selectors, and certain formatting characters.

const a = "hello";
const b = "hel\u200Blo";

console.log(a === b); // false

The strings can look almost identical while containing different Unicode code points. These characters can cause unexpected search results, validation failures, duplicate identifiers, or confusing source-code differences.

UTF-8 and Unicode Normalization

UTF-8 encodes Unicode code points, but it does not decide whether canonically equivalent text should use one particular sequence of code points. Unicode normalization addresses that separate problem.

const a = "é";
const b = "e\u0301";

console.log(a === b); // false

console.log(
  a.normalize("NFC") === b.normalize("NFC")
); // true

Normalization can be important for comparison and searching, but it should be applied deliberately. It is not simply another form of UTF-8 encoding.

UTF-8 and Storage Size

UTF-8 uses one to four bytes per Unicode code point. This means storage requirements depend on the actual text rather than simply the number of characters.

Text TypeTypical UTF-8 Size
ASCII characters1 byte per code point
Many Latin/Cyrillic characters2 bytes per code point
Many BMP characters3 bytes per code point
Many supplementary characters4 bytes per code point

A string containing 100 ASCII characters can therefore require 100 UTF-8 bytes, while another string containing 100 non-ASCII code points may require significantly more.

UTF-8 Is Not Always the Same as String Length

Applications frequently make mistakes when they use a programming language's string length as a substitute for encoded byte length. These values answer different questions.

const text = "😀";

const codeUnits = text.length;
const utf8Bytes = new TextEncoder().encode(text).length;

console.log(codeUnits); // 2
console.log(utf8Bytes); // 4

The JavaScript string contains two UTF-16 code units, while the UTF-8 representation contains four bytes. Neither number should automatically be interpreted as the number of user-perceived characters.

UTF-8 in Network Protocols

When text is sent over a network, the application ultimately transmits bytes. UTF-8 defines how Unicode text becomes those bytes. The receiving side must decode them using the same encoding to reconstruct the intended text.

Sender:
Unicode text
    ↓
UTF-8 encoding
    ↓
Bytes
    ↓
Network
    ↓
Bytes
    ↓
UTF-8 decoding
    ↓
Unicode text

If either side uses the wrong encoding, the resulting text can be corrupted even though the network successfully transmitted every byte.

UTF-8 and APIs

UTF-8 is a natural fit for modern APIs because JSON and many other data formats can represent Unicode text, while UTF-8 provides a standardized byte representation for transport.

HTTP/1.1 200 OK
Content-Type: application/json

{"message":"Привет, мир!"}

When an API returns unexpected characters, inspect the actual response bytes and the decoding performed by the client. Looking only at the rendered text can hide whether the problem originated on the server or client side.

UTF-8 File Conversion Workflow

Suppose a legacy text file uses a non-Unicode encoding and needs to be converted to UTF-8. The correct process is not to reinterpret the existing bytes as UTF-8. The original encoding must first be decoded correctly.

Legacy bytes
    ↓
Decode using original encoding
    ↓
Unicode characters
    ↓
Encode using UTF-8
    ↓
UTF-8 bytes

If the source encoding is unknown, an encoding detector may provide a candidate, but automatic detection should not be treated as infallible. Verification against known text and metadata is often necessary.

Common UTF-8 Mistakes

  • Assuming Unicode and UTF-8 are the same thing
  • Assuming every character occupies one byte
  • Assuming every Unicode character occupies four bytes
  • Confusing Unicode code points with UTF-8 bytes
  • Treating `\uXXXX` escapes as UTF-8 byte sequences
  • Ignoring the difference between UTF-16 code units and UTF-8 bytes
  • Assuming a file extension determines its encoding
  • Adding a BOM to every UTF-8 file without considering compatibility
  • Removing a BOM without checking whether a legacy consumer expects it
  • Decoding UTF-8 bytes using a legacy code page
  • Encoding text as UTF-8 twice
  • Trying to repair mojibake by manually replacing visible characters
  • Assuming visible character count equals byte count
  • Ignoring invisible Unicode characters
  • Assuming every visually identical character has the same code point

Double-Encoding and Double-Decoding Problems

Another common source of corrupted text is applying an encoding or decoding operation at the wrong stage. For example, text that has already been decoded into Unicode should not be treated as if it were still raw UTF-8 bytes.

A well-designed application should have clear boundaries between byte-oriented and text-oriented operations. Decode bytes into text once, process the text, and encode it again only when bytes are required for storage or transmission.

💡 When debugging mojibake, write down each transformation the data went through. Many encoding bugs become obvious when the same data is accidentally encoded or decoded twice.

UTF-8 Best Practices

  • Use UTF-8 consistently for modern text whenever the relevant specification permits it.
  • Keep the distinction between Unicode code points and UTF-8 bytes clear.
  • Declare the encoding where the protocol or file format requires it.
  • Make sure the actual bytes match the declared encoding.
  • Decode external byte data explicitly and consistently.
  • Do not assume string length equals UTF-8 byte length.
  • Inspect raw bytes when debugging corrupted text.
  • Check for a UTF-8 BOM when unexpected invisible data appears at the start of a file.
  • Validate untrusted byte sequences before treating them as text.
  • Consider Unicode normalization when comparing canonically equivalent text.
  • Test with ASCII, accented Latin, Cyrillic, CJK characters, and emoji.
  • Use the same encoding assumptions across applications, APIs, databases, and files.

Testing UTF-8 Correctly

A UTF-8 implementation should not be tested only with English text. ASCII characters exercise the one-byte path, but they do not reveal many encoding problems. A useful test set should contain characters that require two, three, and four UTF-8 bytes.

ASCII:
Hello

Two-byte examples:
Привет
café

Three-byte examples:
€
漢字

Four-byte examples:
😀
🚀

Testing multiple scripts and supplementary characters makes it much more likely that problems involving byte length, decoding, Unicode handling, or incomplete UTF-8 support will be discovered before production.

A Practical UTF-8 Debugging Checklist

QuestionWhat to Inspect
Is the text corrupted?Actual bytes and decoder
Does the file start with an unexpected character?UTF-8 BOM: EF BB BF
Does ASCII work but Unicode fail?Multi-byte UTF-8 handling
Does the byte count look too large?Number of bytes per code point
Does JavaScript report a surprising length?UTF-16 code units vs code points
Does an API return mojibake?HTTP headers, response bytes, and decoder
Does a database contain corrupted text?Connection, storage, and client encoding
Does copied text behave differently?Invisible characters and normalization

Frequently Asked Questions

What is UTF-8 in simple terms?

UTF-8 is a variable-width encoding that converts Unicode code points into one to four bytes. ASCII characters use one byte, while characters outside the ASCII range use multiple bytes.

Is UTF-8 the same as Unicode?

No. Unicode defines a universal character repertoire and assigns code points to characters. UTF-8 is one encoding used to represent those code points as bytes.

How many bytes does a UTF-8 character use?

A UTF-8 encoded Unicode code point uses one, two, three, or four bytes depending on its code point. ASCII-range characters use one byte.

Why is ASCII compatible with UTF-8?

UTF-8 deliberately uses the same byte values as ASCII for Unicode code points U+0000 through U+007F. Therefore standard ASCII text is also valid UTF-8.

Why does Cyrillic use more bytes than English in UTF-8?

Standard Latin English characters are mostly in the ASCII range and require one UTF-8 byte. Cyrillic characters are outside that range and generally require two UTF-8 bytes.

Why does an emoji use four UTF-8 bytes?

Many emoji have Unicode code points above U+FFFF. These code points require four bytes in UTF-8.

Does UTF-8 need a BOM?

No. UTF-8 does not need a BOM because it has no byte-order ambiguity. Some files may still include the UTF-8 BOM sequence EF BB BF for compatibility or encoding-signature purposes.

How can I tell whether bytes are valid UTF-8?

A UTF-8 validator or byte-level inspector can check the structure of leading and continuation bytes, Unicode ranges, and invalid or overlong sequences.

Helpful UTF-8 Tools

Byte-level encoding tools are especially useful when UTF-8 problems are difficult to diagnose visually. A UTF-8 inspector can show how characters are represented as bytes, while an ASCII converter can help compare the ASCII-compatible portion of UTF-8 with traditional ASCII values.

A Unicode escape converter is useful when working with `\uXXXX` representations, while a BOM detector can identify an unexpected UTF-8 signature at the beginning of a file. An invisible-character detector can expose zero-width characters, non-breaking spaces, and other Unicode characters that are difficult to see directly.

Conclusion

UTF-8 is a variable-width Unicode encoding that represents each Unicode code point using one to four bytes. Its most important compatibility feature is that the first 128 Unicode code points use exactly the same byte values as standard ASCII.

Understanding UTF-8 requires keeping several layers separate. Unicode defines code points, UTF-8 converts those code points into bytes, and applications decode those bytes back into text. A visible character, its Unicode code point, and its UTF-8 byte sequence are related but are not the same thing.

For modern web development and general text interchange, UTF-8 is a practical default because it supports international text, preserves ASCII compatibility, and avoids the limitations of older single-byte character encodings. When encoding problems occur, inspecting the actual bytes and tracing the complete decode-and-encode path is usually the fastest way to find the cause.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.