Ctrl + K
Regex22 min read

Regex Character Classes

A practical guide to regex character classes covering brackets, ranges, negated classes, shorthand classes, Unicode properties, escaping, and common patterns.

Published: 2026-10-05

Regex character classes let you describe which characters are allowed at a particular position in a pattern. Instead of matching one specific character, a character class can match any character from a defined set, such as a letter, digit, vowel, hexadecimal character, or whitespace character.

The basic syntax uses square brackets, such as [abc]. Character classes can also contain ranges like [a-z], be negated with [^...], or be replaced by shorthand classes such as \d, \w, and \s. Modern regex engines also provide Unicode property escapes for more precise matching of international text.

This guide explains how character classes work, how ranges and negation behave, how to escape special characters, how character classes interact with quantifiers, and which approach to use for common text-processing tasks.

Character Classes Quick Reference

PatternMeaningExample
[abc]a, b, or c[abc]+
[a-z]Any lowercase ASCII letter[a-z]+
[A-Z]Any uppercase ASCII letter[A-Z]+
[0-9]Any ASCII digit[0-9]+
[a-zA-Z]Any ASCII letter[a-zA-Z]+
[^abc]Any character except a, b, or c[^abc]+
[aeiou]Any listed vowel[aeiou]
[a-z0-9]Lowercase letter or digit[a-z0-9]+
\dDigit according to the regex engine\d+
\DNon-digit character\D+
\wWord character according to the engine\w+
\WNon-word character\W+
\sWhitespace character\s+
\SNon-whitespace character\S+

What Is a Regex Character Class?

A character class matches one character from a specified set. The set is written between square brackets.

[abc]

This pattern matches one character: a, b, or c. It does not mean that the text must contain abc in that order.

a
b
c

For example, the pattern [abc] can find one matching character inside the strings apple, boat, or cat. It matches a single character each time the class is evaluated.

Character Classes Match One Character

One of the most important things to understand is that a character class normally represents exactly one character. Repetition requires a quantifier.

[abc]

The expression above matches one character from the set. To match one or more characters from that set, add +.

[abc]+

Now the expression can match a, ab, abc, cab, and other sequences containing only the characters a, b, and c.

💡 Think of a character class as a choice for one position in the pattern. Quantifiers such as + or * determine whether that choice can repeat.

Simple Character Classes

The simplest class is a list of individual characters.

[aeiou]

This matches one lowercase vowel.

[0123456789]

This matches one of the ten ASCII digits. Although it works, [0-9] is usually shorter and easier to read.

[0-9]

Character Ranges

A hyphen inside a character class can define a range. Instead of listing every character individually, you can specify the first and last character.

RangeMeaning
[a-z]Lowercase ASCII letters
[A-Z]Uppercase ASCII letters
[0-9]ASCII digits
[a-f]Lowercase letters a through f
[A-F]Uppercase letters A through F
[0-9A-F]Digits and uppercase hexadecimal letters

For example, [a-z] represents the lowercase ASCII letters from a through z.

[a-z]+

Adding + allows a sequence of lowercase ASCII letters.

Combining Multiple Ranges

Multiple ranges can appear in the same character class.

[a-zA-Z0-9]

This class allows lowercase letters, uppercase letters, and ASCII digits.

^[a-zA-Z0-9]+$

With anchors and +, the complete input must contain one or more ASCII letters or digits.

Multiple Ranges and Individual Characters

A class can combine ranges with individual characters.

[a-z0-9_-]

This allows lowercase ASCII letters, digits, underscores, and hyphens.

^[a-z0-9_-]+$

A pattern like this is useful for simple slugs, identifiers, or machine-oriented names when those exact ASCII restrictions are desired.

Negated Character Classes

A caret immediately after the opening [ negates a character class. Instead of matching characters inside the class, it matches a character that is not inside the class.

[^0-9]

This matches any character that is not one of the ASCII digits 0 through 9.

[^aeiou]

This matches any character except the listed lowercase vowels.

⚠️ The caret has a special negation meaning only when it appears immediately after the opening bracket. Elsewhere inside a class, its meaning can be different or literal depending on the position and regex engine.

Negated Classes with Quantifiers

Negated classes become particularly useful when you want to consume everything except a specific delimiter.

[^,]+

This matches one or more characters that are not commas. It can be useful for processing a simple comma-separated value when the field itself cannot contain commas.

"[^"]*"

This common pattern matches a double-quoted section without allowing another double quote inside the matched content. It is suitable only for simple formats where escaped quotes and other syntax are not required.

The Difference Between [^...] and .

A dot is a broad wildcard whose behavior depends on the regex engine and flags. A negated character class explicitly describes what must not be matched.

.*

This allows many characters, subject to the dot's rules.

[^,]*

This specifically allows any number of characters except commas. The second expression communicates the intended restriction more clearly.

💡 If you know the delimiter or characters that must stop a match, a negated character class is often more precise than .*.

Shorthand Character Classes

Regex engines provide shorthand classes for common character categories. The three most familiar are \d for digits, \w for word characters, and \s for whitespace.

PatternCommon meaning
\dDigit
\DNon-digit
\wWord character
\WNon-word character
\sWhitespace
\SNon-whitespace

\d and Digits

\d is a shorthand for a digit according to the rules of the regex engine.

\d

For a sequence of one or more digits, combine it with +.

\d+

In many common regex engines, \d corresponds to ASCII digits by default. Some engines or Unicode modes can provide broader Unicode-aware behavior. If exact ASCII behavior is required, [0-9] is often clearer.

\D and Non-Digits

\D is the complementary shorthand for a character that is not considered a digit by the engine.

\D+

This matches one or more non-digit characters according to the engine's definition.

\w and Word Characters

\w is commonly used for word characters. In many traditional regex implementations this means ASCII letters, digits, and underscore.

\w+

The exact behavior of \w can vary between regex engines and Unicode modes. It should not automatically be interpreted as every letter or word character used by every writing system.

\W and Non-Word Characters

\W represents the complement of \w under the regex engine's rules.

\W+

It can be useful for finding punctuation, separators, and other characters outside the engine's word-character definition.

\s and Whitespace

\s matches whitespace characters recognized by the regex engine. Depending on the engine, this can include spaces, tabs, line breaks, and other whitespace characters.

\s+

This is commonly used to find runs of whitespace.

\S+

\S matches one or more non-whitespace characters.

Character Classes vs Shorthand Classes

GoalPossible pattern
ASCII digits[0-9]
Engine-defined digits\d
ASCII lowercase letters[a-z]
ASCII uppercase letters[A-Z]
ASCII letters[A-Za-z]
ASCII letters and digits[A-Za-z0-9]
Whitespace\s
Non-whitespace\S

The choice depends on your requirements. Explicit ranges such as [0-9] make ASCII restrictions obvious. Shorthand classes such as \d and \s are shorter and often more convenient when the engine's category behavior is exactly what you need.

Escaping Characters Inside Classes

Some characters have special meanings inside character classes. If you need one of those characters literally, escape it when required by the regex syntax.

CharacterCommon special roleLiteral example
[Starts a character class\[
]Ends a character class\]
-Defines a range\-
^Negates a class when placed first\^
\Introduces an escape\\

The hyphen is particularly important. In the middle of a class, it can define a range. If you want a literal hyphen, escape it or place it where your regex engine treats it literally, commonly at the beginning or end of the class.

[a-z-]

This class allows lowercase ASCII letters and a literal hyphen.

Literal Special Characters in a Class

Many regex metacharacters lose some of their special meaning inside a character class. For example, a dot normally acts as a wildcard outside a class but represents a literal dot inside a class.

[.]

This matches a literal period.

[+*?]

This matches one of the literal characters +, *, or ?.

Combining Characters in a Class

A character class can contain several different types of allowed characters at once.

[A-Za-z0-9]

This allows uppercase letters, lowercase letters, and digits.

[A-Za-z0-9._-]

This extends the set with a dot, underscore, and hyphen. Such a class can be useful for simple identifiers or filename-like strings when these exact restrictions are appropriate.

Character Classes and Quantifiers

Character classes become much more useful when combined with quantifiers. The class defines what one character can be, while the quantifier defines how many such characters can occur.

[0-9]+

One or more ASCII digits.

[A-Za-z]{3,20}

Between three and twenty ASCII letters.

[^\s]+

One or more non-whitespace characters.

Character Classes with Anchors

Anchors allow a character class to describe the entire input instead of merely finding a matching substring.

^[0-9]+$

This requires the entire input to contain one or more ASCII digits.

^[A-Za-z]+$

This requires the entire input to contain one or more ASCII letters.

⚠️ A pattern that looks like a validator may still have engine-specific anchor behavior, especially in multiline modes. For critical validation, understand how the target regex API interprets the pattern and flags.

Common Character Class Patterns

TaskPatternDescription
ASCII lowercase letters[a-z]+One or more lowercase ASCII letters
ASCII uppercase letters[A-Z]+One or more uppercase ASCII letters
ASCII letters[A-Za-z]+One or more ASCII letters
ASCII digits[0-9]+One or more ASCII digits
Letters and digits[A-Za-z0-9]+One or more ASCII letters or digits
Hexadecimal[0-9A-Fa-f]+One or more hexadecimal characters
Vowels[aeiou]+One or more lowercase vowels
Non-digits[^0-9]+One or more characters other than ASCII digits
Non-whitespace[^\s]+One or more non-whitespace characters
Simple slug[a-z0-9-]+Lowercase ASCII letters, digits, and hyphens

Hexadecimal Character Classes

Hexadecimal values are a common example of a character class because only a limited set of characters is allowed.

[0-9A-Fa-f]

This matches one hexadecimal character.

^[0-9A-Fa-f]{6}$

This checks for exactly six hexadecimal characters, which is a common shape for a full hexadecimal color value after removing the # prefix.

^#[0-9A-Fa-f]{6}$

This version includes the # character and therefore matches values such as #12a4ef.

Character Classes for Whitespace Cleanup

Character classes are useful for text-cleaning tasks. For example, \s can identify whitespace while a literal space inside a class can target spaces specifically.

[ \t]+

This class targets spaces and tabs. By contrast, \s+ can include additional whitespace characters depending on the regex engine.

[^\S\r\n]+

More advanced combinations can be used when you need to distinguish horizontal whitespace from line breaks. Such patterns are more engine-dependent and should be tested carefully.

Unicode Character Classes

ASCII ranges such as [A-Za-z] cover only a small portion of written language. If an application needs to process names, messages, search terms, or other multilingual content, Unicode-aware regex features may be more appropriate.

Some modern regex engines support Unicode property escapes such as \p{Letter}. These expressions match characters belonging to Unicode-defined categories.

\p{Letter}+

The pattern can match one or more Unicode letters in engines that support Unicode property escapes.

\p{Number}+

Similarly, Unicode number properties can be used where supported.

⚠️ Unicode property syntax and supported properties differ between regex engines. In JavaScript, Unicode property escapes require Unicode-aware regex support and are used with the appropriate Unicode-related flag. Always check the target environment before relying on them.

ASCII Ranges vs Unicode Properties

RequirementTypical approach
Only ASCII lowercase letters[a-z]
Only ASCII uppercase letters[A-Z]
Only ASCII digits[0-9]
ASCII letters[A-Za-z]
Unicode letters\p{Letter} when supported
Unicode numbers\p{Number} when supported
Unicode decimal digits\p{Decimal_Number} when supported

Explicit ASCII ranges are useful when an identifier or protocol deliberately defines an ASCII-only format. Unicode properties are more appropriate when the application needs to recognize categories of characters across writing systems.

Case-Insensitive Matching

Instead of writing separate uppercase and lowercase ranges, some regex engines provide a case-insensitive flag such as i.

const pattern = /[a-z]+/i;

This allows the pattern to match uppercase and lowercase forms according to the engine's case-insensitive matching rules.

Whether a case-insensitive pattern behaves exactly like a manually written ASCII range can depend on Unicode and engine rules. For strict protocol or identifier formats, explicit character sets can sometimes communicate the requirement more clearly.

Character Classes and the Dot

The dot and character classes are both ways to match a character, but they communicate different intentions.

PatternPurpose
.Broad wildcard subject to engine and flag rules
[a-z]One lowercase ASCII letter
[0-9]One ASCII digit
[^,]One character other than a comma
[\s\S]A character from whitespace or non-whitespace categories

A specific class is usually easier to understand when the allowed or excluded characters are known. A dot is useful when the exact character set is intentionally broad.

Character Class Intersection and Advanced Set Operations

Some regex engines provide advanced character-set operations such as intersection, subtraction, or Unicode set syntax. These features are not universal, so they should not be treated as portable regex syntax.

For everyday regular expressions, ordinary classes, ranges, negation, shorthand classes, and Unicode properties cover most practical requirements. Advanced set operations are most useful when a specific regex engine supports them and the matching rules genuinely require that level of precision.

Common Mistakes with Character Classes

  • Assuming [abc] matches the complete string abc instead of one character from the set.
  • Forgetting that a hyphen can define a range inside a class.
  • Putting ^ somewhere other than the beginning and assuming it always negates the class.
  • Using [A-Za-z] when the application actually needs Unicode letters.
  • Assuming \d, \w, and \s have identical definitions in every regex engine.
  • Forgetting to escape or correctly position a literal hyphen.
  • Using a broad negated class when a more specific allowed set would be safer.
  • Using .* when a negated character class would describe a delimiter more precisely.
  • Forgetting to combine a class with a quantifier when multiple characters should match.
  • Assuming a character class validates the entire input without using appropriate boundaries or anchors.

How to Build a Character Class

A simple way to design a character class is to start with the exact set of characters that should be allowed at one position.

  • List the characters that should be allowed.
  • Replace long consecutive lists with ranges where appropriate.
  • Combine multiple ranges if they belong to the same allowed set.
  • Escape characters that have special meaning in the class.
  • Decide whether the class should be negated.
  • Choose between an explicit class and a shorthand such as \d or \s.
  • Add a quantifier if the class needs to match multiple characters.
  • Add anchors if the entire input must satisfy the pattern.
  • Test both valid and invalid examples.
^[a-z0-9_-]{3,30}$

For example, this pattern defines an ASCII-only identifier containing lowercase letters, digits, underscores, or hyphens, with a length from three to thirty characters.

Character Classes for Common Tasks

Username Characters

^[A-Za-z0-9_]{3,20}$

This is a simple ASCII username rule. It allows letters, digits, and underscores and limits the length to three through twenty characters.

Slug Characters

^[a-z0-9-]+$

This allows lowercase ASCII letters, digits, and hyphens. Additional application rules may be required if consecutive or leading hyphens are not allowed.

Hexadecimal Characters

^[0-9A-Fa-f]+$

Whitespace-Separated Words

\S+

This can find sequences of non-whitespace characters. It is useful for simple tokenization when whitespace is the only separator.

Non-Comma Text

[^,]+

This matches one or more characters other than commas.

Text Without Quotes

[^"']+

This matches one or more characters other than a double quote or single quote. It can be useful for simple text filtering where quoted characters are delimiters.

Character Classes in Find and Replace

Character classes are especially useful in find-and-replace operations because they can identify variable text without requiring a separate pattern for every possible character.

[ \t]+

For example, a text editor can use a pattern like this to find runs of spaces and tabs. The replacement can then normalize them to a single space, depending on the editor's replacement behavior.

[^\S\r\n]+

More advanced classes can distinguish horizontal whitespace from line breaks, which can be useful when cleaning text without joining separate lines.

Testing Character Classes

Character classes should be tested with values that sit both inside and outside the intended set. This is particularly important for negated classes and Unicode-aware patterns.

  • Test every explicitly allowed character.
  • Test characters immediately outside a defined range.
  • Test the boundaries of ranges such as a, z, 0, and 9.
  • Test punctuation when using \w or custom identifier classes.
  • Test spaces, tabs, and line breaks when using whitespace patterns.
  • Test non-ASCII characters if the application supports international text.
  • Test empty input when a quantifier is involved.
  • Test long sequences when performance matters.
💡 A regex tester makes character-class mistakes easy to spot. Try changing one character at a time around the edge of the allowed set rather than testing only obvious examples.

Character Classes and Performance

Character classes are generally straightforward and often make patterns more precise than broad wildcard expressions. They can also help reduce unnecessary matching choices when the allowed character set is clearly defined.

However, a character class does not automatically make a regex safe or fast. Performance problems can still arise when classes are combined with nested quantifiers, ambiguous alternatives, or other patterns that create extensive backtracking.

⚠️ When a regex processes untrusted input on a server, test worst-case inputs and avoid unnecessarily complex combinations of repetition and alternation. Character classes are useful building blocks, but overall regex structure still determines performance.

When Not to Use a Character Class

A character class is appropriate when the problem is fundamentally about which characters may occur at a position. It is not a substitute for parsing structured data.

For example, if you need to understand nested JSON, HTML, XML, or another structured format, a parser is generally more appropriate than trying to represent the entire grammar with character classes and other regex constructs.

Likewise, if a programming language already provides a clear API for a simple task such as checking whether a string starts with a known prefix, using that API is often easier to maintain than introducing regex.

Character Classes vs String Methods

TaskOften simpler approach
Check exact textString equality
Check a known prefixstartsWith()
Check a known suffixendsWith()
Find a fixed substringincludes()
Split by a fixed delimitersplit()
Match a variable character setRegex character class
Validate a structured patternRegex or a dedicated parser, depending on the format

Regex Character Classes in JavaScript

JavaScript supports ordinary character classes, ranges, negated classes, shorthand classes, and Unicode property escapes in modern environments.

const digits = /^[0-9]+$/;
const letters = /^[A-Za-z]+$/;
const whitespace = /\s+/g;

console.log(digits.test("12345"));
// true

When Unicode property escapes are appropriate, JavaScript can use expressions such as \p{Letter} with Unicode-aware regex syntax.

const letters = /^\p{Letter}+$/u;

The Unicode-aware form is useful when the application should recognize letters beyond the ASCII range. The exact requirements of the application should determine whether Unicode properties or an explicit ASCII class is more appropriate.

A Practical Decision Guide

QuestionUse
Do I need one of a small list of characters?[abc]
Do I need a continuous ASCII range?[a-z] or [0-9]
Do I need several ranges?[a-zA-Z0-9]
Do I need everything except a set?[^...]
Do I need digits?\d or [0-9], depending on requirements
Do I need whitespace?\s
Do I need a Unicode character category?\p{...} when supported
Do I need multiple matching characters?Add a quantifier such as + or *

Frequently Asked Questions

What is a character class in regex?

A character class defines a set of characters that can match one position in a regex. It is usually written inside square brackets, such as [abc], which matches a, b, or c.

What does [a-z] mean in regex?

[a-z] matches one lowercase ASCII letter from a through z. It does not automatically include uppercase or non-ASCII letters.

What does [^...] mean in regex?

A caret immediately after the opening bracket negates the class. For example, [^0-9] matches a character that is not an ASCII digit.

What is the difference between [0-9] and \d?

[0-9] explicitly represents ASCII digits. \d is a shorthand whose exact behavior can depend on the regex engine and Unicode mode. Use the explicit range when an ASCII-only requirement needs to be obvious.

Does a character class match multiple characters?

A character class normally matches one character. To match multiple characters from the same class, add a quantifier such as +, *, or {3,10}.

How do I match a literal hyphen inside a character class?

You can escape the hyphen as \- or place it in a position where the target regex engine treats it literally, commonly at the beginning or end of the class.

Can regex character classes match Unicode letters?

Yes, in regex engines that support Unicode property escapes. A pattern such as \p{Letter} can represent Unicode letters, but the syntax and supported properties depend on the engine.

Should I use [A-Za-z] for names?

Only if the application intentionally restricts names to ASCII letters. For international names and text, [A-Za-z] is too restrictive and a Unicode-aware approach may be more appropriate.

Conclusion

Regex character classes are one of the core building blocks of regular expressions. Square brackets let you define explicit sets, ranges make those sets shorter, and a leading caret creates a negated class. Shorthand expressions such as \d, \w, and \s provide convenient predefined categories.

For simple ASCII formats, explicit classes such as [A-Za-z0-9] make the allowed characters clear. For multilingual applications, Unicode property escapes can provide a more suitable representation when the target regex engine supports them.

The most effective character classes are precise, easy to read, and tested against both valid and invalid input. Combine them with quantifiers and anchors when you need to describe complete values, and prefer dedicated parsers when the data has a structure that goes beyond character-level matching.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.