Regex Cheat Sheet
A practical quick reference for regular expressions with syntax, character classes, quantifiers, anchors, groups, flags, and common regex patterns.
Regular expressions, usually called regex or regexp, provide a compact way to search, match, validate, extract, and replace text. The syntax can look difficult at first, but most patterns are built from a relatively small set of components. This Regex Cheat Sheet collects the syntax you are most likely to need in everyday development.
Use this page as a quick reference when writing a pattern rather than reading a complete tutorial. Each section includes the regex syntax, what it does, and practical examples. The examples use syntax commonly supported by modern JavaScript and many other regex engines, with compatibility differences noted where they matter.
Regex Quick Reference
| Syntax | Meaning | Example |
|---|---|---|
| . | Any character except a line terminator in many engines | a.c |
| ^ | Start of string or line | ^Hello |
| $ | End of string or line | world$ |
| * | Zero or more repetitions | ab* |
| + | One or more repetitions | ab+ |
| ? | Zero or one repetition | colou?r |
| {n} | Exactly n repetitions | \d{4} |
| {n,} | At least n repetitions | \d{2,} |
| {n,m} | Between n and m repetitions | \d{2,4} |
| [] | Character class | [abc] |
| [^] | Negated character class | [^0-9] |
| | | Alternation | cat|dog |
| () | Capturing group | (abc) |
| (?:) | Non-capturing group | (?:abc) |
| \ | Escape special syntax | \. |
| \d | Digit character | \d+ |
| \D | Non-digit character | \D+ |
| \w | Word character | \w+ |
| \W | Non-word character | \W+ |
| \s | Whitespace character | \s+ |
| \S | Non-whitespace character | \S+ |
| \b | Word boundary | \bcat\b |
| \B | Non-word boundary | \Bcat\B |
Literal Characters
Most ordinary characters match themselves. A regex containing the pattern cat normally searches for the sequence c, a, followed by t.
catThe pattern matches the text cat inside a larger string unless additional anchors or boundaries are used.
cat
concatenate
A cat is hereRegex syntax is case-sensitive by default in many engines. A pattern containing cat does not automatically match Cat or CAT. Case-insensitive matching can usually be enabled with a flag such as i.
The Dot Character
The dot matches a single character, although the exact treatment of line terminators depends on the regex engine and flags.
c.tThis can match cat, cot, cut, c9t, and many other three-character sequences beginning with c and ending with t.
Character Classes
Character classes match one character from a defined set. They are written inside square brackets.
| Pattern | Meaning | Example |
|---|---|---|
| [abc] | a, b, or c | [abc]+ |
| [a-z] | Any lowercase letter from a through z | [a-z]+ |
| [A-Z] | Any uppercase letter from A through Z | [A-Z]+ |
| [0-9] | Any ASCII digit | [0-9]+ |
| [a-zA-Z] | Any ASCII letter | [a-zA-Z]+ |
| [aeiou] | Any listed vowel | [aeiou] |
| [^abc] | Any character except a, b, or c | [^abc]+ |
| [a-z0-9] | Lowercase ASCII letter or digit | [a-z0-9]+ |
A hyphen usually defines a range inside a character class. A caret immediately after the opening bracket negates the class.
[0-9]
[^0-9]
[a-zA-Z]
[^\s]Shorthand Character Classes
| Pattern | Meaning |
|---|---|
| \d | A digit |
| \D | A character that is not a digit |
| \w | A word character; exact definition depends on the engine |
| \W | A non-word character |
| \s | Whitespace |
| \S | Non-whitespace |
A common pattern for a sequence of digits is \d+. A common pattern for one or more whitespace characters is \s+.
\d+
\w+
\s+In JavaScript and many other engines, shorthand classes have specific Unicode and flag behavior. If your application handles multilingual data, verify the behavior of the target regex engine instead of assuming that \w represents every possible letter.
Anchors
Anchors describe positions rather than consuming characters. They are useful when you want to control where a match can occur.
| Anchor | Meaning | Example |
|---|---|---|
| ^ | Beginning of input or line depending on mode | ^Hello |
| $ | End of input or line depending on mode | world$ |
| \b | Word boundary | \bword\b |
| \B | Not a word boundary | \Bword\B |
For example, the pattern ^Hello requires Hello to occur at the beginning of the relevant input or line.
^HelloTo require an entire simple value to consist of digits, you can combine start and end anchors.
^\d+$Word Boundaries
The \b assertion identifies a boundary between a word character and a non-word character, or the edge of the input. It is useful when searching for a complete word instead of a substring.
\bcat\bThis pattern can match cat as a separate word but does not normally match the cat portion of concatenate. Exact behavior depends on the engine's definition of word characters.
Quantifiers
Quantifiers specify how many times the preceding token, character class, group, or expression can occur.
| Quantifier | Meaning | Example |
|---|---|---|
| * | Zero or more | a* |
| + | One or more | a+ |
| ? | Zero or one | a? |
| {n} | Exactly n | a{3} |
| {n,} | At least n | a{3,} |
| {n,m} | Between n and m | a{3,5} |
a*
a+
a?
a{3}
a{3,}
a{3,5}The difference between * and + is important. The first allows zero occurrences, while the second requires at least one occurrence.
Greedy and Lazy Quantifiers
Most quantifiers are greedy by default. A greedy quantifier attempts to consume as much matching text as possible while still allowing the overall pattern to succeed.
.*Appending ? to a supported quantifier commonly makes it lazy, meaning it attempts to consume as little as possible while still allowing the rest of the pattern to match.
.*?| Greedy | Lazy |
|---|---|
| * | *? |
| + | +? |
| ? | ?? |
| {n,m} | {n,m}? |
| {n,} | {n,}? |
Alternation
The vertical bar | represents alternation. It means that one alternative or another can match.
cat|dogAlternation can be combined with groups to control its scope.
(cat|dog)s?This can match cat, cats, dog, or dogs. Without grouping, the scope of an alternation may be different from what you intended.
Groups
| Syntax | Purpose |
|---|---|
| (...) | Capturing group |
| (?:...) | Non-capturing group |
| (?<name>...) | Named capturing group in engines that support it |
Capturing groups store the text matched by the group so application code or replacement syntax can access it.
(\d{4})-(\d{2})-(\d{2})For an input such as 2026-09-19, the groups can represent the year, month, and day separately.
Use a non-capturing group when you need grouping for structure or quantification but do not need to capture the matched text.
(?:https?|ftp)://Backreferences
A backreference matches the same text captured by an earlier group. Numeric backreferences commonly use syntax such as \1 or \2.
\b(\w+)\s+\1\bThis pattern can detect a repeated word such as the the. Backreferences are useful for certain text-processing tasks but can make patterns harder to understand.
Lookahead and Lookbehind
Lookarounds test surrounding text without consuming it as part of the match. Support varies between regex engines, especially for lookbehind.
| Syntax | Purpose |
|---|---|
| (?=...) | Positive lookahead |
| (?!...) | Negative lookahead |
| (?<=...) | Positive lookbehind |
| (?<!...) | Negative lookbehind |
\d+(?=\s*USD)The example looks for digits followed by optional whitespace and USD, while the USD portion is not included in the consumed match.
Escaping Special Characters
Some characters have special meanings in regex syntax. To match them literally, escape them with a backslash when required.
| Character | Typical regex meaning | Literal form |
|---|---|---|
| . | Any character | \. |
| * | Repetition | \* |
| + | Repetition | \+ |
| ? | Optionality or assertion syntax | \? |
| ^ | Start anchor or class negation | \^ |
| $ | End anchor | \$ |
| ( | Group | \( |
| ) | Group | \) |
| [ | Character class | \[ |
| ] | Character class | \] |
| { | Quantifier | \{ |
| } | Quantifier | \} |
| | | Alternation | \| |
| \ | Escape character | \\ |
\.com$This matches text ending in the literal .com rather than treating the dot as a wildcard.
Regex Flags
Flags modify how a regular expression is interpreted. The exact set of available flags depends on the regex engine.
| Flag | Common purpose | JavaScript example |
|---|---|---|
| g | Global matching | /cat/g |
| i | Case-insensitive matching | /cat/i |
| m | Multiline anchor behavior | /^cat$/m |
| s | Dot matches line terminators | /a.*b/s |
| u | Unicode-aware matching | /\u{1F600}/u |
| y | Sticky matching from lastIndex | /cat/y |
| d | Match indices in supported JavaScript environments | /cat/d |
| v | Enhanced Unicode character-set behavior in supported JavaScript environments | /[\p{Letter}]/v |
For example, the i flag makes a simple literal match case-insensitive, while g allows repeated matches instead of stopping after the first match in JavaScript's standard regex APIs.
const pattern = /hello/gi;
const matches = text.match(pattern);Common Regex Patterns
The following patterns cover common development and text-processing tasks. They should be treated as starting points rather than universal validators. Real-world requirements often need additional rules.
Digits Only
^\d+$Matches a non-empty string containing only digit characters according to the behavior of \d in the target engine.
Integer Number
^-?\d+$Matches an optional minus sign followed by one or more digits.
Decimal Number
^-?\d+(?:\.\d+)?$Matches integers and decimal values with a dot as the decimal separator, including an optional negative sign.
Basic Email Pattern
^[^\s@]+@[^\s@]+\.[^\s@]+$This is a practical lightweight check for an email-shaped value. It is not a complete implementation of every valid email address defined by relevant standards.
URL Starting with HTTP or HTTPS
^https?:\/\/[^\s]+$This pattern checks for an HTTP or HTTPS scheme followed by a non-whitespace URL-like value. URL validation can become much more complex than this pattern.
Hexadecimal Color
^#(?:[0-9A-Fa-f]{3}|[0-9A-Fa-f]{6})$Matches common three-digit and six-digit hexadecimal color notation such as #fff and #12a4ef.
Whitespace
\s+Matches one or more whitespace characters. It is commonly used when searching for repeated spaces, tabs, or other whitespace recognized by the regex engine.
Repeated Spaces
{2,}Matches two or more literal space characters. Unlike \s+, this does not target tabs or other whitespace characters.
Trim Leading and Trailing Whitespace
^\s+|\s+$This pattern can identify whitespace at the beginning or end of a string. In application code, a native trim method is usually clearer when the goal is simply to trim a string.
Date in YYYY-MM-DD Format
^\d{4}-\d{2}-\d{2}$Time in HH:MM Format
^(?:[01]\d|2[0-3]):[0-5]\d$This pattern checks a 24-hour time from 00:00 through 23:59.
Simple Username Pattern
^[A-Za-z0-9_]{3,20}$Matches usernames containing ASCII letters, digits, and underscores with a length between 3 and 20 characters.
Letters and Spaces
^[A-Za-z ]+$Matches only ASCII letters and literal spaces. For multilingual names, this pattern is usually too restrictive.
Extract Numbers from Text
\d+When used for searching rather than full validation, \d+ can find sequences of digits inside larger text.
Find Hashtags
#[A-Za-z0-9_]+Matches a basic ASCII hashtag-like sequence. Production hashtag extraction may require Unicode-aware rules.
Find Quoted Text
"[^"]*"Matches text surrounded by double quotes when escaped quotes and multiline requirements are not part of the input format.
Find HTML-Like Tags
<[^>]+>Regex for Find and Replace
Regex is especially useful for bulk text replacement. Capturing groups can preserve parts of the original text while changing its structure.
^(\w+),\s*(\w+)$For an input such as Smith, John, the two captured groups can represent the last and first names. A replacement can then rearrange those captured values, depending on the syntax of the editor or programming language being used.
When using regex replacement, always check how the target environment references captures. Some systems use $1, $2, and named references, while others use different replacement syntax.
Regex in JavaScript
JavaScript supports regular expressions through regex literals and the RegExp constructor.
const pattern = /^\d+$/;
pattern.test("12345");
// trueconst pattern = /\d+/g;
const matches = "Order 123 and 456".match(pattern);
// ["123", "456"]When creating a regex dynamically, the RegExp constructor is useful, but remember that JavaScript string escaping and regex escaping are separate layers.
const pattern = new RegExp("\\d+");Regex Flags in JavaScript
| Flag | Use |
|---|---|
| g | Find all matches rather than only the first match in APIs where global matching applies |
| i | Ignore case differences |
| m | Make ^ and $ work with individual lines |
| s | Allow . to match line terminators |
| u | Enable Unicode-aware regex behavior |
| y | Require matching at the current lastIndex |
| d | Expose match indices in supported environments |
| v | Enable newer Unicode set features in supported environments |
Regex Cheat Sheet by Task
| Task | Pattern |
|---|---|
| Digits only | ^\d+$ |
| Integer | ^-?\d+$ |
| Decimal | ^-?\d+(?:\.\d+)?$ |
| Basic email shape | ^[^\s@]+@[^\s@]+\.[^\s@]+$ |
| HTTP/HTTPS URL shape | ^https?:\/\/[^\s]+$ |
| Hex color | ^#(?:[0-9A-Fa-f]{3}|[0-9A-Fa-f]{6})$ |
| YYYY-MM-DD shape | ^\d{4}-\d{2}-\d{2}$ |
| HH:MM time | ^(?:[01]\d|2[0-3]):[0-5]\d$ |
| Whitespace | \s+ |
| Word | \b\w+\b |
| Repeated word | \b(\w+)\s+\1\b |
| Quoted text | "[^"]*" |
| Basic username | ^[A-Za-z0-9_]{3,20}$ |
| HTML-like tag | <[^>]+> |
Regex vs String Methods
Regex is powerful, but it should not automatically be the first choice for every text operation. Simple tasks are often clearer with ordinary string methods.
| Task | Often simpler approach |
|---|---|
| Exact equality | === or equivalent |
| Prefix check | startsWith() |
| Suffix check | endsWith() |
| Simple substring search | includes() |
| Whitespace trimming | trim() |
| Simple replacement | replace() |
| Splitting by a fixed delimiter | split() |
Regex becomes more useful when the pattern contains alternatives, variable-length sections, character classes, repetitions, boundaries, captures, or other conditions that would be cumbersome to express with ordinary string operations.
Common Regex Mistakes
- Using .* when a narrower character class would describe the input more accurately.
- Forgetting that * allows zero matches while + requires at least one.
- Using a regex for full semantic validation when parsing or application logic is more appropriate.
- Assuming \w, \d, or \b have identical Unicode behavior in every regex engine.
- Forgetting to escape regex metacharacters when matching them literally.
- Using capturing groups when non-capturing groups would be sufficient.
- Assuming a regex that checks the shape of a date, URL, or email proves that the value is fully valid.
- Ignoring multiline, Unicode, or dotall behavior when processing multi-line text.
- Writing overly complex patterns without testing representative edge cases.
- Assuming regex syntax and replacement syntax are identical across programming languages and editors.
How to Test a Regex
A regex should be tested against both expected matches and expected non-matches. Testing only a successful example can hide mistakes in boundaries, optional sections, or repeated input.
- Start with a small representative input.
- Test the shortest valid value.
- Test the longest valid value if a length limit exists.
- Test empty input.
- Test missing required parts.
- Test unexpected whitespace.
- Test punctuation and special characters.
- Test uppercase and lowercase variants when relevant.
- Test Unicode or non-ASCII text when the application supports it.
- Test malformed input that should not match.
Regex Performance and ReDoS
Some regex engines use backtracking to explore possible ways a pattern can match. Poorly designed patterns with nested or overlapping quantifiers can create extremely expensive matching behavior on specially constructed input.
^(a+)+$Patterns with nested quantifiers like this are a common example used when discussing catastrophic backtracking. The exact performance depends on the regex engine and input.
Regex and Unicode
Modern regex engines increasingly provide Unicode-aware features, but their syntax and capabilities differ. A pattern designed around ASCII ranges such as [A-Za-z] does not automatically cover every letter used around the world.
\p{Letter}+In engines that support Unicode property escapes, patterns such as \p{Letter} can express character categories more accurately than manually listing ASCII ranges. JavaScript requires the appropriate Unicode-related flag for property escapes.
Regex Engine Differences
Regex is not one completely uniform language. JavaScript, Python, PCRE, .NET, Java, Rust, Go, and command-line tools can support different syntax, flags, Unicode features, and replacement rules.
| Feature | Why it matters |
|---|---|
| Lookbehind | Supported syntax and restrictions vary |
| Named groups | Group naming syntax differs |
| Unicode properties | Available categories and syntax can differ |
| Flags | The available flags are engine-specific |
| Replacement syntax | Capture references can use different forms |
| Backtracking behavior | Performance characteristics vary |
| Multiline behavior | Anchor semantics can differ with flags |
When moving a regex from one language or tool to another, test the pattern instead of assuming that a working expression is automatically portable.
When Not to Use Regex
Regex is excellent for pattern-based text processing, but some tasks are better handled by dedicated parsers or application logic. HTML and XML should generally be processed with parsers when you need to understand document structure. URLs should be parsed with URL-aware APIs when available. Dates should be validated and interpreted with date-aware logic rather than only checking their textual shape.
The goal is not to replace every string operation with regex. A good pattern solves a text-matching problem while remaining understandable, testable, and appropriate for the data being processed.
Helpful Regex Tools
- Regex testers for experimenting with patterns and checking matches.
- Regex generators for creating patterns from natural-language requirements or examples.
- Find and replace tools for applying regex-based transformations to text.
- Text cleaners for normalizing whitespace and removing unwanted text.
- Text diff tools for comparing the result before and after a regex transformation.
Frequently Asked Questions
What is the most important regex syntax to learn first?
Start with character classes such as [a-z] and \d, quantifiers such as *, +, and {n,m}, anchors such as ^ and $, grouping with parentheses, alternation with |, and escaping with a backslash. These components cover a large portion of everyday regex work.
What is the difference between * and + in regex?
The * quantifier matches zero or more occurrences, while + matches one or more. For example, a* can match an empty string, while a+ requires at least one a.
What does \d mean in regex?
\d is a shorthand character class for a digit according to the rules of the regex engine. Its exact Unicode behavior can differ, so verify the target engine when international numeric input matters.
What does ^\d+$ mean?
It combines start and end anchors with \d+. It is commonly used to require the entire input to consist of one or more digit characters.
What is the difference between a capturing and non-capturing group?
A capturing group stores the text it matches so code or replacement operations can reference it. A non-capturing group, written as (?:...), groups an expression without creating a capture.
Are regex patterns portable between programming languages?
Not always. Core syntax overlaps heavily, but engines differ in lookarounds, Unicode support, flags, named groups, replacement syntax, and performance characteristics. Test patterns in the target environment.
Can regex validate an email address or URL perfectly?
A regex can perform a useful shape check, but complete validation can be considerably more complicated. For URLs, URL-aware APIs are usually preferable. For email, applications often combine a practical format check with actual verification or confirmation.
Can regex be dangerous for application performance?
Yes. Some backtracking regex engines can spend excessive time on specially constructed input when patterns contain problematic nested or overlapping quantifiers. This is commonly associated with Regular Expression Denial of Service, or ReDoS.
Conclusion
Most practical regex patterns can be understood as combinations of a small number of building blocks: literal characters, character classes, quantifiers, anchors, groups, alternation, assertions, and flags. Once these pieces become familiar, even complex expressions are easier to read and modify.
Keep this cheat sheet nearby when working with regular expressions, but test every important pattern against realistic input. For production applications, clarity, correctness, engine compatibility, and predictable performance matter more than making a regex as short as possible.