Regular Expressions Explained
A practical introduction to regular expressions covering regex syntax, character classes, quantifiers, groups, anchors, alternation, flags, common patterns, and real-world examples.
Regular expressions, commonly called regex or regexp, are patterns used to search, match, validate, extract, replace, and transform text. Instead of looking only for an exact word or phrase, a regular expression can describe a whole class of strings that follow a particular structure.
For example, a simple regular expression can find every sequence of digits, identify words beginning with a particular letter, extract URLs from text, or validate the general structure of an email address. Regex is supported by many programming languages, command-line tools, editors, databases, and developer utilities.
Regular expressions can look cryptic at first because a small number of characters have special meanings. Once those characters are understood, regex becomes a compact language for describing text patterns.
What Is a Regular Expression?
A regular expression is a sequence of characters that defines a search pattern. A regex engine reads the pattern and determines whether portions of the input text match it.
catThis pattern matches the literal sequence cat. It can match the word cat inside a larger string unless additional boundaries or anchors are used.
The cat is sleeping.
A cat is outside.
catalogA search using the pattern cat can find cat in all three lines because the sequence also appears at the beginning of catalog.
Literal Characters
Many characters in a regex represent themselves. Letters, digits, and several ordinary punctuation characters can be matched literally.
helloThis matches the literal text hello.
2026This matches the sequence 2026.
Special Characters in Regex
Regular expressions also contain metacharacters. These characters have special meanings that allow a pattern to describe more than literal text.
| Character | Common meaning |
|---|---|
| . | Any character except line terminators in common regex modes |
| ^ | Start of input or line |
| $ | End of input or line |
| * | Zero or more repetitions |
| + | One or more repetitions |
| ? | Optional or zero or one repetition |
| {n} | Exactly n repetitions |
| {n,m} | Between n and m repetitions |
| [ ] | Character class |
| ( ) | Group |
| | | Alternation |
| \ | Escape or special sequence prefix |
The exact behavior of some regex features depends on the regular expression engine and its configuration, so patterns should always be tested in the environment where they will run.
The Dot Metacharacter
The dot is one of the most commonly used regex metacharacters. In common regex modes, it matches a single character other than a line terminator.
c.tThis can match cat, cot, cut, or other three-character sequences beginning with c and ending with t.
cat
cot
cut
cartcart does not match because the pattern contains exactly three positions: c, one character, and t.
Character Classes
A character class lets you specify a set or range of characters that can occupy one position in the pattern.
[abc]The class [abc] matches exactly one character that is either a, b, or c.
[a-z]This commonly represents one lowercase ASCII letter from a through z.
[0-9]This matches one ASCII digit from 0 through 9.
Negated Character Classes
A caret immediately after the opening bracket negates a character class.
[^0-9]This matches one character that is not an ASCII digit.
Shorthand Character Classes
Regex engines commonly provide shorthand character classes for frequently used character categories.
| Pattern | Common meaning |
|---|---|
| \d | A digit |
| \D | A non-digit |
| \w | A word character in the engine's defined character set |
| \W | A non-word character |
| \s | Whitespace |
| \S | Non-whitespace |
The exact definition of \w and related shorthand classes can vary between regex engines and Unicode modes. When matching international text, check the behavior of the target engine rather than assuming ASCII-only semantics.
Quantifiers
Quantifiers control how many times a preceding character, character class, group, or other regex element can occur.
| Quantifier | Meaning |
|---|---|
| * | Zero or more |
| + | One or more |
| ? | Zero or one |
| {3} | Exactly three |
| {2,5} | Between two and five |
| {2,} | Two or more |
The * Quantifier
The asterisk allows the preceding element to occur zero or more times.
ab*This matches a followed by zero or more b characters. Examples include a, ab, abb, and abbb.
The + Quantifier
The plus sign requires the preceding element to occur at least once.
ab+This matches ab, abb, abbb, and similar strings, but not a by itself.
The ? Quantifier
The question mark commonly makes the preceding element optional.
colou?rThis can match both color and colour.
Exact and Range Quantifiers
\d{4}This matches exactly four digits.
\d{2,4}This matches between two and four digits.
Greedy and Lazy Quantifiers
Many regex engines use greedy quantifiers by default. A greedy quantifier attempts to consume as much matching text as possible while still allowing the overall pattern to succeed.
<.*>When applied to text containing multiple angle-bracketed sections, the greedy dot-star can consume more text than intended.
A lazy quantifier asks the engine to prefer the shortest match that allows the complete pattern to succeed.
<.*?>The lazy version is often useful when matching the smallest possible section between delimiters, although structured formats such as HTML should generally be parsed with dedicated parsers when the task requires understanding their structure.
Anchors
Anchors match positions rather than ordinary characters. They are especially useful when a pattern should apply to an entire string or a specific location.
^helloThis commonly requires hello to appear at the beginning of the input or line, depending on the regex mode.
world$This commonly requires world to appear at the end of the input or line.
Matching an Entire String
Anchors are particularly useful when validating a whole input instead of searching for a matching substring.
^\d+$This pattern requires the input to consist entirely of one or more digits under the applicable regex semantics.
Word Boundaries
The \b assertion commonly matches a boundary between a word character and a non-word character, or at the beginning or end of the input where such a boundary exists.
\bcat\bThis can match cat as a separate word while avoiding a match inside words such as catalog.
Alternation
The pipe character provides alternatives. It means that one part of the pattern or another can match.
cat|dogThis matches either cat or dog.
Alternation becomes especially useful when combined with groups.
gr(een|ay)This can match green or gray.
Groups
Parentheses group multiple regex elements together. Groups are useful for applying quantifiers to multiple characters, creating alternatives, and capturing matched text.
(ab)+The plus quantifier applies to the entire group, so the pattern can match ab, abab, ababab, and so on.
Capturing Groups
A capturing group stores the portion of the input matched by the group so that the programming environment can retrieve it.
User: (\w+)If the input is User: Alice, the capturing group can contain Alice.
Capturing groups are particularly useful when extracting structured information from larger strings.
Non-Capturing Groups
A non-capturing group groups regex elements without creating a captured value.
(?:cat|dog)This groups the alternatives while avoiding an additional capture.
Escaping Special Characters
If you need to match a regex metacharacter literally, it often needs to be escaped with a backslash.
\.This matches a literal period rather than the dot wildcard.
\+This matches a literal plus sign.
Other characters such as brackets, parentheses, question marks, and backslashes can also require escaping depending on their position and the regex engine.
Regex Flags
Regex engines commonly provide flags that modify how a pattern is interpreted or how matching is performed. The available flags depend on the language or regex engine.
| Flag | Common purpose |
|---|---|
| g | Find multiple matches instead of stopping after the first |
| i | Case-insensitive matching |
| m | Change ^ and $ behavior for multiline input |
| s | Allow . to match line terminators in engines that support dotAll |
| u | Enable Unicode-aware regex behavior in engines that support it |
| y | Use sticky matching in engines that support it |
Flag syntax and exact semantics vary between environments. JavaScript, Python, PCRE, .NET, Java, and other engines share many regex concepts but do not implement every feature identically.
Case-Insensitive Matching
The case-insensitive flag allows letters to match without requiring the same case.
helloWith case-insensitive matching enabled, the pattern can match variations such as Hello, HELLO, and hello according to the engine's case-folding rules.
Multiline Matching
Multiline mode changes the behavior of line anchors in many regex engines. Instead of treating only the entire input as having one beginning and end, ^ and $ can apply to individual lines.
^Error:With the appropriate multiline mode, this can find lines beginning with Error: inside a larger multiline string.
Regex for Finding Numbers
A simple pattern for finding sequences of ASCII digits is:
\d+It can find values such as 42, 100, and 2026.
For a more specific integer format that must occupy the entire input:
^\d+$Real-world numeric validation can be more complicated when negative values, decimal separators, scientific notation, locale-specific formatting, or other requirements are involved.
Regex for Simple Email Patterns
Regex is often used to perform basic email-format checks, but email addresses have complex syntax and a simplistic pattern should not be treated as a complete implementation of email validation.
^[^\s@]+@[^\s@]+\.[^\s@]+$This pattern checks for a simple structure containing a non-whitespace local part, an @ symbol, a non-whitespace domain part, and a dot followed by additional characters.
Regex for URLs
URLs can also be matched with regular expressions, but complete URL syntax is complex. A simple pattern can be useful for detecting likely URLs in ordinary text.
https?://[^\s]+This matches text beginning with http:// or https:// followed by one or more non-whitespace characters.
For complete URL validation and parsing, dedicated URL APIs are usually preferable to maintaining an extremely large regular expression.
Regex for Whitespace
The \s shorthand is commonly used to find whitespace characters.
\s+This can match one or more consecutive whitespace characters. It is useful when normalizing text or replacing repeated spaces.
Regex for Duplicate Whitespace
[ \t]+This pattern specifically targets spaces and tab characters. Depending on the intended behavior, \s+ can be broader because it can include additional whitespace characters.
Regex for Extracting Text Between Delimiters
Regex can extract text between known delimiters when the structure is simple.
\[(.*?)\]This can capture text between square brackets in many common cases. The lazy quantifier helps prevent the match from consuming multiple bracketed sections at once.
Nested or deeply structured delimiters are much harder to handle reliably with ordinary regular expressions and may require a parser or engine-specific recursive features.
Regex for Find and Replace
One of the most useful applications of regex is replacing text based on a pattern instead of replacing one exact string.
\s+A replacement operation can turn every run of whitespace into a single space.
Before:
one two
three four
After:
one two
three fourThe exact replacement syntax depends on the editor, programming language, or regex tool being used.
Regex for Log Files
Regular expressions are useful for extracting structured fields from semi-structured logs.
\[(.*?)\]\s+(ERROR|WARN|INFO)This example can identify a bracketed value followed by one of several log levels. Production log formats should be matched against the actual format used by the application.
Lookahead and Lookbehind
Some regex engines support lookahead and lookbehind assertions. These allow a pattern to require surrounding text without consuming that surrounding text as part of the main match.
\d+(?= USD)This can match digits that are immediately followed by a space and USD while leaving USD outside the matched portion.
Lookbehind provides the corresponding ability to inspect text before the current position. Support varies by regex engine, so portability should be checked before using these features.
Backreferences
Backreferences allow a regex to refer to text previously captured by a group.
\b(\w+)\s+\1\bThis pattern can find repeated adjacent words such as 'the the'. The exact behavior of word characters depends on the regex engine.
Regex Engine Differences
Regular expression syntax is not completely standardized across all programming languages and tools. Different engines support different features and may interpret some constructs differently.
| Environment | Examples of regex context |
|---|---|
| JavaScript | RegExp and String methods |
| Python | re module |
| Java | java.util.regex |
| C# | .NET regular expressions |
| PHP | PCRE-based functions |
| Command line | grep, sed, awk and related tools |
A pattern that works in JavaScript is not automatically guaranteed to work unchanged in every other regex engine. Before moving a pattern between languages, check syntax, flags, Unicode behavior, lookbehind support, named groups, and replacement syntax.
Regex in JavaScript
JavaScript provides the RegExp object and several string methods that work with regular expressions.
const pattern = /\d+/g;
const matches = "Order 42 and order 108".match(pattern);
console.log(matches);The g flag requests multiple matches. JavaScript also supports methods such as test, match, matchAll, replace, replaceAll, search, and split in combination with regular expressions.
Regex in Python
import re
matches = re.findall(r"\d+", "Order 42 and order 108")
print(matches)Python's raw string syntax is commonly used for regex patterns because it reduces confusion caused by Python string escaping and regex escaping occurring at the same time.
Regex Escaping in Programming Languages
When a regex is written inside a programming-language string, two layers of escaping can be involved: the programming language's string syntax and the regex engine's syntax.
const pattern = /\d+/;Regex literals can make patterns easier to read in languages that support them. When a pattern is represented as a string, the backslash may need to be escaped for the programming language as well.
const pattern = "\\d+";This distinction is a common source of bugs when constructing regular expressions dynamically.
Regex for Validation vs Searching
A regex used to find text inside a document has different requirements from a regex used to validate an entire input.
| Task | Typical approach |
|---|---|
| Find a number | \d+ |
| Validate an integer | ^\d+$ or an engine-appropriate full-match operation |
| Find an email-like string | A pattern that searches for the desired structure |
| Validate an email input | A complete input check followed by practical verification |
Anchors or full-match APIs are especially important for validation because a successful substring match does not necessarily mean that the entire input is valid.
Regex and Unicode
Unicode introduces additional considerations for letters, digits, word characters, case conversion, and character properties. A pattern designed around ASCII ranges such as [a-z] does not automatically cover every letter used in human languages.
[a-z]+This is an ASCII-oriented range. For international text, use Unicode-aware features supported by the target regex engine when appropriate.
Some engines support Unicode property escapes such as \p{L} for letters when Unicode mode is enabled.
\p{L}+Support for Unicode properties varies, so the target engine should always be checked.
Regex Performance
Most regex patterns are fast enough for ordinary text processing, but poorly designed expressions can become expensive on large or adversarial inputs. Certain combinations of nested quantifiers and ambiguous alternatives can cause excessive backtracking in backtracking-based engines.
^(a+)+$Patterns with nested repetition can be problematic depending on the regex engine and input. Performance-sensitive applications should test patterns against realistic and deliberately difficult inputs.
How to Write Better Regex
- Start with the simplest pattern that solves the actual problem.
- Use character classes instead of long lists of alternatives when appropriate.
- Use anchors or full-match APIs for whole-input validation.
- Use groups to make repeated or alternative structures explicit.
- Prefer non-capturing groups when captures are not needed.
- Avoid unnecessary .* expressions.
- Use lazy quantifiers when the shortest relevant match is intended.
- Test edge cases rather than only the happy path.
- Check Unicode behavior when processing international text.
- Check regex engine compatibility before sharing a pattern between languages.
- Measure performance for large or untrusted inputs.
Testing a Regular Expression
A regex tester is one of the easiest ways to understand how a pattern behaves. Enter representative input, apply the pattern, and inspect the matches and capture groups.
Good regex tests should include normal cases, boundary cases, invalid input, empty input, unexpected characters, multiline text, and Unicode text when relevant.
^\d{4}-\d{2}-\d{2}$This pattern describes a simple YYYY-MM-DD-shaped string. It does not by itself prove that the date is a real calendar date. For example, additional application logic may be needed to reject impossible dates.
When Not to Use Regex
Regex is powerful, but it is not always the right tool. Structured data is often easier and safer to process with a dedicated parser.
- Use a JSON parser for JSON.
- Use an XML parser for complex XML documents.
- Use a URL API for URL parsing.
- Use a date/time library or built-in date APIs for complex date validation.
- Use an HTML parser for HTML structure.
- Use a programming-language parser for source-code syntax.
Regex is excellent for searching and transforming text with predictable patterns, but increasingly complex structured data often deserves a parser that understands its grammar.
Common Regex Mistakes
Mistake 1: Forgetting Anchors During Validation
A pattern such as \d+ can successfully match digits inside an otherwise invalid string. If the entire input must contain only digits, use appropriate anchors or a full-match API.
Mistake 2: Using .* Everywhere
Dot-star is convenient, but it can make patterns overly broad, difficult to understand, and sometimes expensive to execute. Prefer a more specific character class or structure when possible.
Mistake 3: Ignoring Engine Differences
A regex copied from one programming language may use features that another engine does not support. Always test the final pattern in the target environment.
Mistake 4: Treating a Regex as Complete Validation
A regex can validate a syntactic shape without proving that the underlying value is semantically correct. Date strings, URLs, email addresses, identifiers, and numeric values can require additional validation.
Mistake 5: Forgetting Escaping Layers
When a regex is stored inside a programming-language string, both string escaping and regex escaping can affect the final pattern. Prefer regex literals or raw-string mechanisms when the language provides them and they improve readability.
A Regex Learning Path
A practical way to learn regular expressions is to start with literal matching and gradually introduce more expressive features.
- Learn literal characters and escaping.
- Learn the dot wildcard.
- Learn character classes.
- Learn shorthand classes such as \d and \s.
- Learn quantifiers.
- Learn anchors and boundaries.
- Learn alternation.
- Learn capturing and non-capturing groups.
- Learn flags.
- Learn lookarounds and backreferences.
- Learn engine-specific Unicode features.
- Learn regex performance and backtracking behavior.
Useful Regex Tools
Several types of text-processing tools are useful when working with regular expressions. A regex tester helps verify patterns interactively, while regex generators can help construct initial patterns. Find-and-replace tools are useful for applying a tested pattern to larger text, and text-cleaning tools can simplify repetitive transformations.
- Regex testers for checking matches and capture groups.
- Regex generators for creating initial patterns from examples.
- Find-and-replace tools for applying regex transformations.
- Text cleaners for normalizing and preparing input.
- Duplicate-line removal tools for cleaning repeated text.
Frequently Asked Questions
What is a regular expression?
A regular expression is a pattern language used to search, match, extract, validate, and replace text according to defined rules.
What does \d mean in regex?
\d commonly matches a digit. The exact character set can depend on the regex engine and its Unicode configuration.
What does .* mean in regex?
The dot matches a character in common regex modes, while * allows the preceding element to repeat zero or more times. Together, .* can match a broad sequence of characters, although its exact behavior depends on flags and the engine.
What is the difference between * and +?
* allows zero or more occurrences of the preceding element, while + requires at least one occurrence.
What does ^ mean in regex?
^ commonly represents the beginning of the input or line. Its exact behavior depends on the regex engine and multiline configuration.
What does $ mean in regex?
$ commonly represents the end of the input or line. Multiline behavior and details around line terminators can vary between engines.
Can regex validate an email address?
Regex can check a simplified email-like structure, but complete email syntax is complex. Practical applications should combine format checks with appropriate verification rather than relying on one short regex alone.
Is regex the same in every programming language?
No. Many regex concepts are shared, but engines differ in supported features, Unicode behavior, flags, escaping rules, and replacement syntax.
Conclusion
Regular expressions provide a compact way to describe patterns in text. Character classes define sets of possible characters, quantifiers control repetition, anchors control position, groups organize patterns, alternation provides alternatives, and flags modify matching behavior.
Regex is useful for searching, extracting, replacing, and validating text, but patterns should be designed around the actual problem. Whole-input validation requires anchors or full-match APIs, international text requires attention to Unicode behavior, and complex structured data is often better handled by a dedicated parser.
The most effective way to become comfortable with regex is to build patterns incrementally and test them against realistic input. Once the basic syntax becomes familiar, regular expressions become a practical tool for many everyday development and text-processing tasks.