Regex Lookahead and Lookbehind
A practical guide to regex lookahead and lookbehind, covering positive and negative lookarounds, syntax, practical patterns, JavaScript examples, and common mistakes.
Regex lookahead and lookbehind are advanced features that let a regular expression check surrounding text without including that surrounding text in the main match. Together, these features are commonly called lookarounds.
A normal regex consumes the characters that it matches. A lookaround works differently: it checks whether a condition is satisfied at a particular position, but the characters examined by the assertion are not consumed as part of the main match.
This makes lookarounds useful for matching values only when they appear in a particular context. They can be used to find numbers followed by a unit, words preceded by a specific prefix, values that must not have a certain suffix, and many other patterns that would otherwise require more complicated extraction logic.
Lookaround Quick Reference
| Syntax | Type | Meaning |
|---|---|---|
| (?=pattern) | Positive lookahead | The pattern must follow the current position |
| (?!pattern) | Negative lookahead | The pattern must not follow the current position |
| (?<=pattern) | Positive lookbehind | The pattern must precede the current position |
| (?<!pattern) | Negative lookbehind | The pattern must not precede the current position |
The first two forms look forward in the string, while the last two look backward. Positive lookarounds require a condition to be true; negative lookarounds require it to be false.
What Is a Lookaround?
A lookaround is a zero-width assertion. It checks whether a particular pattern exists at a position without consuming the characters used for that check.
\d+(?=px)This pattern looks for one or more digits that are immediately followed by px. The digits are the main match, while px is only checked by the lookahead.
For example, when applied to 16px, the main match is 16 rather than 16px.
| Input | Main match | Checked context |
|---|---|---|
| 16px | 16 | px |
| 20em | No match | The following text is em, not px |
| 8px | 8 | px |
Positive Lookahead
A positive lookahead uses (?=pattern). It succeeds when the specified pattern can be found immediately after the current matching position.
\d+(?=px)The expression matches digits only when px immediately follows them.
16px 20em 32px 50% 8pxApplied to this text, the expression finds 16, 32, and 8. It does not include px in the returned matches.
Why Use Positive Lookahead?
Positive lookahead is useful when the surrounding text is important for deciding whether something should match, but that surrounding text should remain outside the result.
- Match a number only when it has a specific unit.
- Match a word only when a particular suffix follows it.
- Match text only when a delimiter appears next.
- Validate a condition without consuming the checked characters.
- Apply multiple requirements to the same position.
Positive Lookahead with a Word
\w+(?=\.com)This matches a word-like sequence only when .com follows immediately.
example.com test.org site.comThe main matches are example and site. The .com suffix is checked but is not included in the match.
Positive Lookahead with End-of-String Validation
Lookahead can also be useful when a value must satisfy several independent conditions.
^(?=.*[A-Z])(?=.*\d).+$This pattern requires the entire input to contain at least one uppercase ASCII letter and at least one digit. The lookaheads inspect the input for the required conditions, while .+ performs the main match.
This style is common in password validation. Additional requirements can be represented with additional lookaheads.
Negative Lookahead
A negative lookahead uses (?!pattern). It succeeds when the specified pattern does not occur immediately after the current position.
\d+(?!px)This attempts to match digits that are not immediately followed by px.
Negative lookahead is useful when a value is allowed in general but must be excluded when a particular context follows it.
Negative Lookahead Examples
\w+(?!\.com)The expression checks that .com does not immediately follow the current position. Because regex matching can advance through a string in many ways, patterns using negative lookahead should be tested carefully to make sure the assertion is applied at the intended position.
A more controlled example is an anchored validation pattern that excludes a known prefix.
^(?!admin$)[a-z0-9_-]+$This accepts lowercase letters, digits, underscores, and hyphens, but rejects the complete value admin.
Positive Lookbehind
A positive lookbehind uses (?<=pattern). It requires the specified pattern to appear immediately before the current position.
(?<=\$)\d+This matches digits only when they are immediately preceded by a dollar sign. The dollar sign is not included in the main match.
$25 €30 $100 50The main matches are 25 and 100. The preceding dollar signs are only used as conditions.
Why Use Positive Lookbehind?
Lookbehind is useful when the condition that determines whether a value should match is located before the value rather than after it.
- Match a number only when it has a specific currency symbol before it.
- Match a word only after a particular prefix.
- Extract values without including their preceding delimiter.
- Check a context condition without consuming the context.
- Create cleaner extraction patterns when the required context comes before the target.
Positive Lookbehind with Units
(?<=#)[0-9A-Fa-f]{6}This matches six hexadecimal characters only when they are immediately preceded by #. For example, it can extract 12a4ef from #12a4ef without including the # in the result.
Negative Lookbehind
A negative lookbehind uses (?<!pattern). It succeeds when the specified pattern does not appear immediately before the current position.
(?<!\$)\d+This attempts to match digits that are not immediately preceded by a dollar sign.
$25 30 €40 50The intended use here is to exclude numbers whose immediate preceding character is $. More complicated surrounding text may require a more carefully constrained pattern.
The Four Main Lookaround Types
| Type | Syntax | Condition |
|---|---|---|
| Positive lookahead | (?=X) | X must come next |
| Negative lookahead | (?!X) | X must not come next |
| Positive lookbehind | (?<=X) | X must come before |
| Negative lookbehind | (?<!X) | X must not come before |
The easiest way to remember them is to separate direction from polarity. Lookahead means checking forward; lookbehind means checking backward. A normal equals sign means the condition must exist, while an exclamation mark means it must not exist.
Lookahead vs Lookbehind
| Requirement | Pattern |
|---|---|
| Digits followed by px | \d+(?=px) |
| Digits preceded by $ | (?<=\$)\d+ |
| Text not followed by .com | (?!\.com) |
| Text not preceded by $ | (?<!\$) |
The main practical difference is where the condition is located relative to the text you want to match.
Lookarounds Are Zero-Width Assertions
The term zero-width means that the lookaround checks a condition without advancing the main matching position over the checked characters.
\d+(?=USD)If this matches 100 in 100USD, the regex engine has verified that USD is next, but USD is not part of the returned match.
This is different from a normal sequence of tokens.
\d+USDThe second pattern consumes and returns both the number and USD as part of the complete match.
Lookahead vs Normal Matching
| Pattern | Input | Main match |
|---|---|---|
| \d+USD | 100USD | 100USD |
| \d+(?=USD) | 100USD | 100 |
| (?<=USD)\d+ | USD100 | 100 |
This distinction is one of the main reasons lookarounds are useful for extraction and text transformation.
Multiple Lookaheads
Several lookaheads can be placed at the same position to require multiple conditions simultaneously.
^(?=.*[A-Z])(?=.*[a-z])(?=.*\d).+$This pattern requires at least one uppercase ASCII letter, one lowercase ASCII letter, and one digit somewhere in the input.
The lookaheads do not consume the characters they inspect. The final .+ is responsible for matching the actual content.
Password Validation with Lookahead
Lookaheads are frequently used for password validation because several independent conditions can be checked at the beginning of the string.
^(?=.*[A-Z])(?=.*[a-z])(?=.*\d).{8,}$This example requires at least eight characters, one uppercase ASCII letter, one lowercase ASCII letter, and one digit.
Combining Positive and Negative Lookaheads
Positive and negative assertions can be combined to express both required and forbidden conditions.
^(?!.*password)(?=.*\d).{8,}$This example requires at least eight characters, requires a digit, and rejects an input containing the literal sequence password anywhere in the string.
When combining several assertions, keep each condition understandable. A large collection of lookaheads can become difficult to debug and maintain.
Lookahead with Character Classes
Lookarounds often work together with character classes to describe precise conditions.
[A-Za-z]+(?=\d)This matches letters only when digits immediately follow them.
(?<=\$)[0-9]+This combines a lookbehind with a digit class to find ASCII digits immediately following a dollar sign.
Lookarounds with Quantifiers
Quantifiers can appear inside a lookaround. This lets the assertion describe a larger context.
\w+(?=\s+USD)This matches a word-like sequence only when one or more whitespace characters followed by USD occur next.
(?<=USD\s)\d+This matches digits that are preceded by USD followed by whitespace, in regex engines that support this lookbehind form.
Lookarounds with Alternation
Alternation can be placed inside a lookaround when several possible contexts should satisfy the same condition.
\d+(?=px|em|rem)This matches a number when px, em, or rem follows it.
(?<=\$|€)\d+This illustrates the idea of matching numbers preceded by one of two currency symbols. Lookbehind restrictions vary between regex engines, so verify the exact syntax and supported lookbehind behavior in the target environment.
Lookarounds and Capturing Groups
Lookarounds and capturing groups solve different problems. A capturing group records text matched by the group, while a lookaround primarily asserts that surrounding text satisfies a condition.
(\d+)(?=px)Here, the digits are both the main match and capture group 1. The px suffix is checked by the lookahead but is not consumed.
A useful pattern is to use lookarounds to define context and capturing groups when a specific part of the resulting match must be extracted separately.
Lookarounds in Find and Replace
Lookarounds can be useful in find-and-replace operations when only part of a value should be replaced while surrounding context must remain untouched.
\d+(?=px)Suppose a document contains 10px, 20px, and 30em. This pattern selects only the numbers associated with px, leaving the units outside the match. A replacement can therefore modify the numbers without having to capture and reconstruct the unit.
Removing a Prefix Without Matching It
A positive lookbehind can select text after a known prefix while leaving the prefix untouched.
(?<=ID:)\d+Given ID:12345, the main match is 12345. The ID: prefix is used only as context.
Finding Text Before a Delimiter
A positive lookahead can be used to match text only when a delimiter appears immediately afterward.
[^,]+(?=,)This matches a non-comma sequence when a comma follows it. The comma itself is not consumed by the lookahead.
Lookahead for Excluding a Suffix
Negative lookahead is useful when a particular suffix should prevent a match.
\b\w+\b(?!\.tmp)The intention is to match a word-like token when .tmp does not immediately follow it. The exact behavior depends on where the engine can establish the word boundary, so patterns like this should be tested against representative input.
Lookbehind for Excluding a Prefix
Negative lookbehind performs the corresponding check in the opposite direction.
(?<!#)\b[A-Fa-f0-9]{6}\bThis illustrates matching a six-character hexadecimal-looking value when it is not immediately preceded by #.
Lookarounds and Boundaries
Lookarounds can be combined with anchors and word boundaries to control exactly where a condition applies.
\badmin\b(?!\.)This looks for the complete word admin when a period does not immediately follow it.
Combining zero-width assertions can produce very precise patterns, but it also makes the regex harder to read. Use the simplest expression that clearly communicates the requirement.
Variable-Length Lookbehind
One of the important differences between regex engines is how much flexibility they allow inside lookbehind assertions. Some engines historically required fixed-length lookbehind, while others support broader forms.
For example, a lookbehind containing a fixed number of characters is generally easier to support consistently than one containing arbitrary repetition.
(?<=USD\s)\d+The lookbehind here has a fixed length: USD followed by one whitespace character. More complicated expressions may not be accepted by every engine.
Lookahead and Lookbehind in JavaScript
Modern JavaScript supports positive and negative lookahead as well as positive and negative lookbehind in current environments.
const pattern = /\d+(?=px)/g;
const input = "16px 20em 32px";
console.log(input.match(pattern));
// ["16", "32"]The lookahead checks for px while the match result contains only the numbers.
const pattern = /(?<=\$)\d+/g;
const input = "$25 €30 $100";
console.log(input.match(pattern));
// ["25", "100"]The lookbehind checks for a dollar sign without including it in the returned matches.
Browser and Runtime Compatibility
Positive and negative lookahead have broad support across modern regex implementations. Lookbehind support is more dependent on the target environment, especially when older browsers or specialized runtimes are involved.
If compatibility with older environments matters, a lookbehind-based pattern may need to be rewritten using capturing groups, lookahead, string processing, or another approach.
Replacing Lookbehind with Capturing Groups
When lookbehind is unavailable, a common alternative is to capture the required prefix and then use the captured portion during replacement or post-processing.
(\$)(\d+)Instead of using (?<=\$)\d+, this pattern captures both the dollar sign and the number. Application code can then use the second capture as the value while preserving the first capture when necessary.
The best alternative depends on whether the goal is extraction, validation, replacement, or simply finding a match.
Common Lookaround Mistakes
- Forgetting that lookarounds do not consume the text they inspect.
- Expecting a lookahead's checked text to appear in the main match.
- Confusing positive and negative assertions.
- Using lookbehind without checking whether the target regex engine supports it.
- Assuming all engines support the same variable-length lookbehind behavior.
- Creating too many lookaheads in a single validation pattern.
- Using a negative lookahead without checking exactly where the assertion is evaluated.
- Replacing a simple string operation with a complicated lookaround unnecessarily.
- Using lookarounds to parse complex structured data.
- Failing to test edge cases around boundaries and adjacent characters.
Lookaround Debugging Strategy
Lookarounds can be difficult to debug because the text they inspect does not appear in the main match. Testing a pattern in small steps makes the behavior easier to understand.
- Start by matching the main value without any lookaround.
- Add the positive or negative assertion.
- Test an input where the assertion should succeed.
- Test an input where the assertion should fail.
- Check whether the checked context is immediately adjacent to the match.
- Verify whether the assertion should consume anything. Lookarounds normally should not.
- Test the pattern at the beginning and end of the input.
- Check behavior with punctuation, whitespace, and Unicode characters.
- Verify compatibility with the target regex engine.
Lookarounds and Performance
Lookarounds are powerful, but a large number of nested or ambiguous assertions can make a regex harder for both humans and the regex engine to process. The performance impact depends on the complete expression, input size, regex engine, and possible backtracking paths.
Simple assertions such as \d+(?=px) are usually straightforward. More complicated validation expressions with multiple broad lookaheads and nested repetition deserve additional testing, especially when they process untrusted input.
When Lookarounds Are Useful
- Matching a value only when a specific suffix follows it.
- Matching a value only when a specific prefix precedes it.
- Excluding specific prefixes or suffixes.
- Validating several independent requirements.
- Extracting values without consuming their surrounding delimiters.
- Performing precise find-and-replace operations.
- Checking context without including that context in the result.
When Not to Use Lookarounds
Lookarounds should not be added simply because they are available. If a simple regex or normal string operation communicates the requirement more clearly, it is often preferable.
For example, checking whether a string starts with a known prefix can often be handled more clearly with startsWith() in JavaScript. Similarly, checking whether a string ends with a known suffix may be simpler with endsWith().
Lookarounds are most valuable when the condition and the match need to be kept logically separate inside the same regex operation.
Practical Lookaround Patterns
| Task | Pattern | Main match |
|---|---|---|
| Digits before px | \d+(?=px) | 16 from 16px |
| Digits after $ | (?<=\$)\d+ | 25 from $25 |
| Hex after # | (?<=#)[0-9A-Fa-f]{6} | 12a4ef from #12a4ef |
| Text before comma | [^,]+(?=,) | Text before a comma |
| Number with required digit | ^(?=.*\d).+$ | Complete input |
| Exclude exact value | ^(?!admin$).+$ | Any value except admin |
Lookarounds vs Capturing Groups
| Feature | Capturing group | Lookaround |
|---|---|---|
| Basic syntax | (pattern) | (?=pattern) or related forms |
| Consumes matched text | Yes | No |
| Stores captured value | Yes | Not as the main purpose |
| Useful for extraction | Yes | Yes, indirectly |
| Useful for context checks | Sometimes | Yes |
| Useful for backreferences | Yes | Can contain captures depending on engine |
Capturing groups answer the question: which part of the matched text do I want to save? Lookarounds answer a different question: under what surrounding condition should this position match?
A Practical Workflow for Writing Lookarounds
- Identify the exact text that should appear in the final match.
- Identify the surrounding text that should only act as a condition.
- If the condition comes after the target, consider lookahead.
- If the condition comes before the target, consider lookbehind.
- Use a positive assertion when the context must exist.
- Use a negative assertion when the context must not exist.
- Test the positive and negative cases separately.
- Check compatibility with the target regex engine.
- Prefer a simpler string operation if regex is unnecessary.
Frequently Asked Questions
What is a regex lookahead?
A lookahead checks whether a pattern exists immediately after the current position without consuming that text. Positive lookahead uses (?=...), while negative lookahead uses (?!...).
What is a regex lookbehind?
A lookbehind checks whether a pattern exists immediately before the current position without including that text in the main match. Positive lookbehind uses (?<=...), while negative lookbehind uses (?<!...).
What does zero-width assertion mean?
It means the assertion checks a condition without consuming characters as part of the main match. The regex position is effectively checked without adding the asserted text to the result.
What is the difference between lookahead and lookbehind?
Lookahead checks forward from the current position, while lookbehind checks backward. Both can be positive or negative depending on whether the surrounding pattern must exist or must not exist.
Why does my lookahead not include the checked text in the match?
That is the intended behavior. A lookahead is a zero-width assertion, so it verifies the following text without consuming it. Use ordinary regex tokens if the surrounding text should be part of the match.
Is lookbehind supported in JavaScript?
Modern JavaScript environments support positive and negative lookbehind. Compatibility with older environments should still be checked before using it in a production application.
Can lookarounds be combined?
Yes. Multiple positive and negative lookaheads can be combined to express several conditions. Lookarounds can also contain groups, character classes, alternation, and quantifiers where supported by the target regex engine.
Are lookarounds good for password validation?
They can express multiple format requirements, such as requiring a digit and an uppercase letter. However, regex validation only checks format-related rules; it does not determine password strength or security by itself.
Conclusion
Regex lookahead and lookbehind provide a precise way to match text based on its surrounding context without consuming that context. Positive lookahead checks what comes next, negative lookahead checks what must not come next, positive lookbehind checks what comes before, and negative lookbehind checks what must not come before.
Lookarounds are especially useful for extraction, validation, and find-and-replace tasks. They allow a regex to keep the condition separate from the text that should actually be returned, which can make some patterns considerably cleaner.
At the same time, lookarounds are an advanced feature. Engine compatibility, especially for lookbehind, should always be considered. When a simpler regex or ordinary string method can express the same requirement clearly, the simpler solution is usually easier to maintain.