Ctrl + K
Regex19 min read

Regex Groups and Capturing

A practical guide to regex groups and capturing, including numbered groups, non-capturing groups, named groups, backreferences, replacement patterns, and common mistakes.

Published: 2026-10-05

Regex groups let you treat several parts of a regular expression as a unit and, when needed, capture the text matched by that group. They are one of the most useful features for extracting structured information from strings.

The basic syntax is a pair of parentheses: (pattern). By default, parentheses create a capturing group. The regex engine records the text matched by that group so the application can access it after a successful match.

Groups are useful for much more than extraction. They can control quantifiers, organize alternatives, create optional sections, support backreferences, and work with lookaheads and lookbehinds. When you do not need to capture a value, non-capturing groups such as (?:pattern) avoid creating unnecessary capture results.

Regex Groups Quick Reference

SyntaxPurposeExample
(pattern)Capturing group(\d+)
(?:pattern)Non-capturing group(?:https?://)
(?<name>pattern)Named capturing group(?<year>\d{4})
\1Backreference to group 1(\w+)\s+\1
\k<name>Backreference to named group(?<word>\w+)\s+\k<word>
(?=pattern)Positive lookahead\d+(?=px)
(?!pattern)Negative lookahead\d+(?!px)
(?<=pattern)Positive lookbehind(?<=\$)\d+
(?<!pattern)Negative lookbehind(?<!\$)\d+

What Is a Regex Group?

A regex group is a section of a regular expression enclosed in parentheses. Parentheses tell the regex engine to treat the enclosed expression as a unit.

(abc)

This pattern matches abc and creates a capturing group containing the matched text abc.

Grouping becomes especially useful when a larger expression contains several operations that should be treated together.

(ab)+

Here, the + quantifier applies to the entire group ab. The expression can therefore match ab, abab, ababab, and so on.

Capturing Groups

A normal parenthesized group is a capturing group. The regex engine stores the substring matched by that group separately from the complete match.

(\d{4})-(\d{2})-(\d{2})

This pattern can capture the year, month, and day separately from a date such as 2026-09-03.

ResultValue
Full match2026-09-03
Group 12026
Group 209
Group 303

The exact API used to access these groups depends on the programming language. JavaScript, Python, .NET, Java, and other environments expose capture results through different interfaces, but the underlying regex concept is the same.

Why Capturing Groups Are Useful

Capturing groups are particularly useful when the input contains structured pieces that need to be extracted separately.

  • Extracting parts of dates.
  • Extracting protocol, hostname, and path from simple URLs.
  • Extracting a username from a structured identifier.
  • Extracting numbers and their units.
  • Extracting quoted text.
  • Reusing previously matched text with backreferences.
  • Transforming text during find-and-replace operations.
  • Parsing simple log formats.

Numbered Capture Groups

Capturing groups are normally numbered according to the order of their opening parentheses, starting with 1.

(\d{4})-(\d{2})-(\d{2})

The first opening capturing parenthesis creates group 1, the second creates group 2, and the third creates group 3.

GroupPatternExample
1(\d{4})2026
2(\d{2})09
3(\d{2})03

Numbered groups are convenient for short patterns, but they can become difficult to maintain when a regex contains many nested groups.

Nested Capturing Groups

Groups can be placed inside other groups. Each capturing pair of parentheses still receives its own number.

((\d{4})-(\d{2}))

The outer group captures the complete year-month section, while the inner groups capture the year and month separately.

GroupCaptured value
12026-09
22026
309
⚠️ Adding or removing capturing parentheses can change the numbers of later groups. This is one reason large regex patterns can become difficult to maintain when they rely heavily on numbered backreferences.

Non-Capturing Groups

A non-capturing group uses the syntax (?:pattern). It groups an expression without storing the matched text as a numbered capture.

(?:https?://)([A-Za-z0-9.-]+)

In this example, the protocol is grouped because the ? quantifier and the surrounding structure need it, but only the hostname is captured.

Non-capturing groups are especially useful when parentheses are needed for structure but the application does not need the group's contents.

Capturing vs Non-Capturing Groups

SyntaxCaptures?Typical use
(abc)YesExtract or reuse abc
(?:abc)NoGroup abc for structure
(?<name>abc)YesExtract abc with a meaningful name

A useful rule is simple: use a capturing group when you need its value later; use a non-capturing group when you only need grouping behavior.

Grouping Alternatives

Groups are essential when a quantifier or alternative should apply to more than one token.

(cat|dog)

This matches either cat or dog and captures whichever alternative matched.

(cat|dog)s?

Now the optional s applies to the result of the alternative. The expression can match cat, cats, dog, or dogs.

Grouping Changes Quantifier Scope

Parentheses can completely change what a quantifier applies to.

ab+

Here, + applies only to b, so the pattern matches a followed by one or more b characters.

(ab)+

Here, + applies to the complete sequence ab, so the sequence itself can repeat.

PatternExamples
ab+ab, abb, abbb
(ab)+ab, abab, ababab

Grouping Optional Sections

Groups can make an entire section optional.

https?://

In this example, the s applies only to the character before it. A group becomes useful when multiple characters should be optional together.

(?:www\.)?

This makes the complete www. prefix optional rather than only the final character.

^(?:www\.)?[A-Za-z0-9.-]+$

This illustrates how a non-capturing group can organize an optional section without adding an unnecessary capture.

Named Capturing Groups

Named groups assign a meaningful name to a capture instead of requiring the application to remember its numeric position.

(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})

The groups are now identified as year, month, and day. This can make code easier to understand and reduce the risk of confusing group numbers.

Named group syntax varies between regex engines. The JavaScript syntax above is supported in modern JavaScript environments, while other languages may use different notation.

Named Groups in JavaScript

const pattern = /(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/;
const match = pattern.exec("2026-09-03");

console.log(match.groups.year);
// "2026"

console.log(match.groups.month);
// "09"

Named groups are particularly useful when the regex has several captures and the application needs to refer to them individually.

Named Group Backreferences

A named group can be referenced later in the same regex using the syntax \k<name> in engines that support named backreferences.

(?<word>\w+)\s+\k<word>

This pattern looks for the same word twice with whitespace between the occurrences. It can detect simple repeated words such as hello hello.

What Is a Backreference?

A backreference tells the regex engine to match the exact text previously captured by a group. It is different from simply repeating the same pattern.

(\w+)\s+\1

Here, \1 refers to the text captured by group 1. If group 1 matches hello, the backreference requires hello again.

InputResult
hello helloMatches
hello worldDoes not match
test testMatches
test testingDoes not match

Backreference vs Repeating a Pattern

This distinction is important. A repeated pattern only requires the same rules to be satisfied again. A backreference requires the exact previously captured text.

(\w+)\s+\1

The expression above requires the second word to be identical to the first. A pattern such as \w+\s+\w+ would allow two different words.

Numbered Backreferences

Numbered backreferences use a backslash followed by the capture number.

(.)\1

This looks for two identical consecutive characters. It can match aa, 11, or !!.

(\w+)\s+\1

This looks for a repeated word separated by whitespace.

Named vs Numbered Groups

ApproachExampleAdvantage
Numbered group(\d{4})Short and widely supported
Named group(?<year>\d{4})Meaningful and easier to maintain
Numbered reference\1Compact
Named reference\k<year>Clearer in complex patterns

For short expressions, numbered groups are often sufficient. For larger expressions with several captures, named groups can make the intent much easier to understand.

Groups in Find and Replace

Capturing groups are particularly powerful in find-and-replace operations. The captured text can often be reused in the replacement string.

(\d{4})-(\d{2})-(\d{2})

Suppose the input contains dates in YYYY-MM-DD format and the desired output is DD/MM/YYYY. A replacement can reuse the captured components in a different order.

Input:       2026-09-03
Replacement: $3/$2/$1
Result:      03/09/2026

Replacement syntax differs between tools and programming languages. JavaScript replacement strings commonly use $1, $2, and named replacement forms, while other environments may use different notation.

Capturing Groups for Text Transformation

Groups make it possible to rearrange matched components without manually processing every string. This is useful for simple format conversions.

^(\w+),\s*(\w+)$

This pattern captures two fields separated by a comma. A replacement can then place the second capture before the first.

For example, a value such as Smith, John can be transformed into John Smith. More complicated names and structured data may require a dedicated parser rather than relying on a single regex.

Groups and Alternation

Groups and alternation work together to represent several possible structures.

(https?|ftp)://

The group captures either http, https, or ftp.

(?:https?|ftp)://

If the protocol does not need to be extracted, a non-capturing group avoids creating an unnecessary capture.

Groups and Character Classes

Character classes define a set of possible characters, while groups define a larger expression that can contain classes, literals, quantifiers, and alternatives.

([A-Za-z]+)-([0-9]+)

The first group captures one or more letters, and the second captures one or more digits.

ComponentRole
[A-Za-z]Allowed characters
+One or more
([A-Za-z]+)Capture the complete letter sequence
-Literal separator
([0-9]+)Capture the digit sequence

Groups and Quantifiers

A group can itself be quantified, and a quantifier can also appear inside a group. These produce different results.

(ab){2}

This matches abab. The group ab is repeated twice.

(ab{2})

This instead matches abb because the quantifier applies only to b inside the group.

Repeated Capturing Groups

A subtle behavior appears when a capturing group is repeated. The group usually retains only the text from its final successful iteration in the ordinary match result.

(\d)+

This can match a sequence such as 12345, but the single capture group does not normally provide all five digits as separate captures. It represents the capture from the last iteration according to the regex engine's match-result rules.

⚠️ If you need every repeated item separately, do not assume a repeated capturing group will produce an array of captures. Use an appropriate global matching API, repeated extraction, or a different pattern structure.

Groups That Do Not Participate in a Match

Optional capturing groups may not participate in a particular match. In that situation, the capture result is typically undefined, null, or an equivalent empty state depending on the programming language.

^(\d+)(?:\.(\d+))?$

The second capture exists only when a decimal portion is present. This allows the pattern to match both 42 and 42.75 while capturing the optional fractional part when available.

Groups and Lookarounds

Lookaheads and lookbehinds use parentheses but behave differently from ordinary capturing groups. They assert that a pattern does or does not occur without consuming those characters as part of the overall match.

\d+(?=px)

This can match digits that are immediately followed by px. The px text is checked but is not included in the main match.

Lookarounds are a separate regex feature with their own syntax and behavior. They can contain groups when extraction or more complex conditions are required.

Groups for Simple Structured Extraction

A common use of capturing groups is extracting predictable fields from text such as log entries.

^(\d{4}-\d{2}-\d{2})\s+(\w+)\s+(.*)$

This example captures a date, a word-like field, and the remaining text. The exact pattern should be adapted to the actual log format because \w does not represent every possible character in every regex engine.

Groups for URL Components

Groups can also be used to extract simple components from URLs.

^(https?)://([^/]+)(/.*)?$

The first group captures the protocol, the second captures the host portion, and the third optionally captures the path.

⚠️ URL syntax has many edge cases. A small regex like this can be useful for a controlled format, but it should not be treated as a complete URL parser.

Groups for Numbers and Units

Groups make it easy to separate a numeric value from its unit.

([0-9]+(?:\.[0-9]+)?)\s*(px|em|rem|%)

The first capture contains the numeric value, while the second captures the unit. The decimal portion is grouped without capturing because it is only needed to structure the number.

InputValueUnit
16px16px
1.5rem1.5rem
80%80%

Why Non-Capturing Groups Improve Maintainability

Consider a large regex containing many structural parentheses. If every group captures, the match result can contain many values that the application never needs.

^(?:https?://)?(?:www\.)?([A-Za-z0-9.-]+)$

Only the hostname needs to be extracted, so the protocol and www prefix are non-capturing groups. This makes the capture result smaller and makes the intent of the pattern clearer.

Common Mistakes with Regex Groups

  • Assuming every pair of parentheses should be a capturing group.
  • Using numbered groups in a large pattern and then losing track of their positions.
  • Forgetting that nested capturing groups each receive their own number.
  • Changing earlier groups and unintentionally changing later backreference numbers.
  • Using a repeated capturing group when every iteration needs to be extracted.
  • Confusing grouping with lookahead or lookbehind.
  • Forgetting that replacement syntax differs between regex tools and programming languages.
  • Using backreferences when simply repeating the same pattern would be sufficient.
  • Relying on named-group syntax without checking whether the target regex engine supports it.
  • Using regex groups to parse complex nested formats that should be handled by a parser.

How to Choose Between Capturing and Non-Capturing

  • Use (pattern) when the matched text needs to be extracted or referenced.
  • Use (?:pattern) when parentheses are required only for grouping.
  • Use named groups when several captured values have meaningful names.
  • Use numbered groups for short, stable expressions where the capture positions are obvious.
  • Avoid unnecessary captures in large expressions.
  • Prefer non-capturing groups for structural alternatives that are not part of the extracted result.

Debugging Capturing Groups

When a regex produces unexpected capture results, inspect the groups individually instead of looking only at the complete match.

  • Check where each opening capturing parenthesis occurs.
  • Count the groups in order if numbered captures are used.
  • Look for nested groups that may have shifted the numbering.
  • Check whether an optional group participated in the match.
  • Check whether a repeated group captured only its final iteration.
  • Replace structural captures with non-capturing groups when their values are unnecessary.
  • Use named groups when numeric positions are becoming difficult to track.
💡 A regex tester that displays each capture group separately is especially useful for learning and debugging. Test one group at a time before combining several captures into a large expression.

Regex Groups in JavaScript

JavaScript supports numbered and named capturing groups, non-capturing groups, and backreferences.

const pattern = /^(\w+)-(\d+)$/;
const match = pattern.exec("item-42");

console.log(match[0]);
// "item-42"

console.log(match[1]);
// "item"

console.log(match[2]);
// "42"

The complete match is stored separately from the captured groups. Numbered captures begin at index 1 in the JavaScript match result.

const pattern = /^(?<name>\w+)-(?<id>\d+)$/;
const match = pattern.exec("item-42");

console.log(match.groups.name);
// "item"

console.log(match.groups.id);
// "42"

Groups in Global Matching

When a regex is used to find multiple matches, the programming language's matching API determines how captures are returned. In JavaScript, methods such as matchAll() are useful when you need both multiple matches and their capture groups.

const pattern = /(\w+):(\d+)/g;
const input = "alpha:10 beta:20";

for (const match of input.matchAll(pattern)) {
  console.log(match[1], match[2]);
}

This approach provides each complete match together with its capture groups, making it convenient for extracting repeated structured values from a string.

Groups in Python

Python's re module supports ordinary capturing groups, non-capturing groups, named groups, and backreferences.

import re

pattern = re.compile(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})")
match = pattern.search("2026-09-03")

print(match.group("year"))
# 2026

Python uses a different named-group syntax from JavaScript. This is a good example of why regex syntax should always be checked against the target engine rather than copied between languages without verification.

Groups and Regex Engine Differences

Basic capturing groups are widely supported, but advanced group features can differ significantly between regex engines. Named groups, branch-reset groups, atomic groups, conditional groups, and replacement syntax are not universally identical.

If a regex is intended to run in several environments, keep the expression as portable as practical or explicitly document the target engine and supported features.

Groups and Performance

Capturing groups themselves are not automatically a performance problem. However, large patterns with many captures can create unnecessary match-result data, and complicated grouping combined with ambiguous repetition can contribute to expensive backtracking.

Using non-capturing groups where extraction is unnecessary can make the intention of the regex clearer and may reduce unnecessary capture bookkeeping. More importantly, the overall structure of the regex should avoid unnecessary ambiguity.

⚠️ Do not assume that replacing every capturing group with a non-capturing group will solve regex performance problems. Catastrophic backtracking usually comes from the overall matching structure, especially ambiguous nested repetition and alternatives.

When Regex Groups Are the Right Tool

Groups are a strong choice when the input has a relatively simple and predictable structure. Dates, identifiers, simple log records, version strings, and controlled text formats are common examples.

They become less suitable when the input contains nested structures, escaping rules, recursive syntax, or a full grammar. JSON, HTML, programming languages, and other complex formats are generally better handled by dedicated parsers.

Practical Grouping Workflow

  • Write the basic matching pattern first.
  • Identify which parts need to be extracted.
  • Wrap only those parts in capturing groups.
  • Use non-capturing groups for structural sections that do not need to be returned.
  • Add named groups when the pattern contains several meaningful captures.
  • Add backreferences only when the exact captured text must appear again.
  • Test optional and repeated groups separately.
  • Check the target regex engine's syntax before using advanced group features.
^(?<name>[A-Za-z]+)\s+\((?<id>\d+)\)$

This example captures a name and numeric identifier while treating the surrounding whitespace and parentheses as structural syntax.

Frequently Asked Questions

What are groups in regex?

Groups are expressions enclosed in parentheses. They let you treat multiple regex tokens as one unit and, with ordinary parentheses, capture the text matched by that section.

What is a capturing group?

A capturing group is a parenthesized regex expression whose matched text is stored separately from the complete match. The application can usually access that captured value after matching.

What is a non-capturing group?

A non-capturing group uses (?:pattern). It groups an expression for quantifiers or alternatives without adding its matched text to the normal numbered capture results.

What is the difference between a group and a character class?

A character class such as [abc] defines which individual characters can match one position. A group such as (abc) combines a larger expression and can capture the complete matched substring.

What is a regex backreference?

A backreference tells the regex engine to match the exact text previously captured by a group. For example, (\w+)\s+\1 looks for the same word twice.

Are named groups better than numbered groups?

Named groups can be easier to understand and maintain when a pattern has several captures. Numbered groups are shorter and widely supported. The appropriate choice depends on the pattern and target regex engine.

What happens when a capturing group is repeated?

A repeated capturing group normally does not create a separate capture for every repetition in the ordinary match result. The capture generally represents the value from the final successful iteration according to the engine's rules.

When should I use a non-capturing group?

Use a non-capturing group when you need parentheses for structure, grouping, alternation, or quantifier scope but do not need the matched text as a capture.

Conclusion

Regex groups provide a way to organize expressions, control quantifier scope, extract matched text, and build more powerful matching patterns. Ordinary parentheses create capturing groups, while (?:...) provides grouping without capturing.

Named groups can make complex expressions easier to understand, and backreferences allow a regex to require the same text that was captured earlier. Groups are also extremely useful in find-and-replace operations because captured values can often be rearranged in the replacement.

The most maintainable regexes capture only the information that the application actually needs. Use non-capturing groups for structural parentheses, prefer named groups when a pattern contains many meaningful captures, and always test advanced group syntax against the regex engine where the expression will run.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.