UseToolSuite UseToolSuite
Text Processing 📖 Pillar Guide

Regular Expressions: The Complete Guide

The complete regex reference: syntax, quantifiers, anchors, groups and lookarounds, NFA backtracking and ReDoS, per-language gotchas in JS/Python/Java, plus tested patterns for email, URLs, IPs and dates.

Necmeddin Cunedioglu Necmeddin Cunedioglu 17 min read

Practice what you learn

Regex Tester

Try it free →

Regular expressions (often abbreviated as regex or regexp) represent one of the most powerful and universally adopted tools in a software developer’s arsenal. At first glance, a regex pattern looks like a chaotic string of random punctuation marks. However, once you learn the syntax, regex becomes an irreplaceable superpower that allows you to search, match, validate, and transform large blocks of text with surgical precision.

Whether you are validating user email inputs on a frontend form, parsing terabytes of Nginx access logs on a backend server, executing complex refactoring operations across a large codebase, or building web scrapers, regular expressions are the underlying technology making it possible.

This comprehensive guide will take you from absolute basics to advanced engine mechanics, covering syntax, quantifiers, capture groups, zero-width assertions (lookaheads/lookbehinds), and the dangerous phenomenon of catastrophic backtracking that can crash your servers.

Practice as you read: Open our interactive Regex Tester in another tab. Paste the patterns from this guide and test them against your own sample text to see how they behave in real-time.

What Is a Regular Expression?

At its core, a regular expression is a sequence of characters that defines a specific search pattern. Think of it as a highly advanced “Find and Replace” feature that operates on rules rather than exact string matches.

Pattern: \b\d{3}-\d{4}\b
Matches: "555-1234", "800-9999"
Does not match: "5551234", "A55-1234", "555-12345"

The pattern above tells the engine: “Find a word boundary, followed by exactly three digits, followed by a literal hyphen, followed by exactly four digits, ending with another word boundary.”

Basic Syntax and Literal Characters

The simplest regular expression is just a string of literal characters. The pattern hello matches the exact sequence of letters “h-e-l-l-o” in the target text.

However, the true power of regex comes from Metacharacters — special characters that have a dedicated meaning within the regex engine.

Essential Metacharacters

MetacharacterMeaningExample PatternWhat It MatchesWhat It Doesn’t Match
.Any single character (except a newline)h.that, hit, hot, h5tht, hoot
\dAny digit (0-9)\d\d42, 99, 004a, a9, 4
\wAny word character (a-z, A-Z, 0-9, and underscore _)\w\wa1, _B, ZZa!, a-, a
\sAny whitespace character (space, tab, newline)a\sba b, a\tba-b, ab
\bWord boundary (the position between a word and a non-word)\bcat\b”cat” (as an isolated word)“category”, “tomcat”
^Start of a string (or start of a line in multiline mode)^Error:”Error: File not found""System Error: crash”
$End of a string (or end of a line in multiline mode)failed$”Login failed""failed to login”

Character Classes (Sets)

If you want to match a specific set of characters, you use square brackets [] to define a Character Class.

SyntaxMeaningExample Matches
[aeiou]Matches any single vowel”a”, “e”, “i”
[a-z]Matches any lowercase letter from a to z”f”, “q”, “z”
[A-Z0-9]Matches any uppercase letter OR any digit”G”, “5”, “X”
[^0-9]Matches any character that is NOT a digit (Negation)“a”, ”!”, ” “

Anchors and Boundaries

Anchors are the counterpart to character classes: instead of matching a character, they match a position. Nothing is consumed, and nothing appears in the match result. This is what people mean by a zero-width assertion.

PatternMatches
^Start of the string (start of each line with the m flag)
$End of the string (end of each line with the m flag)
\bA word boundary — the transition between \w and \W
\BA non-word boundary — anywhere \b does not match

Anchors are what stop a pattern from matching in the wrong place, and forgetting them is the single most common reason a “working” regex produces surprise matches in production:

import re

text = "The cat scattered the catch."
# Match 'cat' only as a whole word
print(re.findall(r'\bcat\b', text))   # ['cat']
print(re.findall(r'cat', text))        # ['cat', 'cat', 'cat']

The unanchored version silently matches inside scattered and catch. The same mistake in a validation routine is what lets notanemail@@@x through a check you believed was strict.

Quantifiers: Controlling Repetition

Often, you don’t know exactly how many times a character will appear. Quantifiers allow you to specify the required number of repetitions.

QuantifierMeaningExample PatternMatches
*Zero or more timesab*cac, abc, abbc, abbbbbbc
+One or more timesab+cabc, abbc (Does NOT match ‘ac’)
?Zero or one time (Optional)colou?rcolor, colour
{n}Exactly n times\d{4}2026, 1999
{n,}n or more times\w{3,}cat, doggy, elephant
{n,m}Between n and m times\d{2,4}42, 123, 2026

Greedy vs. Lazy Quantifiers

By default, quantifiers like * and + are greedy. They will consume as much text as possible while still allowing the overall pattern to match.

Imagine you are parsing HTML and want to extract the text inside a div tag: Target string: <div>First</div> <div>Second</div>

Greedy Pattern: <div>.*</div>
Match Result:   <div>First</div> <div>Second</div> (It matched the entire string!)

To fix this, you make the quantifier lazy (or reluctant) by adding a question mark ? after it. This tells the engine to match as little text as possible.

Lazy Pattern:   <div>.*?</div>
Match 1:        <div>First</div>
Match 2:        <div>Second</div>

There is a third mode that JavaScript does not have. Possessive quantifiers (*+, ++, ?+, supported in Java and PCRE) behave like greedy ones but refuse to give characters back: once consumed, they are never released for backtracking. That makes them both a performance optimisation and a structural defence against the catastrophic backtracking covered later in this guide — which is precisely why their absence from JavaScript matters so much.

Capture Groups and Alternation

Alternation (OR logic)

The pipe character | acts as an OR operator. Pattern: cat|dog|bird will match “cat”, “dog”, or “bird”.

Capture Groups

Parentheses () create capture groups. They serve two purposes:

  1. They group multiple characters together so a quantifier can apply to the entire group.
  2. They “capture” the matched text, allowing you to extract it later in your code.
Pattern: (\d{4})-(\d{2})-(\d{2})
Target:  "Born on 2026-03-22"

Match:   2026-03-22
Group 1: 2026 (Year)
Group 2: 03   (Month)
Group 3: 22   (Day)

In programming languages like JavaScript or Python, these extracted groups can be accessed via array indices, making regex an incredibly powerful tool for parsing structured data logs.

Non-Capturing Groups

If you need to group characters for logic but don’t want the engine to store the result (which saves memory and improves performance), use a non-capturing group by adding ?: inside the parentheses.

Pattern: (?:http|https)://(www\.)?example\.com This groups the protocol for alternation, but only captures the “www.” portion.

Capturing costs something: the engine has to store the matched text and track indices for every group. Prefer (?:...) unless you actually need the substring back.

Named Groups and Back-References

Numeric group indices become unreadable the moment a pattern has more than two of them, and they shift whenever someone inserts a group in the middle. Named groups fix both problems.

PatternPurpose
(abc)Capture group, referenced as $1 / \1
(?:abc)Non-capturing group
(?<year>\d{4})Named capture group, referenced as \k<year> or groups.year
\1Back-reference — matches the same text group 1 captured

A back-reference is the one construct here that matches content rather than a pattern, which makes it the natural way to find doubled words:

"the the cat sat sat down".replace(/\b(\w+)\s+\1\b/g, "$1");
// "the cat sat down"

Note that back-references are also one of the two features (with lookarounds) that force an engine to backtrack, so a linear-time engine like RE2 cannot support them at all.

Advanced Techniques: Lookarounds

Lookarounds are zero-width assertions. They check if a pattern exists ahead or behind the current position, but they do not consume any characters and are not included in the final match result.

Lookaround TypeSyntaxMeaningExample Use Case
Positive Lookahead(?=...)”Must be followed by”\d+(?=px) matches “16” in “16px”, but ignores “16” in “16em”
Negative Lookahead(?!...)”Must NOT be followed by”foo(?!bar) matches “foo” in “foobaz”, but not in “foobar”
Positive Lookbehind(?<=...)”Must be preceded by”(?<=\$)\d+ matches “100” in “$100”, but ignores “100” in “€100”
Negative Lookbehind(?<!...)”Must NOT be preceded by”(?<!\w)cat matches “cat” but ignores “cat” inside “tomcat”

Note: While lookaheads are universally supported, lookbehinds have historically lacked support in some older browsers (like Safari pre-2020), though modern environments fully support them.

Regex Execution Flags (Modifiers)

Flags change the global behavior of the regex engine. They are usually appended to the end of the regex literal (e.g., /pattern/g).

FlagNameEffect
gGlobalFinds all matches in the string, rather than stopping after the first match.
iIgnore CaseCase-insensitive matching. /a/i matches both “a” and “A”.
mMultilineChanges the behavior of ^ and $. They now match the start/end of each line rather than the entire string.
sDotAllForces the . metacharacter to match newline characters (\n). By default, . stops at newlines.
uUnicodeEnables full Unicode property escapes and handles surrogate pairs correctly.

Using Regex in JavaScript, Python and Java

The pattern is only half the problem. Each runtime wraps it in an object model with its own trap, and all three traps below produce bugs that look like the regex is wrong when it is not.

JavaScript: lastIndex makes a global regex stateful

A RegExp with the g flag carries a mutable lastIndex across calls. Reusing one instance gives different answers to the same question:

const regex = /test/g;
regex.test("test 1, test 2"); // true  (lastIndex -> 4)
regex.test("test 1, test 2"); // true  (lastIndex -> 12)
regex.test("test 1, test 2"); // false (no match after 12, resets to 0)

Three identical calls, three different results. This bites hardest when a regex is hoisted to module scope and shared — a validator that passes in tests starts rejecting every second request in production. Prefer String.prototype.matchAll(), or construct the regex inside the function, or drop the g flag when you only need a boolean.

Python: compile once, and always use raw strings

Python’s re module caches compiled patterns, but in a hot loop it is idiomatic to compile explicitly. Always use raw strings r"" so Python does not interpret the backslashes before the regex engine ever sees them:

import re

pattern = re.compile(r'\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}\b', re.IGNORECASE)
valid = [e for e in emails if pattern.match(e)]

Without the r prefix, '\b' is a backspace character and the pattern silently stops meaning what you wrote.

Java: Pattern is thread-safe, Matcher is not

public class Validator {
    // Compile once, share freely — Pattern is immutable and thread-safe.
    private static final Pattern EMAIL = Pattern.compile("^[A-Za-z0-9+_.-]+@(.+)$");

    public static boolean isValid(String email) {
        // A new Matcher per call — sharing one across threads corrupts its state.
        return EMAIL.matcher(email).matches();
    }
}

The failure mode here is the nastiest of the three, because a shared Matcher only misbehaves under concurrency and will look perfect in a single-threaded test.

Deep Dive: Engine Mechanics and Catastrophic Backtracking

While Regular Expressions are incredibly powerful, poorly written patterns can inadvertently cause Denial of Service (DoS) vulnerabilities in web applications. This is known as a ReDoS (Regular Expression Denial of Service) attack. To prevent this, developers must understand how the underlying engine operates.

NFA vs. DFA Engines

There are two primary architectures for regex engines: Deterministic Finite Automata (DFA) and Nondeterministic Finite Automata (NFA).

  • DFA Engines (used in command-line tools like grep or awk) are rigid and extremely fast. They evaluate the input string exactly once, tracking all possible matches simultaneously in memory. The tradeoff is that they cannot support advanced features like capture groups or lookarounds.
  • NFA Engines (the standard for JavaScript, Python, PHP, Ruby, and Java) are feature-rich but operate using a “backtracking” approach.

The Mechanics of Catastrophic Backtracking

Catastrophic backtracking occurs exclusively in NFA engines. It happens when a regex contains nested quantifiers (e.g., *, +) and fails to find a match at the very end of a long input string.

Consider this notoriously vulnerable regex pattern: ^(a+)+$

This pattern commands the engine: “Find one or more ‘a’s, group them, repeat that group one or more times, and ensure it touches the end of the string.”

If you evaluate the string "aaaaaaaaaaaaaaaaaaaa!" (20 ‘a’s followed by an exclamation mark), here is what the NFA engine does:

  1. It greedily matches all 20 ‘a’s with the inner a+.
  2. It hits the ! and realizes the $ (end of string) condition fails.
  3. Because of its architecture, the engine assumes it grouped the letters wrong. It backtracks.
  4. It tries matching 19 ‘a’s in the first group, and 1 ‘a’ in the second group. It fails at the !.
  5. It tries 18 ‘a’s and two groups of 1. It fails.
  6. It tries 18 ‘a’s and one group of 2. It fails.

Because of the nested quantifiers (a+)+, the engine attempts to evaluate every single possible mathematical permutation of grouping those 20 letters before finally conceding that no match exists.

For 20 letters, this results in 2^20 (over 1 million) internal evaluations. Adding just 10 more ‘a’s to the input string pushes the evaluation count into the billions. This locks up the CPU thread entirely. In Node.js (which is single-threaded), a single user submitting this payload will crash the entire web server for all users.

Mitigation Strategies for ReDoS

To protect your applications from catastrophic backtracking, implement these rules:

  1. Never nest quantifiers: Avoid patterns like (a+)+, (a*)*, or (a|a+)*. Rewrite them as simplified, flat patterns (e.g., a+).
  2. Beware of overlapping alternations: (a|a)* is functionally identical to nested quantifiers and will cause the same backtracking explosion.
  3. Use anchors immediately: Anchoring your patterns with ^ and $ prevents the engine from attempting to start the match from every single character in a 10,000-word document.
  4. Enforce input length limits: Never run complex regex against an infinitely long user input. Trim or block inputs over a reasonable threshold (e.g., 255 characters for an email) before executing the regex test() or match() function.

What this actually looks like in production

Catastrophic backtracking is usually taught as a curiosity. It is not — it is a recurring cause of real, public, whole-site outages, and the pattern of those incidents is worth knowing because it repeats.

Cloudflare, 2 July 2019. A newly deployed WAF rule intended to catch inline JavaScript contained the fragment .*(?:.*=.*). Nested against certain request bodies it consumed CPU without bound. The rule was pushed globally, and the resulting CPU exhaustion took down HTTP serving across the network for roughly thirty minutes. The regex was not malicious and had passed review; the input that triggered it was ordinary traffic.

Stack Overflow, 20 July 2016. A regex used to trim trailing whitespace from post content — a task nobody would consider dangerous — met a post containing a very long run of spaces. The site was down for 34 minutes.

Two lessons generalise from both. First, the vulnerable pattern was written by competent engineers doing something mundane; “don’t write bad regexes” is not a control. Second, in each case a single request consumed a shared resource, which is what separates ReDoS from an ordinary slow query — on a single-threaded runtime like Node, one request holds the entire event loop.

The mitigations your runtime actually gives you

The advice above — flatten quantifiers, anchor, cap input length — is necessary and not sufficient, because it depends on reviewing every pattern correctly forever. The structural fix is to run untrusted patterns on an engine that cannot backtrack, or to bound them with a timeout. What is available depends entirely on your runtime, and the differences are stark:

RuntimeTimeoutAtomic groups / possessiveLinear-time engine
Gonot needednot supportedregexp is RE2 — immune by default
.NETRegex accepts a matchTimeoutyesno
Javanone built inyesno
Pythonre none; regex module has oneregex module onlyvia google-re2
JavaScript / Nodenonenoneonly via the re2 native module

JavaScript is the worst-placed of the five and the most likely to be handling untrusted input: no timeout, no atomic groups, no possessive quantifiers, and single-threaded. If you are running user-supplied or user-influenced patterns in Node, the practical answer is the re2 binding — RE2 guarantees linear time by refusing to backtrack at all. The price is that it drops backreferences and lookarounds, because those are precisely the features that require backtracking. If your pattern needs them, it cannot be made safe by engine choice, and input limits plus a hard length cap are the remaining defence.

Test any pattern you are unsure about against a deliberately hostile input — a long run of the repeated character followed by one character that cannot match — and watch whether it returns instantly or hangs. That test takes ten seconds and is the whole of the difference between the two outages above and an uneventful deploy.

Case Study: Validating an Email Address

Email is the validation problem every developer meets first and almost everyone over-engineers, so it is worth working through properly — the reasoning generalises to every other “validate this format” task.

The pattern you should almost always use

^[^\s@]+@[^\s@]+\.[^\s@]+$

Read literally: a run of characters that are neither whitespace nor @, then @, then a domain, a dot, and a TLD. That is all. It accepts user@example.com, john.doe+aws@sub.domain.co.uk and 名前@example.jp, and rejects user@example, user @example.com and @example.com.

The philosophy is deliberate ignorance: validate the shape without predicting which characters the IETF will permit next year. This is the same posture the W3C takes for HTML5’s <input type="email">.

Three anti-patterns still being copied off old forum posts

Hardcoding the TLD length. \.[a-z]{2,4}$ was roughly true before 2014. ICANN has since launched hundreds of long gTLDs, so this now rejects .museum, .engineer, .technology and .photography. Never constrain TLD length.

Forbidding +. A pattern whose local-part class omits + breaks plus-addressing (user+newsletter@gmail.com), which power users rely on for trackable aliases. They will not work around it; they will abandon the form.

The “complete RFC 5322” monster. There is an infamous ~6,300-character pattern that claims to implement the full grammar. It is unmaintainable, it drives the engine through nested lookaheads — a textbook ReDoS surface, per the section above — and it happily accepts "quoted string"@example.com and user@[192.168.1.1], which SendGrid, Mailgun and SES reject anyway. It is strictly worse than the four-line version on every axis you care about.

A related trap: the “strict” pattern ^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$ looks professional and is fine for a closed internal system, but it rejects internationalized addresses outright. Non-ASCII mailboxes have been valid since RFC 6531 was ratified in 2012 and are used by tens of millions of people; user@例え.jp fails that pattern for no good reason.

Length limits belong in code, not in the pattern

Regex engines are bad at counting across segments. SMTP’s limits are worth enforcing separately, before the pattern ever runs:

ComponentMaximumReference
Total address254 charactersRFC 5321 §4.5.3.1.3
Local part (before @)64 charactersRFC 5321 §4.5.3.1.1
Domain part (after @)253 charactersRFC 1035 §2.3.4

Rejecting oversized input up front is also your cheapest ReDoS control:

function validateEmail(input) {
  const email = input.trim().toLowerCase();

  // Cheap structural guards first — these also cap ReDoS exposure.
  if (email.length > 254) return { valid: false, reason: "Exceeds the RFC length limit." };

  const [localPart, domainPart] = email.split("@");
  if (!localPart || !domainPart) return { valid: false, reason: "Missing local part or domain." };
  if (localPart.length > 64)     return { valid: false, reason: "Local part exceeds 64 characters." };

  // Only now run the pattern.
  if (!/^[^\s@]+@[^\s@]+\.[^\s@]+$/.test(email)) {
    return { valid: false, reason: "Fails the structural check." };
  }
  return { valid: true, normalized: email };
}

On the front end you often need no JavaScript at all — <input type="email" required maxlength="254"> gives you a localised browser error for free. Treat it as the first line, never the only one.

The rule that makes all of the above secondary

Regex validates format; it cannot validate existence. perfectly.valid.format@this-domain-is-fake.com satisfies the W3C spec, the HTML5 spec and the most elaborate RFC 5322 pattern ever written.

The only real verification is closed-loop: accept the address with a permissive check, mail a time-limited token, and treat the address as confirmed when the user clicks it. Regex is a typo-catcher, not a security boundary — and once you internalise that, the temptation to chase a “perfect” pattern disappears.

Pattern Reference

Battle-tested starting points. Test each against your own edge cases in the Regex Tester before shipping it.

Validating a hex colour code

^#([A-Fa-f0-9]{6}|[A-Fa-f0-9]{3})$

Matches #FFFFFF, #abc, #123456. Rejects #F, #1234567, white.

Extracting a domain name from a URL

^(?:https?:\/\/)?(?:www\.)?([^\/]+)

Given https://www.example.com/page, captures example.com in group 1.

Validating a strong password

^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,}$

Stacked lookaheads require at least one lowercase, one uppercase, one digit and one symbol, with a minimum length of 8.

The rest, in one block

URL (HTTP/HTTPS):
^https?:\/\/(?:www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b(?:[-a-zA-Z0-9()@:%_\+.~#?&\/=]*)$

IPv4 address:
^(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$

Date (YYYY-MM-DD), strict:
^\d{4}-(?:0[1-9]|1[0-2])-(?:0[1-9]|[12]\d|3[01])$

Hex colour (3, 4, 6 or 8 digits):
^#?(?:[0-9a-fA-F]{3,4}|[0-9a-fA-F]{6}|[0-9a-fA-F]{8})$

HTML tag (extraction, non-greedy):
<\/?[\w\s="/.':;#-\/\?]+>

Interactive testing: paste any of these into our Regex Tester to see matches highlighted live, inspect the syntax tree, and run them against your own sample data.

Further Reading


Validate, debug, and test your patterns safely using our completely private, browser-based Regex Tester. It highlights syntax, visualizes capture groups, and provides execution time warnings.

Necmeddin Cunedioglu
Necmeddin Cunedioglu Author
• 17 min read •
-- views

Software developer and the creator of UseToolSuite. I write about the tools and techniques I use daily as a developer — practical guides based on real experience, not theory.