Character Classes in Regex
A character class matches one character out of a set. This is the workhorse of regex — get it right and half your patterns get simpler. Here are the shorthands, custom sets, and the rules that change inside brackets.
Shorthand classes
| Class | Matches | Equivalent |
|---|---|---|
\d | Any digit | [0-9] (ASCII) |
\w | Word character | [A-Za-z0-9_] |
\s | Whitespace | space, tab, newline… |
\D \W \S | The negations | anything but the above |
. | Any char except newline | — |
Custom sets with brackets
| Pattern | Matches |
|---|---|
[aeiou] | Any single vowel |
[a-z] | Any lowercase letter (range) |
[A-Za-z0-9] | Any alphanumeric |
[^0-9] | Any character that is NOT a digit |
[a-fA-F0-9] | A hex digit |
Rules change inside brackets
.,*,+,?,(lose their special meaning inside[ ]—[.+]matches a literal dot or plus.^only negates when it is the first character:[^a]is "not a", but[a^]matchesaor^.- To include
-literally, put it first or last:[-a-z]or[a-z-]. - To include
], put it first:[]a], or escape it:[\]a].
POSIX and Unicode classes
Some engines support named sets like [[:alpha:]] or Unicode property escapes such as \p{L} (any letter) and \p{N} (any number). In JavaScript, \p{...} requires the u flag.
Example: a hex color
#[a-fA-F0-9]{6}\b
A # followed by exactly six hex digits — one compact character class does the work.
FAQ
Does \w match accented letters like é?
By default no — \w is [A-Za-z0-9_]. To match Unicode letters use \p{L} with the unicode flag, or an explicit range that includes the characters you need.
How do I match a literal hyphen inside brackets?
Place it at the start or end of the class — [-a-z] or [a-z-] — or escape it as [a\-z]. In the middle it would be read as a range.
Is [0-9] the same as \d?
For plain ASCII text, yes. Under Unicode mode \d may also match digits from other scripts, while [0-9] always means just those ten characters.