Regex to Match a Roman Numeral
Copy the pattern, paste your text below, and find every valid Roman numeral. Below the tester: how the subtractive rules are encoded and why the anchors matter.
What this pattern matches
This expression validates a single Roman numeral from I up to MMMCMXCIX (3999). It enforces the subtractive rules, so valid forms like XIV and MCMLXXXIV pass while malformed ones like IIII or VX fail.
How it works, part by part
M{0,3} — zero to three thousands.
(CM|CD|D?C{0,3}) — hundreds: 900, 400, or 0–300 with an optional 500.
(XC|XL|L?X{0,3}) — tens, following the same subtractive pattern.
(IX|IV|V?I{0,3}) — units.
Why the anchors are essential
Without ^ and $ the pattern can match an empty string or a valid fragment inside garbage, because every group allows zero characters. Anchoring forces the whole input to be a well-formed numeral. To find numerals inside prose instead, replace the anchors with word boundaries and require at least one character.
Use it in your code
import re
text = "XIV"
pattern = r'''^M{0,3}(CM|CD|D?C{0,3})(XC|XL|L?X{0,3})(IX|IV|V?I{0,3})$'''
for m in re.finditer(pattern, text):
print(m.group())
const text = "XIV";
const re = new RegExp(String.raw`^M{0,3}(CM|CD|D?C{0,3})(XC|XL|L?X{0,3})(IX|IV|V?I{0,3})$`, "g");
console.log(text.match(re));
FAQ
Does it reject IIII?
Yes. The units group allows at most three I characters unless they form IV or IX, so IIII fails validation.
Can it match numbers above 3999?
No. Standard Roman numerals stop at MMMCMXCIX. Larger numbers use overlines that plain text cannot represent.
How do I find numerals inside a sentence?
Replace ^ and $ with \b word boundaries and require one or more characters, otherwise it matches empty spans.