Intro
Quick, what does this match?
That's the official regexp from semver.org. It
validates version numbers like:
Don't get me wrong, I love regexps, but in practice you probably spend
a bunch of time writing one, testing it against some cases, and moving
on, proud of your achievement!
Some time passes and lucky future you (or unlucky someone else) has to
change it. Dramatic pause here.
I bet you've been there. Now your options are probably: decode it
again from the start, rewrite the whole thing, or, in the age of AI,
ask (and hopefully not blindly accept) an LLM for a new recipe.
Emacs has had a nice answer for more readable regexps for a long time:
the rx macro. I started using it all the time in Emacs Lisp, as
reviewers always suggested it to me. Later, I started missing this DSL
in JavaScript and TypeScript, so I wrote a small version of it for my
projects.
So, what about reading that SemVer regexp like semver in the code
below?
The same strings match, and you get named groups as a bonus. By the
end of this post you'll know every piece of it.
TL;DR: jump straight to the cheat sheet,
the side-by-side examples, the full
source, or grab the
gist
to sneak a peek at the result.
NOTE: the RX here has nothing to do with
RxJS, which is an amazing library for reactive
programming with observables.
A taste of rx in Emacs Lisp
With rx you describe a regexp as a tree of named forms, and Emacs
turns it into the regexp string for you:
A few things to notice:
- Strings are literals.
"(" means a parenthesis. You don't
need to escape anything by hand.
- Sequence is implicit. Every form takes a list of things and
matches them one after the other. You don't need to wrap them in a
seq, even though seq exists.
- Groups appear only when needed.
(+ digit) becomes
[[:digit:]]+, not \(?:[[:digit:]]\)+.
The proposed JavaScript/TypeScript version in this post reads like
this:
Under the hood
If you want strings to be literals, you can't represent a regexp piece
as a plain string, otherwise you can't tell "(" (a literal
parenthesis) apart from "(?:...)" (a group you built). So every
piece is a small object:
src is the regexp text. kind records how that text behaves when
you glue it to other things:
atom: a single unit, like a, \d, [a-z] or (...). You can
put a quantifier right after it.
seq: safe to concatenate, but a quantifier needs (?:...) around
it. abc is a seq, and so is a+, since a+? would silently
turn into a lazy quantifier.
alt: has a | at the top level, so it needs (?:...) almost
everywhere.
Plain strings go through literal, which escapes them:
With that in place, seq joins nodes and only brackets alternations:
(That backreference check is one of those bugs you only find by
writing tests, or when it happens to you in prod. backref(1)
followed by the literal "0" gives you backreference number ten.)
Every quantifier is a seq of its arguments plus a suffix, bracketed
only when the body isn't an atom:
Because each quantifier calls seq on its arguments, you get the
implicit sequence for free: optional("-", group(x)) becomes
(?:-(x))?.
And finally, the two entry points. As in Emacs, rx returns a
string. RX returns a RegExp you can use right away:
RX.flags exists because Emacs controls case folding through the
case-fold-search variable, and JavaScript puts it on the regexp
itself.
That's the whole engine! Now, let's build our vocabulary.
Character sets
In Emacs you write (any "a-z" "_"). Inside those strings, a-z is a
range, and a - at either end is a plain dash. I kept the same rule:
The dash comes out escaped because sets can merge. If you combine
anyOf("+-") with anyOf("0-9"), an unescaped - would end up in
the middle and create a range from + to 0. Escaping it costs one
backslash.
And merging is the reason why RxNode has a set field. It is there
to hold the text that goes between [ and ], so anyOf can take
other sets as arguments:
not negates a set, and it knows the shorthand classes:
The simple email check, which most of us have written as
/^[^\s@]+@[^\s@]+\.[^\s@]+$/ at some point, becomes:
The rest of the Emacs character classes are there too: digit,
hexDigit, space, blank, wordChar, notWordChar, alpha,
alnum, lower, upper, punct, control, graphic, printing,
ascii and nonascii. One difference: in Emacs they understand
Unicode, and mine are ASCII only. alpha won't match é.
Two more come from rx's symbol list, and people (me, many times) mix
them up:
In rx, anything really means anything, newlines included. Here is
where the difference shows up:
Alternatives, and the longest match
or works as you'd expect, and gets bracketed when it lands inside a
sequence:
Did you notice the order changed? I copied that behavior from
Emacs. When every branch of an or is a plain string, rx hands them
to regexp-opt, which builds a pattern that prefers the longest
match:
JavaScript alternation takes the first branch that matches, going left
to right. So the naive regexp for a list of keywords has a 'bug':
I don't build a trie like regexp-opt does. Sorting the strings by
length, longest first, is enough to get the same behavior:
As in Emacs, or() with no branches returns unmatchable, which is
(?!) here. It's handy when you build the branch list at runtime and
it might come out empty.
Repetition, greedy and lazy
Emacs has (= n ...), (>= n ...) and (** n m ...). Here they are
repeat, atLeast and between:
The lazy versions *?, +? and ?? are zeroOrMoreLazy,
oneOrMoreLazy and optionalLazy. The classic HTML tag example:
Groups and backreferences
group is a capturing group, and backref points back to it:
Emacs also has (group-n N ...) to pick the group number. JavaScript
can't do that, but it has named groups, which serve the same purpose
and read better:
backref accepts a name as well:
Anchors
rx distinguishes the start of the string (bos) from the start of a
line (bol). In JavaScript both are ^, and the m flag decides
which one you get. I kept both names so the intent shows in the code:
Why not add start and end automatically? Because you only want
them when validating a whole string. When searching inside a text, as
in split, replace or matchAll, a hidden ^ and $ would break
everything. Emacs agrees: bos and eos are explicit in rx too.
wordBoundary and notWordBoundary map straight to \b and \B.
Emacs also has bow and eow (\< and \>), start and end of a
word. JavaScript lacks those, so I combined \b with a lookaround:
Literal and raw
Plain strings are already literals, but rx has an explicit literal
form for strings computed at runtime, and I kept it. It documents that
the value came from somewhere else:
The opposite direction is rx's (regexp ...) form, the escape hatch.
Here it's raw, and it receives either a string or an existing
RegExp. It lets you adopt the DSL in a codebase full of old regexps
without rewriting all of them, like:
raw can't see inside the text it gets, so it adds brackets whenever
it's combined with something else. It's an extra (?:), and the
regexp still works.
Shall we try SemVer again?
Back to the regexp from the intro. In Emacs you would give names to
the pieces with rx-define or rx-let. In TypeScript those are just
consts:
Now you can read the spec in the code. A numeric identifier is 0, or
a non-zero digit followed by any number of digits. A pre-release is a
dotted list of identifiers, and so is build metadata. dotted is a
plain function returning a node, which is as far as abstraction needs
to go here.
It matches the same strings as the official regexp, and the named
groups give you a result like:
Next time the spec changes, you can understand what the current regex
does at a glance, instead of fighting an army of punctuation.
What's missing from Emacs rx
I tried to map every rx form, and a few have no JavaScript equivalent:
point: JavaScript regexps don't know about a cursor.
symbol-start, symbol-end, syntax, category: these depend on
Emacs syntax tables.
intersection: possible with the v flag, but I haven't needed it.
minimal-match / maximal-match: these flip the greediness of
everything inside them. Doable, but it would need a separate pass,
and the *Lazy functions cover my use cases.
eval: TypeScript already evaluates expressions everywhere, so you
get it for free.
And one addition Emacs doesn't need: RX.flags.
Cheat sheet
| Emacs rx |
TypeScript |
JS regexp (roughly) |
seq, :, and |
seq(...), implicit in every form |
ab |
or, | |
or(...) |
a|b |
any, in, char |
anyOf("a-z", "_", digit) |
[a-z_\d] |
not-char |
notChar(...) |
[^...] |
not |
not(charset) |
\D, [^...] |
*, +, ? |
zeroOrMore, oneOrMore, optional |
x*, x+, x? |
*?, +?, ?? |
zeroOrMoreLazy, oneOrMoreLazy, optionalLazy |
x*?, x+?, x?? |
=, >=, ** |
repeat, atLeast, between |
x{n}, x{n,}, x{n,m} |
group |
group(...) |
(...) |
group-n |
named("name", ...) |
(?<name>...) |
backref |
backref(1), backref("name") |
\1, \k<name> |
literal |
literal(s) |
s, escaped: 1\+1 |
regexp, regex |
raw("..."), raw(/.../) |
(?:...), as-is |
rx-define, rx-let |
const |
(none) |
bos, eos |
start, end |
^, $ |
bol, eol |
lineStart, lineEnd (with the m flag) |
^, $ |
bow, eow |
wordStart, wordEnd |
\b(?=\w), \b(?<=\w) |
word-boundary |
wordBoundary |
\b |
not-word-boundary |
notWordBoundary |
\B |
nonl, not-newline |
notNewline |
. |
anychar, anything |
anything |
[\s\S] |
unmatchable |
unmatchable |
(?!) |
digit |
digit |
\d |
hex-digit, xdigit |
hexDigit |
[0-9a-fA-F] |
space, whitespace |
space |
\s |
blank |
blank |
[ \t] |
word, wordchar |
wordChar |
\w |
not-wordchar |
notWordChar |
\W |
alpha, letter |
alpha |
[a-zA-Z] |
alnum |
alnum |
[a-zA-Z0-9] |
lower, upper |
lower, upper |
[a-z], [A-Z] |
punct, punctuation |
punct |
[!-/:-@[-`{-~] |
cntrl, control |
control |
[\x00-\x1f\x7f] |
graph, graphic |
graphic |
[!-~] |
print, printing |
printing |
[ -~] |
ascii, nonascii |
ascii, nonascii |
[\x00-\x7f], [\u0080-\uffff] |
Examples: JS/TS regex vs RX
Each example below shows the goal, the Emacs rx form in a comment,
the regexp you would write by hand, and the RX version. When RX
produces a different regexp text, the // => line shows it. The
results at the bottom come from running both against the same strings.
Digits only
The whole string is digits.
Letters only
The whole string is ASCII letters.
Optional letter
Both spellings, color and colour.
Two words
Two words separated by a space.
Phone number
(123) 456-7890, parentheses and all.
Hex color
#ff00aa-style colors.
Signed integer
An optional sign, then digits.
Simple email
Something@something.something, no spaces.
No digits
A string without any digit.
CSV line
Exactly three comma-separated fields.
One of many
A fixed list of words.
Title and name
mr or ms, then a name, keeping the title.
Between
Two to four digits.
At least
Three or more digits, anywhere.
Repeated word
The same word twice.
Matching tags
An open tag and its own closing tag.
Whole word
cat as a word, not inside another one.
Case-insensitive
hello, in any case.
Full source
It's a single file with no dependencies. Copy it into your project
and start deleting the forms you don't need, or adding the ones you
miss.
You can check the same code, plus all the examples from this post (and
a few more), in this
gist.
If you'd rather not set anything up, paste it into the TypeScript
Playground, hit "Run", and check
the "Logs" tab.
Wrapping up
None of this is new. On the Emacs side, as I said before, rx has
shipped for decades, and the Elisp version is more complete than mine.
The idea of describing patterns with a small DSL instead of raw syntax
isn't new either. Plenty of people have tried it, each in their own
way. One project I like a lot in this space is
Zod, which I wrote about in my Zod quick
tutorial. It's not a regexp builder: you compose
small schema pieces, and Zod gives you back a parser and a TypeScript
type from the same construction. It follows the same spirit, though:
build big things out of small named pieces you can read.
If you write Elisp and have never tried rx, open *scratch*, type
(rx (+ digit)), and C-x C-e it. If you write JavaScript or
TypeScript, the file above is yours. And if you port it to another
language, send me a link.