What this regex matches
The canonical slug pattern is short and easy to read:
^[a-z0-9]+(?:-[a-z0-9]+)*$
Broken down piece by piece:
^and$anchor the match to the whole string. Without them, the pattern would happily match a valid slug buried inside a longer, invalid string.[a-z0-9]+matches one or more lowercase letters or digits. This is the "first word" of the slug and it must exist, so the slug cannot start with a hyphen.(?:-[a-z0-9]+)*is a non-capturing group repeated zero or more times. Each repetition is a single hyphen followed by one or more letters or digits. Because the hyphen is always followed by a letter or digit, you can never have two hyphens in a row, a trailing hyphen, or an empty segment.
The result is exactly the shape you want for a URL segment: my-first-post, product-123, a-b-c. Anything with a leading hyphen, trailing hyphen, double hyphen, uppercase, space, or symbol is rejected.
Why slugs exist in the first place
A slug is the human-readable part of a URL that identifies a piece of content: /blog/my-first-post or /shop/product-123. Slugs show up in blog posts, product pages, category pages, documentation, and anywhere a stable, readable address matters for SEO.
The constraints on a slug do not come from a style guide. They come from how URLs actually behave:
- URLs are case-sensitive in the path. On most servers
/Postand/postare different pages. Lowercasing removes an entire class of "why is my link 404ing" bugs. - Readability. A slug is meant to be read, typed, and shared. Spaces and punctuation make that harder and get percent-encoded into garbage like
%20and%2C. - Avoid percent-encoding. Every character outside the safe set becomes
%XX. A clean slug keeps the URL short and trustworthy-looking.
That safe set is, essentially, lowercase letters, digits, and hyphens. Hence the pattern.
The job is to generate, not validate
Here is the uncomfortable truth: a slug regex is useful for rejecting bad input, but it cannot create a good slug. If you only validate, you force the user to invent my-first-post by hand, and they will invent My First Post!, post (2), or café-au-lait instead.
The correct design is to generate the slug from a source string (usually a title) and show it to the user for confirmation. Validation then becomes a last-line guard, not the primary mechanism. This is why every CMS, every static site generator, and every e-commerce platform ships a "slugify" function rather than a "slug validation" dialog.
A slug generation algorithm
A robust slugifier does the following, in order:
- Lowercase the entire string.
- Normalize Unicode (NFKD) so accented characters split into base letter plus combining mark, then strip the marks:
cafébecomescafe. - Remove every character that is not a letter or digit in your target script.
- Convert whitespace (including runs of it) into a single hyphen.
- Merge adjacent hyphens into one.
- Trim leading and trailing hyphens.
A minimal Python version:
import re, unicodedata
def slugify(text):
text = text.lower()
text = unicodedata.normalize('NFKD', text)
text = ''.join(c for c in text if not unicodedata.combining(c))
text = re.sub(r'[^a-z0-9]+', '-', text)
return text.strip('-')
This produces my-first-post from "My First Post!", and cafe-au-lait from "Café au lait".
Handling non-Latin scripts
The algorithm above assumes a Latin alphabet. For Chinese, Japanese, Arabic, or Cyrillic titles you have two honest options, and you should pick deliberately:
- Transliterate to Latin before slugifying (pinyin for Chinese, romanization for others). This keeps URLs short and globally readable, but loses the original wording.
- Keep the original script and percent-encode it. Modern browsers and servers handle UTF-8 in paths fine, but the URL will contain long
%E6%...sequences. That is valid, just ugly, and some老旧 proxies mishandle it.
Many sites do both: a transliterated slug for the URL plus the original title shown on the page. Do not silently pick one and surprise your users.
Length limits and duplicates
A slug regex says nothing about length. Two practical rules:
- Cap the length. 50-80 characters is plenty. Long slugs get truncated in search results and look spammy. Truncate on a word boundary when possible.
- Detect duplicates. When two posts both generate
my-post, append a counter:my-post-2,my-post-3. This is a database uniqueness check, not something a regex can do.
What this regex cannot tell you
Even a perfect slug regex validates only shape, never meaning:
- It cannot know if
product-123is already taken. - It cannot tell a reserved word (
admin,api,www) from a normal slug. - It accepts
aaaaaaas happily asuseful-title. - It says nothing about the rest of the URL;
/blog//my-postis a routing problem, not a slug problem.
Use the regex as a cheap filter at the edge, and do the real work (uniqueness, routing, reserved words) in application code.
Common mistakes
| Wrong | Why it fails | Right |
|---|---|---|
^[a-z0-9-]+$ | Allows leading, trailing, and -- double hyphens. | ^[a-z0-9]+(?:-[a-z0-9]+)*$ |
^[\w-]+$ | \w includes uppercase and underscore, and still allows bad hyphen placement. | Explicit [a-z0-9] classes. |
| Validating, never generating | User enters garbage; you reject and they are stuck. | Generate from title, let user edit. |
| Forgetting Unicode | café survives as a combining-mark mess. | NFKD normalize, strip marks. |
Variants and stricter versions
Depending on your stack you may want to tighten the pattern:
- Disallow fully numeric slugs so
123is rejected (useful when slugs and numeric IDs could collide): add a negative lookahead^(?![0-9]+$)[a-z0-9]+(?:-[a-z0-9]+)*$. - Require at least two segments for hierarchical slugs; build it from the router, not the slug itself.
- Allow a dot for file-like slugs only if your routing genuinely needs it, then document it.
The base pattern stays the same. Reach for stricter rules only when a concrete requirement forces you to.