You have seen them everywhere: 550e8400-e29b-41d4-a716-446655440000. UUIDs are the default identifier in databases, message queues, and distributed systems. But "just use a UUID" hides real choices — there are several versions, and the version you pick changes everything from privacy to database performance. This article lays it all out.
What a UUID is
A UUID (Universally Unique Identifier) is a 128-bit value. Written in its canonical form it is 36 characters: 32 hex digits in five groups separated by hyphens.
xxxxxxxx-xxxx-Mxxx-Nxxx-xxxxxxxxxxxx
8 - 4 - 4 - 4 - 12 = 36 chars
Two bits are reserved to encode the version (M) and the variant (N). The rest is filled by whichever method the version specifies. The 128-bit width is what makes "universally unique" a reasonable claim rather than marketing.
The versions
| Version | How it is generated | Traits |
|---|---|---|
| v1 | Timestamp + MAC address | Sortable by time; leaks the machine MAC |
| v3 | MD5 namespace hash | Deterministic for a given name |
| v4 | Random (122 bits) | No order; the common default |
| v5 | SHA-1 namespace hash | Deterministic, stronger than v3 |
| v7 | Unix time (ms) + random | Time-ordered, DB-friendly |
v1 concatenates a timestamp with the host's MAC address. That makes collisions effectively impossible and the values sortable, but it also broadcasts a hardware identifier — a privacy problem in some contexts. v3 and v5 are namespace hashes: feed the same name and namespace and you get the same UUID every time, which is handy for idempotent lookups. v4 is pure randomness and the version most libraries return by default. v7 puts a millisecond timestamp at the front so values sort chronologically while keeping randomness in the tail.
Why distributed systems use UUIDs
With a single database, an auto-incrementing integer primary key is simple and compact. The moment you have more than one writer, it breaks down:
- Coordination overhead. Two nodes inserting at once will fight over the next integer unless you add a central sequencer — which is a single point of failure and a bottleneck.
- Merge conflicts. If you ever sync two databases, their
id = 42rows collide and you cannot tell them apart. - Information leakage. Sequential ids reveal how many users you have and let an attacker enumerate records at
/users/43,/users/44.
A UUID generated on any node is unique without coordination, merges cleanly, and gives away nothing about your volume. That is why distributed logs, event stores, and multi-region databases reach for them.
The collision probability, intuitively
v4 uses 122 random bits, so there are 2^122 possible values — about 5 × 10^36. The birthday math says you need to generate roughly 2.6 × 10^18 v4 UUIDs before a collision becomes likely. To put that in perspective: generating a billion UUIDs per second for a hundred years still leaves you enormously far from that number. For any normal application, v4 collisions are not a real risk. If you are worrying about v4 collisions, you have bigger architectural problems than UUIDs.
The database performance cost
Here is the catch that surprises teams moving from integers to UUIDs: a random UUID is a terrible primary key for a B+ tree index.
An auto-increment key always inserts at the "right edge" of the index — the newest row is physically next to the last one, so inserts are cheap and the working set stays in cache. A v4 UUID is random, so every insert lands somewhere different in the key space. The B+ tree pages are scattered, the cache locality is destroyed, and on a large table you get index fragmentation and more disk I/O.
The fix is to make the identifier time-ordered. UUID v7 (or ULID, which is a similar 128-bit time-prefix design) puts the timestamp first, so inserts cluster at the recent edge just like an auto-increment key — but without giving up decentralization. If you must use UUIDs as keys, prefer v7 or ULID over v4.
When NOT to use a UUID
- As a primary key on a huge, write-heavy table without time-ordering — you will pay in index fragmentation. Use v7/ULID or a composite key.
- Where compactness matters. A UUID is 16 bytes (36 as text). An integer is 4 or 8. For a table with billions of rows and many secondary indexes, that size adds up.
- When a human must type or read it. Nobody should ever transcribe a UUID by hand; use a shorter slug or code for anything user-facing.
Generating one
Most languages ship a generator. In Node:
import { randomUUID } from 'crypto';
randomUUID(); // v4
// for v7 you typically use a small library, e.g. uuidv7()
For quick experiments and bulk generation in the browser, the UUID generator on this site produces v4 and v7 values locally, with no round trip to a server.