A hash function is a machine that takes anything you feed it — a password, a 4 GB video file, a single byte — and spits out a fixed-length string of characters. Same input always produces the same output. Different input almost always produces a different output. That output is called a digest, a fingerprint, or just a hash. Nothing about it is reversible, and that one-way property is the whole point.
What a hash function actually does
You do not call a hash function to "scramble" data the way you encrypt it. You call it to summarise data into something small and stable. The mechanics stay the same regardless of input size:
input: "The quick brown fox" -> digest: 37ab... (64 hex chars for SHA-256)
input: a 2-hour movie -> digest: 9f2c... (still 64 hex chars)
The digest never grows with the input. That fixed size is what makes hashes cheap to store and compare.
Four properties you can rely on
- Deterministic. Hash the same bytes twice and you get byte-identical output. This is why a checksum can ever be trusted.
- Fixed-length output. SHA-256 always returns 256 bits no matter what you give it.
- One-way. Given the digest, you cannot recover the input. There is no "unhash" button, and anyone who claims otherwise is selling something.
- Avalanche effect. Change a single bit of the input and roughly half the output bits flip. The two digests look completely unrelated.
You can watch the avalanche effect with two almost identical strings:
$ echo -n "hello" | sha256sum
2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
$ echo -n "hellp" | sha256sum
1dd9f635fc9d209cbfa643af332d5b9f6d4b66cb3c8c1756a65d9d19c8a6833c
One letter changed in a five-letter word, and the entire 64-character output is different. That is the avalanche effect doing its job.
Hash vs encryption vs encoding
These three get confused constantly, usually because all three turn readable data into something that looks like noise. They are solving three different problems.
| Property | Hash | Encryption | Encoding |
|---|---|---|---|
| Reversible? | No (one-way) | Yes, with the key | Yes, no key |
| Needs a key? | No | Yes | No |
| Output size | Fixed | Roughly input size | Roughly input size |
| What it proves | Integrity / identity | Confidentiality | Transport safety |
| Typical use | Passwords, checksums, signatures | Storing secrets, TLS | Base64 in JSON, URLs |
The fastest way to tell them apart: encoding is a format change with no secrecy (anyone can decode it), encryption hides data behind a key, and hashing throws the original away on purpose.
The algorithm lineup: who is still standing
Not every hash algorithm is safe to use, and the gap has widened a lot over the last two decades.
| Algorithm | Digest size | Status in 2026 |
|---|---|---|
| MD5 | 128 bits | Broken for collisions. Checksums only, never for security. |
| SHA-1 | 160 bits | Broken for collisions (SHAttered, 2017). Retire it. |
| SHA-256 | 256 bits | Solid. The default for checksums and signatures. |
| SHA-512 | 512 bits | Solid. Faster than SHA-256 on 64-bit hardware. |
| BLAKE3 | configurable | Modern, extremely fast, parallel. Great for new code. |
MD5 and SHA-1 are the two you must stop using for anything security-sensitive. They are still fine for detecting an accidental download corruption, because an attacker is not actively trying to forge a colliding file there. But the moment someone has an incentive to craft a second file with the same digest, both algorithms fall over.
Collisions and the length-extension trap
A collision is when two different inputs produce the same digest. With a 256-bit output there are far more possible inputs than possible digests, so collisions mathematically must exist — the question is whether an attacker can find one on purpose. For MD5 and SHA-1, they can, in seconds on a laptop.
A separate gotcha is the length-extension attack. SHA-256 and SHA-1 are built on the Merkle–Damgård construction, which means that if you know H(message) but not the message itself, you can often append data and compute H(message || extra) without ever seeing the original message. This is exactly why you must not build a "MAC" by just doing hash(secret || data). Use HMAC instead (covered in the next article).
Salt and pepper: two different things
These two words get used interchangeably by people who should know better. They are not the same.
- Salt is a random value generated per user, stored next to the hash. It stops attackers from using precomputed rainbow tables and stops two users with the same password from getting the same hash. It is not secret — it sits in the database in plain sight.
- Pepper is a single secret value shared across your whole application, kept outside the database (an environment variable, a KMS). It protects you in the scenario where the database leaks but the pepper does not. It does nothing if the attacker has both.
If you are storing passwords, you want a proper password hash (bcrypt, scrypt, Argon2id) that already salts internally — you do not hand-roll salt and pepper yourself. For message authentication with HMAC, the key plays the pepper role.
Which one should you use
| Scenario | Reach for | Why |
|---|---|---|
| Verify a downloaded file | SHA-256 | Standard, collision-resistant checksum |
| Store user passwords | Argon2id or bcrypt | Slow + salted on purpose |
| Authenticate a message | HMAC-SHA256 | Keyed, immune to length extension |
| Just make binary fit in text | Base64 (encoding!) | Not a hash at all |
| Hide data from others | AES-GCM (encryption!) | A hash cannot do this |
The single most common mistake is grabbing MD5 "because it is fast" and using it for something that needs to resist an attacker. Speed is a virtue for checksums and a liability for passwords. Pick the algorithm based on the adversary, not the benchmark.
The mistake: using a hash where you needed encryption
A surprisingly common bug is "encrypting" something by hashing it, only to find later that the business actually needs the original back. A payments team hashes card numbers to "protect" them, then cannot bill the customer because the number is gone forever. Hashing proves a thing is unchanged; it does not hide it reversibly. When you need to recover the data, hashing is the wrong tool and encryption is the right one:
# wrong: you can never get the number back
stored = sha256(card_number)
# right: encrypt if you must recover it later
from cryptography.fernet import Fernet
key = Fernet.generate_key()
token = Fernet(key).encrypt(card_number.encode())
This is the single cleanest way to tell hash from encryption in practice: if you need the original, a hash cannot help you, and pretending otherwise just loses data.
Want to compute a few digests yourself and see the avalanche effect firsthand? The hash generator will hash any text with SHA-256, SHA-512 and BLAKE3 side by side.