The argument "YAML or JSON?" is mostly the wrong question. They describe the same data shapes - maps, lists, scalars - but they were built for two different readers: JSON for programs, YAML for the humans who have to edit the file at 2am. Once you stop treating them as rivals and start treating them as two views of the same information, most of the hand-wringing disappears. This article puts them side by side so you can see exactly where each one pulls its weight.
Readability: who is the file for?
Start with the obvious difference. Here is one tiny service definition written both ways.
{
"service": "checkout",
"replicas": 3,
"ports": [8080, 9090],
"database": {
"host": "db.internal",
"tls": true
}
}
service: checkout
replicas: 3
ports:
- 8080
- 9090
database:
host: db.internal
tls: true
Both say the same thing. The JSON version is honest about its structure: every brace, every comma, every quote is a delimiter you can see. The YAML version drops all of that and leans on indentation. For a human scanning it, the YAML is faster to read. For a program, neither is easier - the parser does the work either way.
Comments: the feature JSON simply does not have
This is the single biggest reason people reach for YAML in configuration. JSON has no comment syntax at all. If you want to record why a timeout is 30 seconds instead of 5, you either abuse a field like "_comment" or you put it in a separate wiki that nobody reads six months later.
YAML lets you write:
timeout: 30 # raised after the latency incident on 2026-04-11
retries: 2
That line is worth more than it looks. Comments turn a config file from a snapshot into a conversation with your future self.
Multi-line strings: where YAML is quietly excellent
JSON has no clean way to embed a paragraph. You either escape every newline by hand or let it all collapse onto one line. YAML gives you two operators:
|(literal block) keeps every line break exactly as written.>(folded block) joins the lines with spaces, which is what you want for prose.
script: |
#!/bin/sh
echo "starting"
./run.sh
description: >
A long human-readable explanation that
wraps naturally across several lines
without you fighting the format.
Try doing that readably in JSON and you will see the appeal immediately.
Anchors: YAML's killer feature, and its sharpest edge
YAML can define a block once and reuse it, which removes an enormous amount of copy-paste in real configs:
defaults: &defaults
memory: 512Mi
cpu: 250m
worker:
<<: *defaults
replicas: 4
api:
<<: *defaults
replicas: 2
The &defaults anchor and the *defaults alias mean "paste that map here". The << merge key pulls the fields in. JSON has no equivalent - you repeat yourself or you generate the file.
Implicit vs explicit types: the trap that bites newcomers
JSON is explicit about types. If you want a string, you quote it. If you want a number, you don't. There is no guessing.
YAML guesses, and it guesses from the bare text:
country: NO # boolean false, not the string "NO"
version: 1.20 # number 1.2, not the string "1.20"
flag: yes # boolean true
ratio: 1:20 # sexagesimal, parsed as 80
Country codes, version strings, ratios - all of these get silently reinterpreted. The only safe rule is: quote anything you want treated as a literal string. JSON never makes you think about this.
Parsing speed: JSON is in another league
Raw parsing performance is not close. JSON's grammar is tiny - a few production rules - so a parser can be a tight state machine. YAML's grammar is a large specification with anchors, tags, flow and block styles, custom types, and a resolver that has to sniff types on every scalar. In practice a well-tuned JSON parser is often one to two orders of magnitude faster than a YAML parser on the same data.
For a config file you read once at startup that barely matters. For a hot path that parses millions of documents, it absolutely does.
The security surface: YAML can execute things
This is the part that surprises people. A YAML document is not just data - it can carry type tags that tell the loader to construct objects. The classic dangerous example:
!!python/object/apply:os.system
- "rm -rf /"
A naive loader that honours tags (Python's yaml.load without restrictions) will happily construct and run that. The fix is to use yaml.safe_load, which refuses anything but the basic types. JSON has no such footgun - there is no mechanism in the format to instantiate objects, so a JSON parser is safe by construction.
If you ever parse YAML from an untrusted source, safe_load is not optional. It is the difference between "I read a config" and "I ran someone's code."
Converting between them: what you lose
You can round-trip data from one format to the other, but the transformation is lossy in one direction:
- JSON to YAML keeps all the data and adds readability, but it also inherits YAML's type-guessing. A JSON string
"1.20"can become a YAML number if you are not careful. - YAML to JSON always drops comments and anchors. The merged result stays, but the reusable definitions and the explanations are gone for good.
Side-by-side summary
| Dimension | JSON | YAML |
|---|---|---|
| Readability for humans | Verbose | Clean, indentation-based |
| Comments | None | Native # |
| Multi-line strings | Awkward | | and > blocks |
| Reuse | None | Anchors and aliases |
| Type handling | Explicit | Implicit, can surprise |
| Parse speed | Very fast | 10-100x slower |
| Security | Safe by design | Tags can run code; use safe_load |
| Niche | API / data exchange | Human-edited config |
So which one?
Both. Keep JSON on the wire and on disk where programs talk to programs, and reach for YAML where a person has to read, reason about, and edit the file. A common and sensible pattern is to author in YAML, validate it, then compile to JSON before it reaches anything that needs to be fast or safe.
When you need to move between them, the JSON to YAML converter handles either direction. Just remember to re-check numbers and booleans after a YAML-to-JSON conversion - version strings are the usual casualty.