YAML is everywhere in infrastructure: Kubernetes, CI pipelines, Docker Compose, GitHub Actions, Ansible. It is also the format developers love to hate. The honest answer to "when should I use YAML" is not "always" or "never" - it is "when the file is written and read by humans who value comments and structure, and you do not need the file to compute." This article gives you a decision frame, concrete comparisons, and an exit path.
What YAML is good at
YAML was designed to be pleasant for people. It has comments, it uses indentation instead of braces, and it maps cleanly onto nested data. For a config that a human edits by hand - a deployment manifest, a pipeline definition - those properties matter. A teammate can read a YAML file and understand it without a parser in their head.
database:
host: db.prod.internal
port: 5432
pool:
max: 20
idle: 5
Compare that to the same intent in JSON and the difference is mostly braces and quotes; the win is readability and the # comment you lose in JSON. YAML also supports anchors and aliases (& and *) so you can define a block once and reference it elsewhere, which is a genuine convenience for repeated structure - as long as you remember the reference is a copy, not a pointer you can later diverge.
The same config in four formats
To see the trade, here is one small configuration expressed in each format. The data is identical; the ergonomics differ.
// JSON - no comments, strict, parser-friendly
{
"database": {
"host": "db.prod.internal",
"port": 5432,
"pool": { "max": 20, "idle": 5 }
}
}
# TOML - flat, comment-friendly, great for app config
[database]
host = "db.prod.internal"
port = 5432
[database.pool]
max = 20
idle = 5
# CUE - schema AND data, validates as you write
database: {
host: string | *"db.prod.internal"
port: int & >= 1 & <= 65535 | 5432
pool: { max: int | 20, idle: int | 5 }
}
JSON wins when the consumer is a program and you never hand-edit it, and its strictness is a feature: a missing comma fails loudly instead of being silently coerced. TOML wins for application settings files (Cargo, pyproject) because it stays flat and readable. CUE wins when you want the config to be checked against a schema at authoring time, catching a bad port before it ever reaches a server. YAML sits in the middle: human-friendly, but with footguns.
YAML's footguns
The indentation model hides real risks. The most famous is the Norwegian town NO parsed as boolean false, and unquoted 012 sometimes read as octal. Strings that look like timestamps (2026-09-14) get coerced to dates by some parsers. The fix is to quote everything that is not a number, but that defeats the readability point.
# These surprises are real across parsers
active: no # becomes boolean false
value: 012 # may become octal 10
date: 2026-09-14 # may become a Date object
There are quieter traps too. A multiline string uses | or >, and a stray extra space changes block vs folded semantics. Mixing tabs and spaces for indentation fails in most parsers, because YAML explicitly forbids tabs. And anchors are copy-by-reference: edit one aliased block and every alias moves with it, which surprises people who expected independent copies. Defense: quote string values, pin a YAML version and parser, and validate the parsed result against a schema rather than trusting inferred types.
Should a config language be Turing complete?
This is the deeper argument behind templating. One side says config should be data only - pure, declarative, diffable. The other says you need computation: loops over environments, derived values, conditionals. The pro side points at DRY: without loops you repeat the same block for dev, staging, prod, and the copies drift. The con side points at the disasters: a config that calls a function, hits a network, or behaves differently per run is no longer reviewable. A reviewer cannot tell what a Turing-complete manifest will produce without executing it, which means CI cannot show a clean diff of the final file.
The pragmatic middle is a two-layer model: keep the base config as plain data, and apply computation in a separate, bounded step (templating) that emits plain data before it reaches the system. The emitted YAML stays reviewable in the artifact store; the computation stays out of the hot path and can be unit-tested on its own. You get DRY where you need it without turning every config into a program.
The pain of large manifests
At scale, YAML stops being friendly. A single Kubernetes deployment plus its services, configmaps, and ingress can run to hundreds of lines; a real application with several components reaches thousands. Reviewing a 4,000-line diff is brutal, copy-paste drifts between environments, and a single bad indent takes down a rollout. This is where teams reach for templating - not because YAML is wrong, but because hand-maintaining thousands of near-identical lines is.
Templating: Helm vs Kustomize
| Tool | Model | Strength | Weakness |
|---|---|---|---|
| Helm | Go templates over YAML | Reusable charts, rich logic | Templating can produce invalid YAML; logic hides in templates |
| Kustomize | Patch-based overlays | Plain YAML, no new language | Overlays get complex; limited computation |
| CUE/Jsonnet | Code that emits data | Types and validation | Learning curve; another tool to run |
Helm gives you full computation but pushes you toward the Turing-complete trap, and a Helm template that emits a syntax error only fails at apply time. Kustomize keeps you closer to plain YAML with overlays per environment, which stays reviewable and fails fast on bad patches. For most teams the question is not "which is best" but "how much logic do we actually need" - start with Kustomize, reach for Helm only when you are shipping a chart to others, and reach for CUE when validation matters more than templating.
When JSON or TOML is the better call
Pick JSON when the config is generated and consumed by programs (build outputs, API payloads) and you want strict, unambiguous parsing. Pick TOML for application settings that humans edit but stay small and flat. Pick YAML when humans edit nested, comment-rich definitions and you accept the validation overhead. If you find yourself writing loops or conditionals in YAML, that is the signal you have outgrown it and should move the logic to a generator.
YAML versus the alternatives at a glance
| Need | YAML | JSON | TOML | CUE |
|---|---|---|---|---|
| Comments | Yes | No | Yes | Yes |
| Human editing | Good | Poor | Good | Good |
| Schema validation | External | External | External | Built in |
| Computation | None | None | None | Yes (bounded) |
A decision table
| If you need... | Use |
|---|---|
| Human-edited nested config with comments | YAML |
| Machine-generated, parsed-only config | JSON |
| Flat app settings, simple types | TOML |
| Config validated against a schema | CUE or Jsonnet |
| Many environments from one source | Kustomize (YAML) or Helm |
Migrating away from YAML
If YAML is costing you more than it helps, move in steps. First, add a schema (CUE or native Kubernetes validation) so the current files are at least checked, which also gives you a contract to migrate against. Second, generate the YAML from a typed source for the parts that hurt most - your CI can render Jsonnet or CUE into YAML that the cluster consumes unchanged, so the runtime does not change and you can diff the generated output. Tools like yq and kpt help you refactor and pipeline YAML without hand-editing thousands of lines. Third, migrate one component at a time; do not rewrite the whole repo in a weekend. Measure whether the new format actually reduces diff size and review time before committing the rest of the team. The goal is less YAML where it hurts, not zero YAML everywhere.
Summary
Use YAML for human-edited, nested, commented configuration, and accept that you must quote strings and validate the result. Reach for JSON when programs own the file, TOML for flat settings, and CUE when you want schemas. When manifests grow past a few hundred lines, add Kustomize or Helm rather than hand-copying, and keep the emitted YAML reviewable. If the pain keeps rising, migrate the worst components to a typed, code-generated source one at a time.