Skip to content
d.devtul.fun
中文
Config · 2026-08-16

When to Use YAML (and When You Shouldn't)

YAML is everywhere in infrastructure: Kubernetes, CI pipelines, Docker Compose, GitHub Actions, Ansible. It is also the format developers love to hate. The honest answer to "when should I use YAML" is not "always" or "never" - it is "when the file is written and read by humans who value comments and structure, and you do not need the file to compute." This article gives you a decision frame, concrete comparisons, and an exit path.

What YAML is good at

YAML was designed to be pleasant for people. It has comments, it uses indentation instead of braces, and it maps cleanly onto nested data. For a config that a human edits by hand - a deployment manifest, a pipeline definition - those properties matter. A teammate can read a YAML file and understand it without a parser in their head.

database:
  host: db.prod.internal
  port: 5432
  pool:
    max: 20
    idle: 5

Compare that to the same intent in JSON and the difference is mostly braces and quotes; the win is readability and the # comment you lose in JSON. YAML also supports anchors and aliases (& and *) so you can define a block once and reference it elsewhere, which is a genuine convenience for repeated structure - as long as you remember the reference is a copy, not a pointer you can later diverge.

The same config in four formats

To see the trade, here is one small configuration expressed in each format. The data is identical; the ergonomics differ.

// JSON - no comments, strict, parser-friendly
{
  "database": {
    "host": "db.prod.internal",
    "port": 5432,
    "pool": { "max": 20, "idle": 5 }
  }
}

# TOML - flat, comment-friendly, great for app config
[database]
host = "db.prod.internal"
port = 5432
[database.pool]
max = 20
idle = 5

# CUE - schema AND data, validates as you write
database: {
  host: string | *"db.prod.internal"
  port: int & >= 1 & <= 65535 | 5432
  pool: { max: int | 20, idle: int | 5 }
}

JSON wins when the consumer is a program and you never hand-edit it, and its strictness is a feature: a missing comma fails loudly instead of being silently coerced. TOML wins for application settings files (Cargo, pyproject) because it stays flat and readable. CUE wins when you want the config to be checked against a schema at authoring time, catching a bad port before it ever reaches a server. YAML sits in the middle: human-friendly, but with footguns.

YAML's footguns

The indentation model hides real risks. The most famous is the Norwegian town NO parsed as boolean false, and unquoted 012 sometimes read as octal. Strings that look like timestamps (2026-09-14) get coerced to dates by some parsers. The fix is to quote everything that is not a number, but that defeats the readability point.

# These surprises are real across parsers
active: no        # becomes boolean false
value: 012        # may become octal 10
date: 2026-09-14  # may become a Date object

There are quieter traps too. A multiline string uses | or >, and a stray extra space changes block vs folded semantics. Mixing tabs and spaces for indentation fails in most parsers, because YAML explicitly forbids tabs. And anchors are copy-by-reference: edit one aliased block and every alias moves with it, which surprises people who expected independent copies. Defense: quote string values, pin a YAML version and parser, and validate the parsed result against a schema rather than trusting inferred types.

Should a config language be Turing complete?

This is the deeper argument behind templating. One side says config should be data only - pure, declarative, diffable. The other says you need computation: loops over environments, derived values, conditionals. The pro side points at DRY: without loops you repeat the same block for dev, staging, prod, and the copies drift. The con side points at the disasters: a config that calls a function, hits a network, or behaves differently per run is no longer reviewable. A reviewer cannot tell what a Turing-complete manifest will produce without executing it, which means CI cannot show a clean diff of the final file.

The pragmatic middle is a two-layer model: keep the base config as plain data, and apply computation in a separate, bounded step (templating) that emits plain data before it reaches the system. The emitted YAML stays reviewable in the artifact store; the computation stays out of the hot path and can be unit-tested on its own. You get DRY where you need it without turning every config into a program.

The pain of large manifests

At scale, YAML stops being friendly. A single Kubernetes deployment plus its services, configmaps, and ingress can run to hundreds of lines; a real application with several components reaches thousands. Reviewing a 4,000-line diff is brutal, copy-paste drifts between environments, and a single bad indent takes down a rollout. This is where teams reach for templating - not because YAML is wrong, but because hand-maintaining thousands of near-identical lines is.

Templating: Helm vs Kustomize

ToolModelStrengthWeakness
HelmGo templates over YAMLReusable charts, rich logicTemplating can produce invalid YAML; logic hides in templates
KustomizePatch-based overlaysPlain YAML, no new languageOverlays get complex; limited computation
CUE/JsonnetCode that emits dataTypes and validationLearning curve; another tool to run

Helm gives you full computation but pushes you toward the Turing-complete trap, and a Helm template that emits a syntax error only fails at apply time. Kustomize keeps you closer to plain YAML with overlays per environment, which stays reviewable and fails fast on bad patches. For most teams the question is not "which is best" but "how much logic do we actually need" - start with Kustomize, reach for Helm only when you are shipping a chart to others, and reach for CUE when validation matters more than templating.

When JSON or TOML is the better call

Pick JSON when the config is generated and consumed by programs (build outputs, API payloads) and you want strict, unambiguous parsing. Pick TOML for application settings that humans edit but stay small and flat. Pick YAML when humans edit nested, comment-rich definitions and you accept the validation overhead. If you find yourself writing loops or conditionals in YAML, that is the signal you have outgrown it and should move the logic to a generator.

YAML versus the alternatives at a glance

NeedYAMLJSONTOMLCUE
CommentsYesNoYesYes
Human editingGoodPoorGoodGood
Schema validationExternalExternalExternalBuilt in
ComputationNoneNoneNoneYes (bounded)

A decision table

If you need...Use
Human-edited nested config with commentsYAML
Machine-generated, parsed-only configJSON
Flat app settings, simple typesTOML
Config validated against a schemaCUE or Jsonnet
Many environments from one sourceKustomize (YAML) or Helm

Migrating away from YAML

If YAML is costing you more than it helps, move in steps. First, add a schema (CUE or native Kubernetes validation) so the current files are at least checked, which also gives you a contract to migrate against. Second, generate the YAML from a typed source for the parts that hurt most - your CI can render Jsonnet or CUE into YAML that the cluster consumes unchanged, so the runtime does not change and you can diff the generated output. Tools like yq and kpt help you refactor and pipeline YAML without hand-editing thousands of lines. Third, migrate one component at a time; do not rewrite the whole repo in a weekend. Measure whether the new format actually reduces diff size and review time before committing the rest of the team. The goal is less YAML where it hurts, not zero YAML everywhere.

Summary

Use YAML for human-edited, nested, commented configuration, and accept that you must quote strings and validate the result. Reach for JSON when programs own the file, TOML for flat settings, and CUE when you want schemas. When manifests grow past a few hundred lines, add Kustomize or Helm rather than hand-copying, and keep the emitted YAML reviewable. If the pain keeps rising, migrate the worst components to a typed, code-generated source one at a time.

Keep reading