Skip to content
d.devtul.fun
中文
DevOps · 2026-08-20

YAML for Kubernetes: Writing Manifests That Don't Break at 3am

Kubernetes manifests are just YAML. That is both the good news and the bad news: the format is simple, but every field carries meaning, and a single mismatch can leave you staring at a pod that refuses to start at 3am. This is a field-by-field walk through a minimal, working setup - a Deployment and the Service in front of it - plus the mistakes that actually happen in production.

A minimal Deployment

apiVersion: apps/v1
kind: Deployment
metadata:
  name: checkout
  namespace: shop
  labels:
    app: checkout
spec:
  replicas: 3
  selector:
    matchLabels:
      app: checkout
  template:
    metadata:
      labels:
        app: checkout
    spec:
      containers:
        - name: checkout
          image: registry.internal/shop/checkout:1.4.2
          ports:
            - containerPort: 8080
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              cpu: 500m
              memory: 256Mi
          readinessProbe:
            httpGet:
              path: /healthz
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 10
          livenessProbe:
            httpGet:
              path: /live
              port: 8080
            periodSeconds: 15

That is a complete, deployable object. Now the parts people get wrong.

apiVersion is not decoration

apiVersion tells the API server which schema to validate against, and it is tied to kind. A Deployment lives under apps/v1; a bare v1 is for core objects like ConfigMap and Service. Write apps/v1 for a Deployment and v1 for a Service, or the server rejects the object before it ever looks at the rest.

labels and selector must match - the classic "nothing happens" bug

The selector.matchLabels block says which pods this Deployment owns. The pod template's labels must contain every key-value pair in the selector. If they drift apart, the Deployment creates pods it does not manage, or, more commonly, you edit the template labels and kubectl apply silently refuses because the selector is immutable.

The wrong way:

selector:
  matchLabels:
    app: checkout
template:
  metadata:
    labels:
      app: checkout
      tier: web     # harmless-looking, but now selector no longer matches

The right way is to keep the template labels a strict superset of the selector labels. When this breaks, the symptom is usually "I changed the config and nothing happened" - because the controller rejected the apply and nobody read the error.

requests and limits are not optional in production

The resources block looks optional and the cluster will let you omit it. Don't. Without requests, the scheduler has no idea how much CPU or memory to reserve, so pods pile onto one node until it tips over. Without limits, one runaway process can starve everything else on the node.

requests are what the scheduler guarantees; limits are the ceiling. Set them per container, and set memory limits especially - memory is not compressible, so an over-limit container gets OOM-killed rather than throttled.

liveness, readiness, and startup probes are three different things

Mixing these up is how you get services that serve errors while "running."

  • livenessProbe answers "is this process alive?" If it fails, Kubernetes restarts the container. Point it at something that only fails when the process is truly wedged.
  • readinessProbe answers "can this pod take traffic?" If it fails, the pod stays in the Service's endpoints but receives no requests. Use it for "database connection not ready yet."
  • startupProbe answers "has the app finished booting?" It protects slow starters from being killed by the liveness probe before they are up.

The common mistake: pointing liveness at the same endpoint as readiness. If the database is briefly down, readiness correctly says "not ready," but liveness also fails and Kubernetes hard-restarts a container that was actually fine - turning a brief blip into a full outage.

ConfigMap and Secret: two ways to mount

Both can be injected as environment variables or as files in a volume. Files are usually safer because they don't require a restart to change in many setups and avoid leaking secrets into printenv output.

env:
  - name: LOG_LEVEL
    valueFrom:
      configMapKeyRef:
        name: checkout-config
        key: log_level
volumeMounts:
  - name: secrets
    mountPath: /etc/secrets
    readOnly: true

Secrets are base64-encoded, not encrypted at rest by default - treat them like plaintext and restrict who can kubectl get secret.

kubectl apply vs create vs replace

  • kubectl create -f creates from scratch and fails if the object exists. Good for one-shot setup, bad for iteration.
  • kubectl apply -f is declarative: it computes a diff against the live object and patches only what changed. This is what you want for day-to-day config.
  • kubectl replace -f overwrites the entire object with what's in the file. If someone changed something live and it's not in your file, that change is gone.

Test before you ship: --dry-run and diff

Two commands save more 3am pages than any dashboard:

kubectl apply -f deploy.yaml --dry-run=client    # validates locally, no cluster change
kubectl diff -f deploy.yaml                      # shows exactly what will change

--dry-run=client checks the manifest parses and is well-formed before you touch the cluster. kubectl diff shows the delta against what is live, so you see the blast radius before you press enter.

Don't copy-paste for multiple environments

The moment you maintain dev.yaml, staging.yaml, prod.yaml by hand, they drift. Use Kustomize overlays or Helm values to keep one base and layer differences:

apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
  - ../base
patches:
  - target:
      kind: Deployment
      name: checkout
    patch: |-
      - op: replace
        path: /spec/replicas
        value: 5

The namespace trap

If you omit namespace from metadata, the object lands in whatever your current context's default namespace is - usually default, which has no NetworkPolicy, no resource quotas, and no naming discipline. Always set it explicitly, or better, put it in the Kustomization context so you can't forget.

A Service to match

apiVersion: v1
kind: Service
metadata:
  name: checkout
  namespace: shop
spec:
  selector:
    app: checkout
  ports:
    - port: 80
      targetPort: 8080

Note the Service selector must match the pod labels (not the Deployment selector) - same rule, different object. Get that wrong and the Service has zero endpoints and every request times out.

If you want to draft and sanity-check manifests without a cluster handy, the JSON to YAML converter lets you flip a file back and forth, which at least catches indentation and syntax problems before they reach kubectl.

Keep reading