Skip to content
d.devtul.fun
中文
Workflow · 2026-08-24

How to Compare Large JSON Files Locally

Comparing two big JSON files is a different problem from comparing two short snippets. The moment a file hits a few megabytes, the obvious move — paste both into a web tool — stops working: the tab hangs, and you have just shipped potentially sensitive data to a third-party server. This guide is about doing it correctly, on your own machine, for free.

Why you should not use an online tool

Two reasons, and the second is the deal-breaker:

  • Privacy. JSON exports are where access tokens, customer records and internal schemas hide. Uploading them to a random site is how leaks start.
  • Scale. Most browser-based diff UIs choke on a few megabytes. They were built for a paste, not for a payload.

Everything below runs locally. The diff checker and the JSON viewer on this site are both client-side, so even the in-browser path keeps your data on the device.

Normalise before you compare

The biggest source of false differences is formatting, not content. Two files can describe identical data and still differ because of key order, indentation, or line endings. Text-diffing such files produces noise that buries the real change.

# sort keys and pretty-print both, then compare as text
jq -S . a.json > a.norm.json
jq -S . b.json > b.norm.json
diff a.norm.json b.norm.json

jq -S ("sort keys") re-emits the object with keys in a stable order. After that, a text diff reports only genuine content differences. This single step resolves the majority of "why is everything different" cases.

Text diff vs structural diff

Consider two files that are the same data with different key order:

// a.json
{"name": "x", "id": 1}
// b.json
{"id": 1, "name": "x"}

A text diff flags the whole line as changed, even though semantically nothing did. A structural comparison parses both into objects and compares values, reporting "equal". So the rule is: normalise, then text-diff for a quick answer; use a structured comparison when key order is unreliable.

Ignoring array order

Arrays are ordered in JSON, so [1,2,3] and [3,2,1] are textually and semantically different. Sometimes order is meaningless to you — a list of tags, say. Decide per case:

  • If order matters (a sequence of events), keep it and let the diff show the reordering.
  • If order is noise, sort the array during normalisation so the comparison ignores it.

A small Python script for semantic diff

For real structure-aware comparison — including aligning array elements by an id rather than by position — a short script beats any text tool. This one recurses and reports paths that differ:

import json

def load(p):
    with open(p, encoding="utf-8") as f:
        return json.load(f)

def diff(a, b, path=""):
    if type(a) != type(b):
        print(f"{path}: type {type(a).__name__} != {type(b).__name__}")
        return
    if isinstance(a, dict):
        for k in set(a) | set(b):
            diff(a.get(k), b.get(k), f"{path}.{k}")
    elif isinstance(a, list):
        if len(a) != len(b):
            print(f"{path}: length {len(a)} != {len(b)}")
        for i, (x, y) in enumerate(zip(a, b)):
            diff(x, y, f"{path}[{i}]")
    elif a != b:
        print(f"{path}: {a!r} != {b!r}")

diff(load("a.json"), load("b.json"))

To compare array elements by id instead of position, replace the list branch with a lookup keyed on each element's id field before recursing. That turns "element 4 moved" into a true match.

Memory-friendly streaming for huge files

When a file is too large to load whole, use a streaming parser. ijson yields items one at a time so memory stays flat:

import ijson

with open("huge.json", encoding="utf-8") as f:
    for item in ijson.items(f, "item"):
        process(item)   # compare item-by-item

For many top-level records, stream both files and compare records as you go rather than building two giant lists. The --stream style also appears in tools that process JSON line by line; the principle is the same — never hold more than one record in memory at a time.

Producing a readable report

A raw path-by-path dump is hard to read. Collect differences into a list and write a small summary:

report = []
def diff(a, b, path=""):
    ...
    else:
        report.append((path, a, b))

diff(load("a.json"), load("b.json"))
with open("report.txt", "w", encoding="utf-8") as out:
    for path, x, y in report:
        out.write(f"{path}: {x!r} -> {y!r}\n")
print(f"{len(report)} differences written to report.txt")

The -> above is an escaped arrow so it survives the code block; in your own script you can write a literal -> or just use a plain hyphen.

Performance traps

  • Diffing before normalising. Key-order and whitespace noise inflate the diff and hide the real edit. Always jq -S first.
  • Loading everything into memory. A 500 MB file as one Python object can exhaust RAM; stream it.
  • Pretty-printing gigantic output. Pretty JSON can be 2–3x larger; diff the compact sorted form, not a reformatted one.
  • Repeated parsing. Parse each file once and reuse the object; re-reading per comparison wastes time.

Putting it together

The reliable workflow is: sort keys with jq -S, text-diff for a fast verdict, and reach for a small structural script when order or arrays mislead the text view. For everyday files the diff checker handles it in the browser without uploading a byte; for truly huge exports, the streaming Python approach above keeps both your data and your memory safe.

Keep reading