Comparison layers
Nextflow already records everything about a run in .nextflow/history and the per-session LevelDB cache under .nextflow/cache/. That data is rich, but it’s not built for eyeballing two runs side by side. nf-diff reads that same data — reusing Nextflow’s own internal cache and history APIs — and turns it into a readable comparison. Each question below maps to a layer of the report.
- Were the runs launched differently? A structured, flag-by-flag diff of each run’s launch command — Nextflow options (
-profile,-r) and pipeline params (--genome,--input) side by side — instead of eyeballing two opaque command strings. Params passed via-params-file(JSON/YAML) are parsed and merged in too, tagged by source (CLI,file, orCLI+file) so you can see where each value came from. - Did the configuration change? A diff of the resolved
nextflow.config— flattened to dotted keys likeprocess.cpus,executor.name,docker.enabled— with each run’s-profile/-coptions applied. This is what catches the classic case where two identical commands still behave differently because-profile dockerand-profile testresolve to different process resources, executors, or container settings. - Did the topology change? Which processes gained or lost tasks between the two runs.
- Did the wiring change? With
--diff-dag, the process→process edges each run actually ran are reconstructed and diffed, so a rewired pipeline (A → CbecomingA → B → C) shows up as added/removed edges — something the per-process task counts cannot reveal. When a Nextflow data-lineage store (.lineage/,lineage.enabled = true) recorded the run, edges are read from that authoritative provenance — the input LID references each task recorded — so no work directories are needed. Without a lineage store, edges are inferred from the input symlinks each task staged into its work directory (a link resolving into another task’s work dir is a producer→consumer edge); that fallback needs the work directories to still exist locally and is best-effort (a note reports which source was used and how many task work dirs were missing). Either way this layer is informational only — it never affects the “identical” verdict or--fail-on-change. - Did the tools change? A per-process software & versions diff of the container image(s) and Conda package spec(s) each process ran with — read straight from the run cache, so no work directories are needed. This is the layer that catches
biocontainers/fastqc:0.11.9→biocontainers/fastqc:0.12.1(or a bumpedbioconda::salmon=pin) directly, rather than leaving it buried in per-task detail. A software change counts toward the “identical” verdict and--fail-on-change. - Did a task change? Per-task diffs of status, exit code, container, script, requested resources, and measured usage.
- What failed, and why? A top-level failure rollup that gathers every failed task — detected from the cached
status/exitfields (an explicitFAILED/ABORTEDstatus or a non-zero exit code), so no work directories are needed — and rolls them up by(process, status, exit)signature, counted per run and sorted by biggest blast radius first. A signature seen only in Run B is flagged new (a regression), one in Run A but gone in Run B is resolved, and one in both is persistent, so you can tell at a glance whether a re-run introduced, fixed, or carried over a failure — e.g. aCALLprocess that started exiting137(OOM-killed) only in Run B. A run-level error state is surfaced too, even when no individual task failure was recorded. This is a summary of the per-task status/exit already shown in the task layer, so it is informational only and never separately affects the “identical” verdict or--fail-on-change. - Did anything get slower or heavier? A dedicated performance-regressions layer flags matched tasks whose runtime (
realtime) or peak memory (peak_rss,peak_vmem) changed by at least--perf-thresholdpercent (default25), sorted worst-regression first. Tasks that share the same cache hash are marked “same work”, so a+200%realtime on identical work stands out from a slowdown that also changed what ran. Improvements (Run B faster/leaner) are shown too, but only true regressions are counted. - Was anything over- or under-provisioned? A resource-efficiency layer that, for each process in each run, compares what it requested (
cpus,memory) against what it actually peaked at (%cpu,peak_rss) — both read straight from the run cache, so no work directories are needed and it is always computed. Each process is classified per run: over-provisioned (used under 50% of the reservation — e.g. “requested 32 GB, peaked at 4 GB” — so the allocation, and often the cost, was wasted), tight (used 90%+ of it, a risk of OOM kills or CPU throttling), or right-sized in between. Because provisioning is a tuning signal rather than a correctness change, this layer is informational only — it never affects the “identical” verdict or--fail-on-change. - Was work reused? Cached-task counts and total task realtime, plus a recompute count — matched tasks (present in both runs) whose cache hash differs, i.e. work that was re-executed rather than resumed — so you can see whether a re-run actually recomputed anything, and why.
- Did the results change? With
--diff-outputs, the files each matched task wrote to its work directory are compared — by size first, then a streamed SHA-256 for same-size files — and classified as added / removed / changed / unchanged. For a changed file that is text on both sides, it also produces a line-level diff (reusing the sameLineDiffengine as--diff-logs), bounded to the first--outputs-max-lineslines (default1000) and a hard byte cap; binary files (a NUL byte in the head) show only the size/hash change. Tasks that resumed from cache share a work directory, so they short-circuit to “identical”; the interesting cases are recomputed tasks whose outputs actually differ. This layer needs the work directories to still exist locally, and (unlike the always-changing performance metrics) an output change does count toward the “identical” verdict and--fail-on-change. - Did the published results change? With
--published-a=<dir>and--published-b=<dir>, the two runs’ durable published result trees (theiroutdir/publishDirdirectories) are compared directly. Files are keyed by their path relative to each published root and classified as added / removed / changed / unchanged, using the same size → SHA-256 → line-diff engine (FileContentComparator) as--diff-outputs, and honouring--outputs-max-bytes/--outputs-max-lines. Symlinks are followed, so it works whetherpublishDircopied or symlinked. Unlike--diff-outputs, this reads the durable published results instead of the work directories, so it still works after the work dirs are gone (cleaned up, or on remote object storage). A published-file change does count toward the “identical” verdict and--fail-on-change. - Why did a task fail? With
--diff-logs, the standard log files (.command.out,.command.err,.command.log) each matched task wrote are compared line by line, with the task’s exit code and status surfaced alongside — so a task that flipped from exit 0 to exit 1 shows both the change and the stderr that explains it. Reads are bounded (tailed to--logs-max-lines, default200, and a hard byte cap) so a huge log never blows up memory, and cache-resumed tasks sharing a work directory short-circuit as identical. Because stdout/stderr legitimately varies between runs (timestamps, paths, ordering), this layer is informational only — it never affects the “identical” verdict or--fail-on-change(the exit-code change already does).
Meaningful changes vs. everything
Section titled “Meaningful changes vs. everything”By default the report highlights meaningful changes and treats fields that always differ between two distinct runs (run name, session id, launch time, work directory, wall-clock time, measured resource usage, and each task’s cache hash) as context rather than “changes”. The task cache hash is treated this way because Nextflow folds the per-run session UUID into every task’s hash, so two independent runs always compute a different hash for every task even when the script, inputs and container are byte-identical — that difference is surfaced separately as the recompute count rather than flipping the verdict. The same rule applies to the parameters layer: launch options that routinely differ without changing what was executed (-name, -resume, -ansi-log, -with-tower, -with-weblog, -bg) are shown for context but not flagged as changes. Use --verbose when you want everything flagged.