Skip to content
nf-diff

Comparison layers

Nextflow already records everything about a run in .nextflow/history and the per-session LevelDB cache under .nextflow/cache/. That data is rich, but it’s not built for eyeballing two runs side by side. nf-diff reads that same data — reusing Nextflow’s own internal cache and history APIs — and turns it into a readable comparison. Each question below maps to a layer of the report.

  • Were the runs launched differently? A structured, flag-by-flag diff of each run’s launch command — Nextflow options (-profile, -r) and pipeline params (--genome, --input) side by side — instead of eyeballing two opaque command strings. Params passed via -params-file (JSON/YAML) are parsed and merged in too, tagged by source (CLI, file, or CLI+file) so you can see where each value came from.
  • Did the configuration change? A diff of the resolved nextflow.config — flattened to dotted keys like process.cpus, executor.name, docker.enabled — with each run’s -profile/-c options applied. This is what catches the classic case where two identical commands still behave differently because -profile docker and -profile test resolve to different process resources, executors, or container settings.
  • Did the topology change? Which processes gained or lost tasks between the two runs.
  • Did the wiring change? With --diff-dag, the process→process edges each run actually ran are reconstructed and diffed, so a rewired pipeline (A → C becoming A → B → C) shows up as added/removed edges — something the per-process task counts cannot reveal. When a Nextflow data-lineage store (.lineage/, lineage.enabled = true) recorded the run, edges are read from that authoritative provenance — the input LID references each task recorded — so no work directories are needed. Without a lineage store, edges are inferred from the input symlinks each task staged into its work directory (a link resolving into another task’s work dir is a producer→consumer edge); that fallback needs the work directories to still exist locally and is best-effort (a note reports which source was used and how many task work dirs were missing). Either way this layer is informational only — it never affects the “identical” verdict or --fail-on-change.
  • Did the tools change? A per-process software & versions diff of the container image(s) and Conda package spec(s) each process ran with — read straight from the run cache, so no work directories are needed. This is the layer that catches biocontainers/fastqc:0.11.9 → biocontainers/fastqc:0.12.1 (or a bumped bioconda::salmon= pin) directly, rather than leaving it buried in per-task detail. A software change counts toward the “identical” verdict and --fail-on-change.
  • Did a task change? Per-task diffs of status, exit code, container, script, requested resources, and measured usage.
  • What failed, and why? A top-level failure rollup that gathers every failed task — detected from the cached status/exit fields (an explicit FAILED/ABORTED status or a non-zero exit code), so no work directories are needed — and rolls them up by (process, status, exit) signature, counted per run and sorted by biggest blast radius first. A signature seen only in Run B is flagged new (a regression), one in Run A but gone in Run B is resolved, and one in both is persistent, so you can tell at a glance whether a re-run introduced, fixed, or carried over a failure — e.g. a CALL process that started exiting 137 (OOM-killed) only in Run B. A run-level error state is surfaced too, even when no individual task failure was recorded. This is a summary of the per-task status/exit already shown in the task layer, so it is informational only and never separately affects the “identical” verdict or --fail-on-change.
  • Did anything get slower or heavier? A dedicated performance-regressions layer flags matched tasks whose runtime (realtime) or peak memory (peak_rss, peak_vmem) changed by at least --perf-threshold percent (default 25), sorted worst-regression first. Tasks that share the same cache hash are marked “same work”, so a +200% realtime on identical work stands out from a slowdown that also changed what ran. Improvements (Run B faster/leaner) are shown too, but only true regressions are counted.
  • Was anything over- or under-provisioned? A resource-efficiency layer that, for each process in each run, compares what it requested (cpus, memory) against what it actually peaked at (%cpu, peak_rss) — both read straight from the run cache, so no work directories are needed and it is always computed. Each process is classified per run: over-provisioned (used under 50% of the reservation — e.g. “requested 32 GB, peaked at 4 GB” — so the allocation, and often the cost, was wasted), tight (used 90%+ of it, a risk of OOM kills or CPU throttling), or right-sized in between. Because provisioning is a tuning signal rather than a correctness change, this layer is informational only — it never affects the “identical” verdict or --fail-on-change.
  • Was work reused? Cached-task counts and total task realtime, plus a recompute count — matched tasks (present in both runs) whose cache hash differs, i.e. work that was re-executed rather than resumed — so you can see whether a re-run actually recomputed anything, and why.
  • Did the results change? With --diff-outputs, the files each matched task wrote to its work directory are compared — by size first, then a streamed SHA-256 for same-size files — and classified as added / removed / changed / unchanged. For a changed file that is text on both sides, it also produces a line-level diff (reusing the same LineDiff engine as --diff-logs), bounded to the first --outputs-max-lines lines (default 1000) and a hard byte cap; binary files (a NUL byte in the head) show only the size/hash change. Tasks that resumed from cache share a work directory, so they short-circuit to “identical”; the interesting cases are recomputed tasks whose outputs actually differ. This layer needs the work directories to still exist locally, and (unlike the always-changing performance metrics) an output change does count toward the “identical” verdict and --fail-on-change.
  • Did the published results change? With --published-a=<dir> and --published-b=<dir>, the two runs’ durable published result trees (their outdir / publishDir directories) are compared directly. Files are keyed by their path relative to each published root and classified as added / removed / changed / unchanged, using the same size → SHA-256 → line-diff engine (FileContentComparator) as --diff-outputs, and honouring --outputs-max-bytes / --outputs-max-lines. Symlinks are followed, so it works whether publishDir copied or symlinked. Unlike --diff-outputs, this reads the durable published results instead of the work directories, so it still works after the work dirs are gone (cleaned up, or on remote object storage). A published-file change does count toward the “identical” verdict and --fail-on-change.
  • Why did a task fail? With --diff-logs, the standard log files (.command.out, .command.err, .command.log) each matched task wrote are compared line by line, with the task’s exit code and status surfaced alongside — so a task that flipped from exit 0 to exit 1 shows both the change and the stderr that explains it. Reads are bounded (tailed to --logs-max-lines, default 200, and a hard byte cap) so a huge log never blows up memory, and cache-resumed tasks sharing a work directory short-circuit as identical. Because stdout/stderr legitimately varies between runs (timestamps, paths, ordering), this layer is informational only — it never affects the “identical” verdict or --fail-on-change (the exit-code change already does).

By default the report highlights meaningful changes and treats fields that always differ between two distinct runs (run name, session id, launch time, work directory, wall-clock time, measured resource usage, and each task’s cache hash) as context rather than “changes”. The task cache hash is treated this way because Nextflow folds the per-run session UUID into every task’s hash, so two independent runs always compute a different hash for every task even when the script, inputs and container are byte-identical — that difference is surfaced separately as the recompute count rather than flipping the verdict. The same rule applies to the parameters layer: launch options that routinely differ without changing what was executed (-name, -resume, -ansi-log, -with-tower, -with-weblog, -bg) are shown for context but not flagged as changes. Use --verbose when you want everything flagged.