Skip to content
nf-diff

How it works

nf-diff never re-executes anything. It reconstructs each run entirely from local state:

  1. Resolve the run — RunLoader looks up the identifier (name or session-id prefix) in .nextflow/history via Nextflow’s HistoryFile, capturing run-level metadata: run name, session id, status, revision, command, launch time, and duration.
  2. Hydrate tasks — it opens the run’s LevelDB cache (.nextflow/cache/<sessionId>) read-only through Nextflow’s CacheDB, reading every TraceRecord into a TaskInfo (both human-formatted and raw values). Cache opens hold an exclusive lock, so reads retry with backoff when the lock is briefly contended (e.g. by nextflow log or an IDE indexing the cache).
  3. Compare across the layers — RunComparator produces (the first six are always computed; performance regressions and resource efficiency are derived, and outputs / logs / DAG are opt-in):
    • Metadata — run-level field diffs. The raw launch command is kept here as context, since the structured Parameters layer is now the authoritative view of what changed on the command line.
    • Parameters — each run’s launch command is parsed by CommandParams into Nextflow options (single-dash, e.g. -profile) and pipeline params (double-dash, e.g. --genome), then diffed flag by flag. When the command referenced a -params-file, that JSON/YAML is read (relative paths resolved against --dir), flattened into dotted keys (--genome.build), and merged underneath the command-line flags — Nextflow’s precedence, so an explicit --flag overrides the file. Each value is tagged with its source (CLI, file, CLI+file). A flag present in only one run shows a missing value on the other side; noisy options (-name, -resume, …) are flagged “obvious” and excluded from change detection unless --verbose is set.
    • Configuration — ConfigLoader rebuilds each run’s effective config with Nextflow’s own ConfigBuilder (the same machinery behind nextflow config), applying the run’s recorded -profile/-c options, then flattens it to dotted keys and diffs the two. The params scope is excluded here — it belongs to the Parameters layer. Because Nextflow does not persist the fully merged config in its history/cache, this is resolved from the project’s config files as they exist now; it is authoritative for -profile/-c-driven differences rather than a byte-for-byte snapshot at launch time, and the report says so.
    • Processes — task counts per process, classified as added / removed / changed / unchanged.
    • Software & versions — for each process, the distinct container image(s) and Conda package spec(s) its tasks ran with are collected from the cached trace records (container, conda) and diffed. A process present in only one run is added / removed; a process in both whose container or Conda set differs is changed. Because it reads only the cache, it needs no work directories and is always computed. A changed software environment does count toward “identical” and --fail-on-change — which also means a Conda-only change (not part of the per-task field set) now breaks identity.
    • Tasks — matched across runs by task name (falling back to process + tag), with per-field diffs. “Obvious” always-changing fields are shown for context but excluded from change detection unless --verbose is set. The cache hash is one of those obvious fields — Nextflow mixes the per-run session UUID into it, so it differs between any two independent runs even for byte-identical work — so a bare hash change never flips the “identical” verdict on its own; it is instead tallied as the recompute count (matched tasks whose cache hash differs — work that was re-executed rather than resumed). Under --verbose the hash is flagged like any other obvious field.
    • Performance regressions — derived from the matched tasks’ numeric trace metrics (realtime, peak_rss, peak_vmem). For each metric that moved by at least --perf-threshold percent, a signed delta is recorded and the list is sorted worst-regression first; each entry notes whether both tasks shared a cache hash (“same work”). This layer is computed from fields that change between runs, so by default it is informational only — it never affects the “identical” verdict or --fail-on-change. Under --verbose, where always-changing fields become meaningful, a flagged regression does count as a difference.
    • Outputs (only with --diff-outputs) — OutputComparator walks each matched task’s work directory, skipping staged inputs (symlinks) and Nextflow control files (.command.*, .exitcode), and compares the remaining real files by path. Files are matched by size first; same-size files are then compared by a streamed SHA-256 (capped by --outputs-max-bytes, which leaves oversized same-size files content-unverified). For a changed file that is text on both sides (no NUL byte in its head), it then computes a line-level diff via the shared LineDiff engine, bounded to the first --outputs-max-lines lines and a hard byte cap. Tasks that resolved to the same physical work directory (a cache resume) short-circuit as identical. Unlike performance regressions, an output-file change does count toward “identical” and --fail-on-change.
    • Published outputs (only with --published-a + --published-b) — PublishedComparator walks each run’s published result tree (its outdir / publishDir directory), keying files by their path relative to each root, and classifies them as added / removed / changed / unchanged. It reuses the same FileContentComparator engine as the work-dir outputs layer: files are matched by size first, then a streamed SHA-256 for same-size files (capped by --outputs-max-bytes), and a changed text file gets a line-level diff bounded to --outputs-max-lines. Symlinks are followed, so it works whether publishDir copied or symlinked. Unlike --diff-outputs, this reads the durable published results rather than the work directories, so it still works after the work dirs are cleaned up. A published-file change does count toward “identical” and --fail-on-change.
    • Logs (only with --diff-logs) — LogComparator reads the standard log files (.command.out, .command.err, .command.log) from each matched task’s work directory and produces a line-level diff of each via LineDiff, alongside the task’s exit code and status. Reads are bounded (only the tail up to --logs-max-lines and a hard byte cap are loaded) so large logs never exhaust memory, and tasks sharing a physical work directory short-circuit as identical. Like performance regressions, this layer is informational only — task stdout/stderr legitimately varies between runs, so it never affects “identical” or --fail-on-change.
    • DAG / wiring (only with --diff-dag) — DagComparator reconstructs each run’s process→process graph and diffs the two edge sets, preferring the authoritative Nextflow data-lineage store and falling back to work-dir symlinks. When the run’s project directory has a .lineage/ store (from lineage.enabled = true), LineageStore reads it directly off disk: it indexes the run’s TaskRun records (matched by session id) by their LID hash, then reconstructs a producer→consumer edge for every input LID reference (lid://<producerTaskHash>/…) a consumer task recorded. This is exactly the provenance Nextflow persisted, so it needs no work directories and is not best-effort. Without a lineage store, it falls back to symlinkGraphOf: it walks each task’s staged input symlinks and, for any that resolve into another task’s work directory (walking the parent chain so a link into a nested output subdir still attributes to its producer), records an edge; links resolving outside every work dir are external inputs and yield no edge — reporting how many task work dirs were missing so a partial reconstruction isn’t read as authoritative. Each run’s graph records its source (LINEAGE/SYMLINK/NONE) and the layer’s note says which was used. Because the fallback is best-effort and the layer is a summary regardless, it is informational only — it never affects “identical” or --fail-on-change.
  4. Render — HtmlReportRenderer emits the standalone HTML document, JsonReportRenderer the structured JSON (--format=json), or MarkdownReportRenderer the Markdown report (--format=md).

--only / --exclude narrow the comparison via a ProcessFilter before rendering, so process- and task-level output is restricted to the processes you care about.