fix(jobs): flush the checkpoint tail, so progress reflects reality

All three import jobs only checkpointed on a full batch, so the
remainder after the last one was never counted. A run shorter than
BATCH_SIZE never checkpointed at all: `scanned_count` stayed 0 against
a known `total_rows`, and the admin progress bar sat at zero for the
whole run and finished there.

Seen on a transcode_import run over 20 entries — 13 imported, 5
negatives, 2 already present, progress 0/20 throughout. The thumbnail
imports had it too, just less visibly: a 105-file run reported
`scanned_count: 100`, losing the tail rather than all of it.

Cursor-wise the final checkpoint is a no-op — the walk is finished, so
nothing resumes from it — but the scanned delta is what the progress
display reads, and it has to include the last partial batch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-08-30 19:41:41 +02:00
parent bf2f0dc2b2
commit 10c362a94a
3 changed files with 51 additions and 0 deletions
@@ -734,6 +734,19 @@ impl RecoverableJobHandler for ThumbDerivedImport {
}
}
// Flush the tail. The loop only checkpoints on a full batch, so the
// remainder after the last one was never counted — a run of fewer
// than BATCH_SIZE files reported `scanned_count: 0` against a known
// total and left the admin progress bar at zero for its whole life.
// Same fix in both imports and in transcode_import.
if since_checkpoint > 0
&& let Err(e) = store.checkpoint(Vec::new(), since_checkpoint as u64).await
{
return RunOutcome::Failed {
message: format!("final checkpoint: {e}"),
};
}
// Remove the size directories once genuinely empty, because ABSENCE
// is what step 10e gates the fallback removal on — not emptiness.
// Empty is momentary: an on-demand render can repopulate it the next