fix(migration): counters must describe the run, not the current segment
Ed's completed migration reported `copied: 0` beside `scanned_count: 2522`. Both numbers were accurate; they were measuring different things and neither said which. `scanned_count` was cumulative because `checkpoint` had been persisting it after every batch. `copied` / `skipped` / `failed` / `source_missing` were plain locals initialised to zero at the top of the handler, written to `stats` only via `merge_stats` — which is engine-only and fires on `Completed`, a state a paused run never reaches. So every pause threw them away and every resumed segment started counting from nothing. ## The fix has two halves, and only one is the obvious one Restoring on resume is the obvious half: the four counters now seed from `stats` exactly as `already_scanned` already did. The half that actually matters is WHEN they are written. Restoring is useless if nothing durable exists to restore from, so counters are persisted per batch through a new handler-callable `checkpoint_counters`, immediately after the cursor checkpoint. `merge_stats` stays engine-only; the end-of-run summary write is unchanged. Two deliberate choices: * **Absolute values, not deltas.** The merge is last-write-wins and the handler owns the running total. Deltas would double-count on exactly the replay path that produced 2522 scanned against 2022 rows. * **A counter-write failure warns, it does not fail the run.** The cursor is the correctness-critical write; these are reporting. Losing a migration to a hiccuping stats merge is the wrong trade. `scanned_count()` is now a default method over the new generic `stat_u64(key)` rather than a second near-identical query. ## Not fixed, and not claimed to be The 2522-vs-2022 overshoot itself. This makes it legible — cumulative and per-segment values now both land on the row — but whether the final segment re-walked rows it had already counted is a cursor question that needs reproducing, not inferring. The counters should let it be observed next time rather than reconstructed afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -557,16 +557,33 @@ impl RecoverableJobHandler for BackendMigrationService {
|
||||
},
|
||||
};
|
||||
|
||||
let mut copied_count = 0u64;
|
||||
// Restored on Resume, exactly like `already_scanned` above.
|
||||
//
|
||||
// These used to start at zero on every segment while
|
||||
// `scanned_count` was restored, so one counter described the
|
||||
// migration and the other four described the current segment.
|
||||
// A run that paused and resumed then reported `copied: 0`
|
||||
// beside a `scanned_count` in the thousands — the numbers were
|
||||
// measuring different things and only one of them said so.
|
||||
// `checkpoint_counters` below persists them per batch so a
|
||||
// pause cannot discard them.
|
||||
let restore = |key: &'static str| async move {
|
||||
if is_fresh {
|
||||
0
|
||||
} else {
|
||||
store.stat_u64(key).await.unwrap_or(0)
|
||||
}
|
||||
};
|
||||
let mut copied_count = restore("copied").await;
|
||||
// Populated by the smart-skip probe below: target blob
|
||||
// already exists at the current head format+key, so a
|
||||
// rewrite would be identical bytes. Cheap (15-byte range
|
||||
// read via `is_at_head_format`), massive latency win on
|
||||
// resume + on backends where the source was rotated to the
|
||||
// same key as the target already had.
|
||||
let mut skipped_count: u64 = 0;
|
||||
let mut failed_count = 0u64;
|
||||
let mut source_missing_count = 0u64;
|
||||
let mut skipped_count: u64 = restore("skipped").await;
|
||||
let mut failed_count = restore("failed").await;
|
||||
let mut source_missing_count = restore("source_missing").await;
|
||||
|
||||
loop {
|
||||
// Cooperative cancel poll between batches.
|
||||
@@ -923,6 +940,29 @@ impl RecoverableJobHandler for BackendMigrationService {
|
||||
message: format!("checkpoint: {e}"),
|
||||
};
|
||||
}
|
||||
// Persist the counters alongside the cursor. Absolute
|
||||
// values, not deltas — the merge is last-write-wins, and
|
||||
// the checkpoint above already made this batch's work part
|
||||
// of the durable position. A failure here is logged but
|
||||
// does NOT fail the run: the cursor is the correctness-
|
||||
// critical write, these are reporting.
|
||||
let counters: serde_json::Map<String, serde_json::Value> = serde_json::json!({
|
||||
"copied": copied_count,
|
||||
"skipped": skipped_count,
|
||||
"failed": failed_count,
|
||||
"source_missing": source_missing_count,
|
||||
})
|
||||
.as_object()
|
||||
.cloned()
|
||||
.unwrap_or_default();
|
||||
if let Err(e) = store.checkpoint_counters(&counters).await {
|
||||
tracing::warn!(
|
||||
target: "oxicloud::migration",
|
||||
event = "backend_migration.counter_persist_failed",
|
||||
error = %e,
|
||||
"could not persist per-batch counters; totals may under-report after a resume"
|
||||
);
|
||||
}
|
||||
// Bump the shared progress snapshot so the server-status
|
||||
// header middleware surfaces fresh numbers on every
|
||||
// user's next API call. Guard is held only for a struct
|
||||
|
||||
Reference in New Issue
Block a user