feat(jobs): a transient backend failure pauses at its cursor instead of failing

Step 3 of docs/plan/jobs-handling-recoverable-error.md, and it
deliberately does NOT add the retry loop the plan sketched. Reasoning
below.

`RunOutcome::from_domain_error(cursor, context, err)` routes a failed
operation to `PausedRetryable` when the error is transient and `Failed`
otherwise. Handlers call it instead of reaching for `Failed`, so an
outage stops a long scan at its cursor rather than discarding it —
`Failed` is terminal, and only `Paused` resumes.

Applied to `backend_consistency`'s enumeration failure first, because
that is the case with the most to lose: the job fails the whole run on
an enumeration error, so a brief 503 partway through a million-object
bucket used to throw away the entire audit.

## Why no bounded retry loop in the engine

The plan said "bounded exponential backoff, ~5 attempts" in
`run_or_resume`, and also warned "do not double-retry — the AWS SDK
already retries internally, so a second layer above it multiplies".
Checking before writing it, there are already TWO layers:

  * the AWS SDK retries internally;
  * `RetryBlobBackend` wraps every remote backend with exponential
    backoff — 3 retries, 100 ms initial, ×2, 10 s cap, all tunable via
    OXICLOUD_STORAGE_RETRY_*, and applied in di.rs for non-Local
    backends.

A third layer multiplies rather than adds: one logical operation could
span SDK × decorator × engine attempts, turning a brief outage into
minutes of held `migration_readonly` — the precise failure this plan
exists to stop.

Retrying here would also re-run a SCAN, not an operation. The retrying
belongs where it already is, per request; what was genuinely missing is
the conversion of an exhausted-retry failure into a resumable pause with
a reason, which is what this commit adds. If the attempt budget needs
tuning, `OXICLOUD_STORAGE_RETRY_MAX_RETRIES` is the knob, and it applies
to every backend call rather than only to jobs.

## Tests

`transient_failure_pauses_with_a_reason_and_keeps_the_cursor` asserts
the three things that matter: status Paused, cursor preserved,
`error_message` naming the cause. `permanent_failure_still_fails_terminally`
is the control — without it the classification could be inert and
everything would simply pause, which would look like success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-09-07 19:37:16 +02:00
parent a7e25eea76
commit 303a0421c2
2 changed files with 180 additions and 6 deletions
@@ -506,12 +506,18 @@ impl RecoverableJobHandler for BackendConsistencyCheck {
// Whether the enumeration died on page 1 or page 900,
// the audit did not complete, and the operator needs to
// know that rather than read a green run.
return RunOutcome::Failed {
message: format!(
"backend enumeration failed on {}: {e}",
backend.backend_type()
),
};
// Transient (throttle, 5xx, connection reset) pauses at
// the cursor so a resume continues the sweep;
// everything else fails terminally. Losing a
// half-finished audit of a million-object bucket to a
// brief 503 is the case this distinction exists for —
// the retry decorator has already given up by the time
// the error arrives here.
return RunOutcome::from_domain_error(
cursor.as_ref().map(|s| s.as_bytes()),
&format!("backend enumeration failed on {}", backend.backend_type()),
&e,
);
}
};