feat(jobs): a transient backend failure pauses at its cursor instead of failing
Step 3 of docs/plan/jobs-handling-recoverable-error.md, and it
deliberately does NOT add the retry loop the plan sketched. Reasoning
below.
`RunOutcome::from_domain_error(cursor, context, err)` routes a failed
operation to `PausedRetryable` when the error is transient and `Failed`
otherwise. Handlers call it instead of reaching for `Failed`, so an
outage stops a long scan at its cursor rather than discarding it —
`Failed` is terminal, and only `Paused` resumes.
Applied to `backend_consistency`'s enumeration failure first, because
that is the case with the most to lose: the job fails the whole run on
an enumeration error, so a brief 503 partway through a million-object
bucket used to throw away the entire audit.
## Why no bounded retry loop in the engine
The plan said "bounded exponential backoff, ~5 attempts" in
`run_or_resume`, and also warned "do not double-retry — the AWS SDK
already retries internally, so a second layer above it multiplies".
Checking before writing it, there are already TWO layers:
* the AWS SDK retries internally;
* `RetryBlobBackend` wraps every remote backend with exponential
backoff — 3 retries, 100 ms initial, ×2, 10 s cap, all tunable via
OXICLOUD_STORAGE_RETRY_*, and applied in di.rs for non-Local
backends.
A third layer multiplies rather than adds: one logical operation could
span SDK × decorator × engine attempts, turning a brief outage into
minutes of held `migration_readonly` — the precise failure this plan
exists to stop.
Retrying here would also re-run a SCAN, not an operation. The retrying
belongs where it already is, per request; what was genuinely missing is
the conversion of an exhausted-retry failure into a resumable pause with
a reason, which is what this commit adds. If the attempt budget needs
tuning, `OXICLOUD_STORAGE_RETRY_MAX_RETRIES` is the knob, and it applies
to every backend call rather than only to jobs.
## Tests
`transient_failure_pauses_with_a_reason_and_keeps_the_cursor` asserts
the three things that matter: status Paused, cursor preserved,
`error_message` naming the cause. `permanent_failure_still_fails_terminally`
is the control — without it the classification could be inert and
everything would simply pause, which would look like success.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -506,12 +506,18 @@ impl RecoverableJobHandler for BackendConsistencyCheck {
|
||||
// Whether the enumeration died on page 1 or page 900,
|
||||
// the audit did not complete, and the operator needs to
|
||||
// know that rather than read a green run.
|
||||
return RunOutcome::Failed {
|
||||
message: format!(
|
||||
"backend enumeration failed on {}: {e}",
|
||||
backend.backend_type()
|
||||
),
|
||||
};
|
||||
// Transient (throttle, 5xx, connection reset) pauses at
|
||||
// the cursor so a resume continues the sweep;
|
||||
// everything else fails terminally. Losing a
|
||||
// half-finished audit of a million-object bucket to a
|
||||
// brief 503 is the case this distinction exists for —
|
||||
// the retry decorator has already given up by the time
|
||||
// the error arrives here.
|
||||
return RunOutcome::from_domain_error(
|
||||
cursor.as_ref().map(|s| s.as_bytes()),
|
||||
&format!("backend enumeration failed on {}", backend.backend_type()),
|
||||
&e,
|
||||
);
|
||||
}
|
||||
};
|
||||
|
||||
|
||||
Reference in New Issue
Block a user