feat(jobs): PausedRetryable — an outcome the engine can act on

Step 2 of docs/plan/jobs-handling-recoverable-error.md. A handler could
say `Completed`, `Paused` or `Failed`, so a transient backend failure was
flattened into `Failed` before the engine saw it — "the provider is
down" and "this data is wrong" were indistinguishable, and `Failed` is
terminal, so an outage threw away a partially-complete migration.

`PausedRetryable { cursor, reason }` lands as `Paused` in the row, so
resume is unchanged. What differs is `error_message`:

  | outcome           | meaning                          | resumes?     |
  |-------------------|----------------------------------|--------------|
  | Failed            | the data or request is wrong     | no, terminal |
  | Paused            | an operator asked it to stop     | yes          |
  | PausedRetryable   | the environment failed           | yes, + why   |

Without the reason a paused run is an unexplained one — and a paused
`backend_migration` still holds `migration_readonly`, refusing writes
application-wide, so "why is this app read-only" has to be answerable
from the row.

`mark_paused_retryable` is a separate store method rather than an extra
argument on `mark_paused`: only one of them writes `error_message`, and
a `reason: Option<&str>` parameter would let a caller produce a Paused
row carrying an error message and no error — the exact state this exists
to distinguish from.

Reported as `JobOutcome::ok`, not `err`. The run did not fail; it
stopped and can be resumed. A red job in the panel that a Resume click
fixes reads as a bug rather than as a decision waiting to be made. The
`extra` carries `retryable: true` and the reason so the panel can say
which kind of pause it was. Audited too, since a run that stopped on an
outage is an operational event someone has to act on.

## Also: Azure now classifies its errors

The previous commit said Azure could wait for the official-SDK
migration. That was wrong — `azure_core::error::ErrorKind::HttpResponse`
carries the status on the archived 0.21, so `azure_domain_error` works
today. It matters because Azure is the backend this whole plan was
written for.

Applied at five sites including the 256-shard enumeration walk, where
`backend_consistency` fails the entire run on an error, so a throttle
partway through should be retryable rather than discarding the sweep.

Per Ed's call on the ambiguous case: a deterministic 500 — Azurite
answering the CRC64 ranged GET, every time — classifies as transient
because nothing at this layer can tell it from a passing one. Retry as
if transient, let the bounded cap convert the difference into a Paused
run, and let Ops decide to resume or cancel.

Not yet wired: the engine's bounded backoff (step 3). Note for that
work — backoff already exists in the AWS SDK internally AND in
`RetryBlobBackend` (100 ms, ×2, 10 s cap, 3 retries). A third naive
layer would multiply, so the plan's "do not double-retry" needs
measuring before adding one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-09-07 19:31:16 +02:00
parent 465fbe2480
commit a7e25eea76
3 changed files with 221 additions and 15 deletions
@@ -143,9 +143,10 @@ impl BlobStorageBackend for AzureBlobBackend {
})?;
let file_size = data.len() as u64;
client.put_block_blob(data).await.map_err(|e| {
DomainError::internal_error("Azure", format!("Failed to upload blob {hash}: {e}"))
})?;
client
.put_block_blob(data)
.await
.map_err(|e| azure_domain_error(format!("Failed to upload blob {hash}"), &e))?;
let _ = fs::remove_file(&source_path).await;
Ok(file_size)
@@ -169,9 +170,10 @@ impl BlobStorageBackend for AzureBlobBackend {
// `Bytes` converts into `azure_core::Body` by reference count —
// the old `data.to_vec()` copied every chunk once more.
client.put_block_blob(data).await.map_err(|e| {
DomainError::internal_error("Azure", format!("Failed to upload blob {hash}: {e}"))
})?;
client
.put_block_blob(data)
.await
.map_err(|e| azure_domain_error(format!("Failed to upload blob {hash}"), &e))?;
Ok(size)
})
@@ -190,9 +192,10 @@ impl BlobStorageBackend for AzureBlobBackend {
Box::pin(async move {
let client = self.blob_client(&hash);
let size = data.len() as u64;
client.put_block_blob(data).await.map_err(|e| {
DomainError::internal_error("Azure", format!("Failed to upload blob {hash}: {e}"))
})?;
client
.put_block_blob(data)
.await
.map_err(|e| azure_domain_error(format!("Failed to upload blob {hash}"), &e))?;
Ok(size)
})
}
@@ -371,9 +374,9 @@ impl BlobStorageBackend for AzureBlobBackend {
if status == Some(azure_core::StatusCode::NotFound) {
Ok(())
} else {
Err(DomainError::internal_error(
"Azure",
format!("Failed to delete blob {hash}: {e}"),
Err(azure_domain_error(
format!("Failed to delete blob {hash}"),
&e,
))
}
}
@@ -558,12 +561,16 @@ impl BlobStorageBackend for AzureBlobBackend {
while let Some(page) = pages.next().await {
let page = page.map_err(|e| {
DomainError::internal_error(
"Blob",
// Classified: `backend_consistency` fails the whole
// run on an enumeration error, so a throttle
// partway through the 256-shard walk should be
// retryable rather than discarding the sweep.
azure_domain_error(
format!(
"Azure ListBlobs failed on shard {shard:02x} of container '{}': {e}",
"Azure ListBlobs failed on shard {shard:02x} of container '{}'",
self.container_name
),
&e,
)
})?;
@@ -654,6 +661,50 @@ impl BlobStorageBackend for AzureBlobBackend {
}
}
/// Wrap an `azure_core` error as a `DomainError` that says whether
/// retrying it could help. Azure counterpart of `s3_domain_error`.
///
/// `azure_core::error::ErrorKind::HttpResponse` carries the status, so
/// this works on the archived 0.21 SDK — no need to wait for the
/// official-crate migration. That matters because Azure is the backend
/// the retry-then-pause plan was written for: a ranged GET carrying
/// `x-ms-range-get-content-crc64` that Azurite answers 500 to, retried
/// forever by `azure_core` while `migration_readonly` refused writes
/// application-wide.
///
/// Transient: 5xx, 429, 408. Also `Io` — connection resets, DNS, TLS.
/// Permanent: other 4xx (credentials, missing container, malformed
/// request), `DataConversion`, `Credential`.
///
/// **A deterministic 500 still classifies as transient**, and that is
/// deliberate rather than an oversight. Nothing at this layer can tell
/// "this provider is briefly unwell" from "this provider will answer
/// 500 to this exact request forever" — the Azurite CRC64 case is the
/// second wearing the clothes of the first. So the policy is to retry
/// as if transient and let the bounded attempt cap turn the difference
/// into a Paused run an operator can act on.
pub(crate) fn azure_domain_error(context: String, err: &azure_core::Error) -> DomainError {
use azure_core::error::ErrorKind as AzKind;
let transient = match err.kind() {
AzKind::HttpResponse { status, .. } => {
let code = u16::from(*status);
code >= 500 || code == 429 || code == 408
}
AzKind::Io => true,
AzKind::DataConversion | AzKind::Credential | AzKind::MockFramework | AzKind::Other => {
false
}
};
let message = format!("{context}: {err}");
if transient {
DomainError::transient_backend("Azure", message)
} else {
DomainError::internal_error("Azure", message)
}
}
#[cfg(test)]
mod tests {
use super::*;