feat(storage): a backend that never answers is now a transient failure
Ed pulled the network mid-migration and got nothing: no log, no pause,
after more than two minutes. The cause is not the classification work
that preceded this — it is that there was no error to classify.
Pull a network on an ESTABLISHED TCP connection and there is no RST and
no ICMP. The peer simply stops answering and the socket read blocks
until the OS abandons retransmission, on the order of fifteen minutes.
For that whole window the job is neither running nor failed. Nothing
retries, because nothing failed. It looks exactly like a slow migration.
A refused connection is instant and does surface, which is what made
the earlier `127.0.0.1` test look reassuring. It exercised the one
network failure that cannot hang.
## Two layers, because one does not fit
`TimeoutBlobBackend` is innermost, below retry — a hang has to become an
error before any layer above can react to it. Bounds are per operation
class, because one number cannot fit both a HEAD and a 5 GB upload:
metadata 30s exists / size / delete / init / health / list
open 60s time to FIRST BYTE, not transfer duration
write off the whole transfer is inside the future, so any
bound here is also a maximum upload duration
Write is unbounded by default deliberately: guessing it wrong truncates
legitimate uploads, which is worse than the hang it would prevent. All
three are configurable (`OXICLOUD_STORAGE_TIMEOUT_*_MS`, 0 = unbounded).
The S3 client also gets what it could always have had. It was built from
a bare `config::Builder::new()`, which carries NO `TimeoutConfig` at
all — so `SdkError::TimeoutError`, an arm `s3_domain_error` already
handles, was unreachable. It now sets connect/read timeouts plus
stalled-stream protection, which measures throughput rather than
elapsed time and is therefore the correct instrument for a stream: it
bounds a stalled upload without capping how long a large one may take.
## Local is not the justification
Ed's correction, and it is right: a local path is reached through the
kernel, and the kernel owns that timeout. iSCSI gives up after
`replacement_timeout` (120s default) and returns an I/O error; NVMe-oF
and soft-mounted NFS behave the same. Those arrive as `io::Error` and
`local_io_error` already classifies them. Local passes through the
decorator only because a uniform chain beats a conditional one, and a
bound that never fires costs nothing.
The real asymmetry is Azure: its 0.21 client has no timeout knob short
of a custom transport, and the SDK migration is deferred. That is why
this lives in the chain rather than being configured per SDK.
Also fixes the log gap: the timeout warns with the wrapper, backend,
operation and bound, so a stalled layer is visible before the pause
rather than only afterwards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -2,6 +2,8 @@ use std::env;
|
||||
use std::path::PathBuf;
|
||||
use std::time::Duration;
|
||||
|
||||
use crate::infrastructure::services::timeout_blob_backend::TimeoutPolicy;
|
||||
|
||||
/// Cache configuration
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct CacheConfig {
|
||||
@@ -268,6 +270,9 @@ pub struct StorageConfig {
|
||||
pub encryption: EncryptionConfig,
|
||||
/// Retry policy for remote backends.
|
||||
pub retry: RetryConfig,
|
||||
/// Wall-clock bounds on backend calls, so a stalled endpoint
|
||||
/// surfaces as a transient error instead of hanging indefinitely.
|
||||
pub timeout: TimeoutPolicy,
|
||||
}
|
||||
|
||||
/// Which blob storage backend to use.
|
||||
@@ -1295,6 +1300,7 @@ impl Default for StorageConfig {
|
||||
cache: BlobCacheConfig::default(),
|
||||
encryption: EncryptionConfig::default(),
|
||||
retry: RetryConfig::default(),
|
||||
timeout: TimeoutPolicy::default(),
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -3722,6 +3728,29 @@ impl AppConfig {
|
||||
config.storage.retry.backoff_multiplier = n;
|
||||
}
|
||||
|
||||
// Backend call timeouts. `0` means "unbounded" for that class,
|
||||
// which is the default for writes — see `TimeoutPolicy`.
|
||||
for (var, slot) in [
|
||||
(
|
||||
"OXICLOUD_STORAGE_TIMEOUT_METADATA_MS",
|
||||
&mut config.storage.timeout.metadata,
|
||||
),
|
||||
(
|
||||
"OXICLOUD_STORAGE_TIMEOUT_OPEN_MS",
|
||||
&mut config.storage.timeout.open,
|
||||
),
|
||||
(
|
||||
"OXICLOUD_STORAGE_TIMEOUT_WRITE_MS",
|
||||
&mut config.storage.timeout.write,
|
||||
),
|
||||
] {
|
||||
if let Ok(v) = env::var(var)
|
||||
&& let Ok(n) = v.parse::<u64>()
|
||||
{
|
||||
*slot = (n > 0).then(|| std::time::Duration::from_millis(n));
|
||||
}
|
||||
}
|
||||
|
||||
// OIDC configuration
|
||||
if let Ok(v) = env::var("OXICLOUD_OIDC_ENABLED") {
|
||||
config.oidc.enabled = v.parse::<bool>().unwrap_or(false);
|
||||
|
||||
+28
-1
@@ -329,7 +329,8 @@ impl AppServiceFactory {
|
||||
// just before the struct init.
|
||||
let active_backend_name = Arc::new(std::sync::RwLock::new(active_backend_name));
|
||||
|
||||
// Stack decorators: retry → encryption → cache (inner-to-outer).
|
||||
// Stack decorators: timeout → retry → encryption → cache
|
||||
// (inner-to-outer).
|
||||
//
|
||||
// Encryption is applied INSIDE build_entry_backend (per-entry
|
||||
// key), so it's already on the base returned above when the
|
||||
@@ -343,6 +344,32 @@ impl AppServiceFactory {
|
||||
// gates the "remote-only" decorators the same as before.
|
||||
let mut blob_backend: Arc<dyn BlobStorageBackend> = base_backend;
|
||||
|
||||
// Timeout decorator — INNERMOST, and applied to every backend
|
||||
// kind including Local.
|
||||
//
|
||||
// It has to sit below retry: a call that never returns produces
|
||||
// no error, so retry has nothing to react to and the job never
|
||||
// pauses. Converting the hang into a transient error first is
|
||||
// what gives every layer above it something to act on.
|
||||
//
|
||||
// Unconditional by design. Retry is gated on "not Local" because
|
||||
// the kernel already retries local I/O, but a bound that never
|
||||
// fires is free, and keeping the chain uniform avoids a class of
|
||||
// backend-specific surprise.
|
||||
{
|
||||
use crate::infrastructure::services::timeout_blob_backend::TimeoutBlobBackend;
|
||||
let policy = self.config.storage.timeout.clone();
|
||||
if policy.is_enabled() {
|
||||
blob_backend = Arc::new(TimeoutBlobBackend::new(blob_backend, policy.clone()));
|
||||
tracing::info!(
|
||||
metadata_ms = policy.metadata.map(|d| d.as_millis() as u64),
|
||||
open_ms = policy.open.map(|d| d.as_millis() as u64),
|
||||
write_ms = policy.write.map(|d| d.as_millis() as u64),
|
||||
"Blob storage timeout decorator enabled"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// Retry decorator (for remote backends)
|
||||
if self.config.storage.retry.enabled && active_backend_kind != StorageBackendType::Local {
|
||||
use crate::infrastructure::services::retry_blob_backend::{
|
||||
|
||||
Reference in New Issue
Block a user