perf(blobs): batch chunk fsyncs into one durability sweep per upload

Storing a new file through CDC dedup issued sync_all + a parent-dir
fsync for every ~256 KB chunk (~8,200 fsyncs for a 1 GB upload), plus
one PG INSERT round-trip per chunk. The actual durability boundary is
the manifest INSERT: chunks only need to be durable before any PG row
references them, not one by one.

- BlobStorageBackend grows put_blob_from_bytes_unsynced + sync_blobs
  with conservative defaults (unsynced delegates to the synced write,
  sync_blobs is a no-op) so backends that don't opt in keep the
  per-write durability semantics. Remote stores are durable on PUT.
- LocalBlobBackend writes chunks without fsync and implements
  sync_blobs as a parallel sweep: every listed blob file (hard
  requirement) plus each distinct prefix directory exactly once
  (best-effort, same tier as fsync_parent_dir).
- DedupService::store_chunks writes new chunks unsynced, runs one
  sync_blobs sweep, then registers all new chunks in ONE batched
  UNNEST INSERT - durability before visibility, and the per-chunk PG
  round-trips collapse into one.
- Encrypted/Migration decorators forward both methods so the
  optimization survives encrypted-local and live-migration stacks.

https://claude.ai/code/session_013Bk4BMQEvR9QxCU7QXLRwv
This commit is contained in:
Claude
2026-06-10 09:55:02 +00:00
parent 2161292e2c
commit 9a181053bd
5 changed files with 367 additions and 69 deletions
@@ -60,6 +60,38 @@ pub trait BlobStorageBackend: Send + Sync + 'static {
/// without overwriting. Returns the number of bytes stored.
fn put_blob_from_bytes(&self, hash: &str, data: Bytes) -> BoxFut<'_, Result<u64, DomainError>>;
/// Store a blob from in-memory bytes **without forcing durability**.
///
/// Same idempotency contract as [`Self::put_blob_from_bytes`], but the
/// bytes may still sit in volatile caches (e.g. the OS page cache) when
/// the future resolves. Durability is only guaranteed after a subsequent
/// [`Self::sync_blobs`] covering this hash returns `Ok`. Callers MUST NOT
/// record a durable reference to the blob (e.g. a PostgreSQL row) before
/// that sync completes.
///
/// Default: delegates to `put_blob_from_bytes` (immediately durable),
/// pairing with the no-op `sync_blobs` default so backends that don't
/// opt in keep today's per-write durability semantics.
fn put_blob_from_bytes_unsynced(
&self,
hash: &str,
data: Bytes,
) -> BoxFut<'_, Result<u64, DomainError>> {
self.put_blob_from_bytes(hash, data)
}
/// Make previously written blobs durable in one batched operation.
///
/// Durability barrier for blobs written via `put_blob_from_bytes_unsynced`:
/// when this returns `Ok`, every listed blob is crash-safe. Local
/// filesystem backends fsync each listed blob file plus each distinct
/// parent directory once — one sweep per upload instead of two fsyncs
/// per chunk. Remote object stores are durable on PUT, so the default
/// is a no-op.
fn sync_blobs(&self, _hashes: &[String]) -> BoxFut<'_, Result<(), DomainError>> {
Box::pin(async { Ok(()) })
}
/// Stream the full blob content in chunks.
fn get_blob_stream(&self, hash: &str) -> BoxFut<'_, Result<BlobStream, DomainError>>;