9c63f9969a
Step 7 of docs/plan/derived-blobs.md, write path first — the plan is explicit that fixing it before the import means transcode_import only has to handle history, not a moving target. ImageTranscodeService now reads and writes storage.content_derived_blobs under kind='transcode', keyed by the BLAKE3 of the SOURCE content. The hash is threaded in from file_retrieval_service, which already holds it as dto.content_hash; hashing here would be a BLAKE3 over the whole file on every request. Callers without one (external mounts) keep the local cache untouched, which is what the service did before this tier existed. Negative verdicts become rows rather than zero-byte .skip files. A transcode that came out larger is deterministic in the content, so it is worth remembering; the row survives moka eviction, a restart, and the deletion of .transcoded/, none of which the marker does. Only that verdict is persisted — a timeout or a read error returns Err and is recorded nowhere, because a momentary failure written here would mark a perfectly transcodable image hopeless with nothing to retry it. Representation is a NULL blob_hash, per the plan: a sentinel hash would stop blob_hash naming a real Blob and every consumer would need to learn the exception. A CHECK keeps blob_hash and content_type NULL together — a type without bytes describes nothing, bytes without a type cannot be served. Two consumers had to be corrected for NULLs first, both of which would have broken on the first negative row ever written: * satellites_consistency reported them as derived_dangling_blob at data_loss severity. SQL comparison against NULL is NULL, so EXISTS was false and a row correctly pointing at nothing read as an artifact that had gone missing. * blob_reference_sources::list_referenced_blobs decodes blob_hash into String, so the first NULL would have failed the decode and taken the whole enumeration down. It would also have been wrong if it decoded — a negative row holds no reference, which is why the counting forms (WHERE blob_hash = <hash>) already exclude it for free. lookup_derived returns a three-way answer because Option collapses the two cases a caller deciding whether to spend a decode most needs apart: never attempted, versus attempted and known not worth it. DedupService is attached after construction via a OnceLock. DI builds the transcode service ~240 lines before DedupService exists, and the retrieval path that needs it is wired earlier still, so a constructor argument would mean reordering more than this is worth. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
111 lines
3.5 KiB
Rust
111 lines
3.5 KiB
Rust
//! Image Transcode Port - Application layer abstraction for image transcoding.
|
|
//!
|
|
//! This module defines the port (trait) for on-demand image format conversion
|
|
//! (e.g., JPEG/PNG → WebP), keeping the application and interface layers
|
|
//! independent of specific image processing implementations.
|
|
|
|
use crate::common::errors::DomainError;
|
|
use bytes::Bytes;
|
|
|
|
/// Supported output formats for image transcoding.
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
|
|
pub enum OutputFormat {
|
|
/// WebP format — best current browser support with good compression.
|
|
WebP,
|
|
// Future: Avif, JpegXl
|
|
}
|
|
|
|
impl OutputFormat {
|
|
/// Get the file extension for this format.
|
|
pub fn extension(&self) -> &'static str {
|
|
match self {
|
|
OutputFormat::WebP => "webp",
|
|
}
|
|
}
|
|
|
|
/// Get the MIME type for this format.
|
|
pub fn mime_type(&self) -> &'static str {
|
|
match self {
|
|
OutputFormat::WebP => "image/webp",
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Browser image format capabilities detected from the Accept header.
|
|
#[derive(Debug)]
|
|
pub struct BrowserCapabilities {
|
|
pub supports_webp: bool,
|
|
pub supports_avif: bool,
|
|
}
|
|
|
|
impl BrowserCapabilities {
|
|
/// Parse the HTTP Accept header to determine browser image format support.
|
|
pub fn from_accept_header(accept: Option<&str>) -> Self {
|
|
let accept = accept.unwrap_or("");
|
|
Self {
|
|
supports_webp: accept.contains("image/webp"),
|
|
supports_avif: accept.contains("image/avif"),
|
|
}
|
|
}
|
|
|
|
/// Get the best output format supported by the browser.
|
|
pub fn best_format(&self) -> Option<OutputFormat> {
|
|
if self.supports_webp {
|
|
Some(OutputFormat::WebP)
|
|
} else {
|
|
None
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Statistics about transcoding operations.
|
|
#[derive(Debug, Default, Clone)]
|
|
pub struct TranscodeStatsDto {
|
|
pub cache_hits: u64,
|
|
pub disk_hits: u64,
|
|
pub transcodes: u64,
|
|
pub bytes_saved: u64,
|
|
pub transcode_errors: u64,
|
|
}
|
|
|
|
/// Port for image transcoding operations.
|
|
///
|
|
/// Implementations handle the actual image conversion, caching,
|
|
/// and format detection, while the application layer only interacts
|
|
/// through this abstraction.
|
|
pub trait ImageTranscodePort: Send + Sync + 'static {
|
|
/// Check if a MIME type can be transcoded.
|
|
fn can_transcode(&self, mime_type: &str) -> bool;
|
|
|
|
/// Check if transcoding should be attempted based on file size and type.
|
|
fn should_transcode(&self, mime_type: &str, file_size: u64) -> bool;
|
|
|
|
/// Get a transcoded version of an image.
|
|
///
|
|
/// Returns `(content, mime_type, was_transcoded)`.
|
|
/// If transcoding is not beneficial (output larger than input), returns the
|
|
/// original content with `was_transcoded = false`.
|
|
///
|
|
/// `source_hash` is the BLAKE3 of the original content, which is how the
|
|
/// durable derived tier is keyed. `None` restricts the implementation to
|
|
/// its local cache — correct for callers with no hash (external mounts),
|
|
/// and the behaviour of every caller before that tier existed.
|
|
async fn get_transcoded(
|
|
&self,
|
|
file_id: &str,
|
|
source_hash: Option<&str>,
|
|
original_content: Bytes,
|
|
original_mime: &str,
|
|
target_format: OutputFormat,
|
|
) -> Result<(Bytes, String, bool), DomainError>;
|
|
|
|
/// Invalidate cached transcodes for a file.
|
|
async fn invalidate(&self, file_id: &str);
|
|
|
|
/// Get transcoding statistics.
|
|
async fn get_stats(&self) -> TranscodeStatsDto;
|
|
|
|
/// Clear all caches.
|
|
async fn clear_cache(&self) -> Result<(), DomainError>;
|
|
}
|