feat(jobs): jobs declare their own run parameters

`JobRunArgs` was a fixed struct — `force`, `deep`, `storage`, `repair` —
and six places hardcoded that same list: the engine's persist/restore,
the trigger endpoint's query type, the OXICLOUD_STARTUP_JOBS parser, the
frontend API wrapper, the panel's checkboxes, and `StartupTrigger` on
the wire.

Two costs. Adding a parameter meant editing all six, and forgetting one
dropped it silently — most damagingly in persist/restore, where a
resumed run lost it and a `?repair=true` migration came back as
discovery-only after a restart. And the panel offered the same knobs on
every job: only two jobs read `deep`, six read `repair`, so most of
those controls did nothing with no way to tell which.

Now `JobHandler::parameters()` returns `&'static [JobParam]` — name,
type (boolean/string/number), default, and the job's own description of
what it does. `JobRunArgs` holds a map keyed by those names.

Everything reads the declaration:

* `run_or_resume` iterates it to persist and restore, replacing
  `const FLAGS` plus a `storage` special case. `storage` stops being
  special — it was the one Option<String> among three bools.
* `dispatch` normalises every run against it, which is what makes "a
  handler sees its declared parameters with their declared defaults"
  true rather than usual. The periodic tick passes an empty
  `JobRunArgs::default()`, so a `default: true` parameter would
  otherwise read false on every scheduled run.
* The trigger endpoint takes free-form query params and rejects
  undeclared ones with a 400 naming the real set, instead of ignoring
  them.
* OXICLOUD_STARTUP_JOBS keeps raw pairs (config is parsed before the
  registry exists) and validates at dispatch, where the error can name
  the job's actual parameters. Still a boot panic, same as an unknown
  job name — a typo'd `?repare=true` must not leave a migration
  importing forever in discovery mode.
* `JobSummary.parameters` carries it to the panel, whose `supportsDeep`
  was a hardcoded name allowlist (`consistency_batch ||
  backend_consistency`). A job gaining a deep mode needed a frontend
  release; one losing it left a button that silently did nothing. The
  menu now renders from the declaration, so a newly-declared boolean
  appears with no frontend change.

Three consistency tenants were hand-rolling persist-on-fresh /
restore-on-resume for their own flag, under the same `params` key the
engine already used. Deleted — they read `args.get_bool(…)` now.

Fresh runs also filter to the declaration. `consistency_batch` forwards
its args verbatim to sub-jobs, so a tenant's `params` row could grow
`deep` with no deep mode, and the run-detail view would claim a mode the
job never had.

Two things found while wiring it, both worth knowing:

`RecoverableAdapter` bridges the two traits, and `parameters` has to be
forwarded there or the registry sees `&[]`. Both traits have defaults,
so omitting it compiled cleanly — and the trigger endpoint then rejected
`?repair=true` on the very jobs that declare it, with
OXICLOUD_STARTUP_JOBS panicking at boot. Now covered by
`adapter_forwards_job_metadata_from_inner_handler`.

`TriggerJobQuery` was briefly a newtype over the map. `serde_urlencoded`
cannot deserialize a newtype struct at the top level, so axum's `Query`
rejected EVERY trigger with a 400 — even one with no query string —
before the handler ran. It reads exactly like the new validation
rejecting something, which sent the first diagnosis to the wrong layer.
Now covered by `trigger_query_extracts_from_every_url_shape`.

Wire names are a compatibility surface: `params` rows are keyed by them
and the panel switches on them, so a rename breaks existing run history
the same way renaming a `Mutates` variant does. The JSON shape is pinned
in `snapshot_carries_job_metadata`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-09-07 12:30:40 +02:00
parent ff286f8159
commit a4101743e0
23 changed files with 1242 additions and 397 deletions
@@ -199,6 +199,25 @@ impl RecoverableJobHandler for BackendConsistencyCheck {
deleted."
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[
JobParam::boolean(
"deep",
false,
"Read every matched blob back and re-hash it, catching \
silent bit-rot. A full read of storage — can take hours.",
),
JobParam::string(
"storage",
"Name of the storage entry to audit. Absent audits the \
active backend; naming an entry is how either side of a \
migration gets audited directly.",
),
];
PARAMS
}
/// Approximate total: on a healthy install every backend blob
/// has a `storage.blobs` row, so the DB count is a proxy for
/// the backend count. The fraction deviating from 1.0 at run
@@ -240,26 +259,40 @@ impl RecoverableJobHandler for BackendConsistencyCheck {
// `blobs_consistency` uses — Fresh + args.storage=Some stamps
// probed_storage into params; Resumed reads it back so a
// mid-audit restart re-uses the same target.
let is_fresh = resume_cursor.is_none();
let probed_storage: Option<String> = if is_fresh {
let name = args.storage.clone();
if let Some(n) = &name
&& let Err(e) = store.set_string_param(PROBED_STORAGE_PARAM, n).await
{
return RunOutcome::Failed {
message: format!("persist {PROBED_STORAGE_PARAM} to params: {e}"),
};
}
name
} else {
match store.get_string_param(PROBED_STORAGE_PARAM).await {
Ok(v) => v,
Err(e) => {
return RunOutcome::Failed {
message: format!("read {PROBED_STORAGE_PARAM} from params: {e}"),
};
// `run_or_resume` persists and restores `storage` for us now, so
// the normal path is a plain read.
//
// The fallback is a MIGRATION concern, not defensiveness. This job
// used to persist the same value under its own
// `probed_storage` key; a run paused before this change has that
// key and no `storage` one. Without the fallback such a run would
// resume against the ACTIVE backend instead of the entry it was
// auditing — silently auditing the wrong thing, which is worse
// than failing. Removable once no pre-upgrade paused runs remain.
let probed_storage: Option<String> = match args.get_str("storage") {
Some(name) => Some(name.to_string()),
None if resume_cursor.is_some() => {
match store.get_string_param(PROBED_STORAGE_PARAM).await {
Ok(legacy) => {
if legacy.is_some() {
tracing::info!(
target: "oxicloud::consistency",
event = "backend_consistency.legacy_storage_param",
run_id = %store.run_id(),
"resumed a run that recorded its target under the pre-declaration \
`probed_storage` key"
);
}
legacy
}
Err(e) => {
return RunOutcome::Failed {
message: format!("read {PROBED_STORAGE_PARAM} from params: {e}"),
};
}
}
}
None => None,
};
let backend: Arc<dyn BlobStorageBackend> = match &probed_storage {
None => self.backend.clone(),
@@ -313,30 +346,12 @@ impl RecoverableJobHandler for BackendConsistencyCheck {
// that tenant to carry a backend for one flag, which is the
// overlap this split removes.
//
// Persisted to `params.deep` on a Fresh run so a Resume picks up
// the same mode (a Paused deep scan must not silently continue
// shallow) and the admin run-detail view can show what the scan
// actually verified. Written BEFORE the walk so a crash mid-batch
// still leaves the marker.
let deep = if is_fresh {
let v = if args.deep { "true" } else { "false" };
if let Err(e) = store.set_string_param("deep", v).await {
return RunOutcome::Failed {
message: format!("failed to persist deep flag to params: {e}"),
};
}
args.deep
} else {
match store.get_string_param("deep").await {
Ok(Some(v)) => v == "true",
Ok(None) => false,
Err(e) => {
return RunOutcome::Failed {
message: format!("read `deep` from params: {e}"),
};
}
}
};
// Persisted to `params.deep` and restored on resume by
// `run_or_resume`, so a Paused deep scan does not silently
// continue shallow and the run-detail view can show what the scan
// actually verified.
let deep = args.get_bool("deep");
if deep {
tracing::info!(
target: "oxicloud::consistency",
@@ -196,6 +196,21 @@ impl RecoverableJobHandler for BackendMigrationService {
restarting."
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
// Required in practice, though the declaration cannot express
// that: a Fresh run without it fails with a message naming the
// proper entrypoint, while a Resume legitimately omits it and
// reads the target back from `params.target_name`.
const PARAMS: &[JobParam] = &[JobParam::string(
"storage",
"Name of the storage entry to copy blobs INTO. Required on a \
fresh run; ignored on a resume, which reuses the recorded \
target.",
)];
PARAMS
}
/// Writes bytes to the target backend. Source bytes are left in place —
/// the copy is additive, so an aborted migration loses nothing.
fn mutates(&self) -> Mutates {
@@ -245,7 +260,7 @@ impl RecoverableJobHandler for BackendMigrationService {
// into the wrong entry.
let is_fresh = resume_cursor.is_none();
let target_name = if is_fresh {
let Some(name) = args.storage.clone() else {
let Some(name) = args.get_str("storage").map(str::to_string) else {
return RunOutcome::Failed {
message:
"backend_migration requires `target_name` on a fresh run — trigger via \
@@ -143,6 +143,17 @@ impl RecoverableJobHandler for BackendRotateService {
so re-running after a key change is cheap."
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::string(
"storage",
"Name of the storage entry whose blobs to rewrite. Required on \
a fresh run; ignored on a resume, which reuses the recorded \
target.",
)];
PARAMS
}
/// Rewrites blobs **in place**. Unlike a migration this has no additive
/// fallback — the previous ciphertext is gone once a blob is rewritten.
fn mutates(&self) -> Mutates {
@@ -178,7 +189,7 @@ impl RecoverableJobHandler for BackendRotateService {
// Resolve target entry name — same shape as `backend_migration`.
let is_fresh = resume_cursor.is_none();
let target_name = if is_fresh {
let Some(name) = args.storage.clone() else {
let Some(name) = args.get_str("storage").map(str::to_string) else {
return RunOutcome::Failed {
message: "backend_rotate requires `target_name` on a fresh run — trigger via \
POST /api/admin/storage/entries/{name}/rotate"
@@ -267,6 +267,21 @@ impl RecoverableJobHandler for BlobsConsistencyCheck {
)
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
// No `deep` — re-reading bytes is backend work and moved to
// `backend_consistency`. Declaring it here would put a knob in
// the panel that this job ignores, which is the thing the
// declaration exists to stop.
const PARAMS: &[JobParam] = &[JobParam::boolean(
"repair",
false,
"Rewrite drifted ref_count values to the recomputed truth. \
Without this the run only reports them.",
)];
PARAMS
}
/// Definitive count. `storage.blobs` PK scan is index-only;
/// even at millions of rows it's sub-second on modern PG.
async fn count_total(&self) -> Option<u64> {
@@ -299,12 +314,6 @@ impl RecoverableJobHandler for BlobsConsistencyCheck {
// `backend_consistency`, which finds it in one enumeration pass
// instead of one probe per row.
//
// Snapshot "is this a Fresh run?" BEFORE the resume_cursor
// match consumes it — otherwise the `is_none()` check later
// borrows a partially-moved value. Fresh = no cursor bytes
// at all; Resumed = cursor bytes present (possibly empty).
let is_fresh = resume_cursor.is_none();
// Cursor = the last-visited `hash` string, UTF-8-encoded. On
// resume, we walk `WHERE hash > $cursor` in ASC order. First
// batch: NULL cursor → start from the smallest hash.
@@ -340,30 +349,15 @@ impl RecoverableJobHandler for BlobsConsistencyCheck {
// matched key pairs worth verifying. A deep flag on this tenant
// would be a flag with nothing to do.
// Repair mode persisted to `params.repair` so the admin run-detail
// view can display it. Fresh persists what the trigger asked for;
// Resume reads back so a paused repair scan stays a repair
// scan (a mid-scan crash mustn't silently downgrade to
// discovery-only for the remaining rows).
let repair = if is_fresh {
let v = if args.repair { "true" } else { "false" };
if let Err(e) = store.set_string_param("repair", v).await {
return RunOutcome::Failed {
message: format!("failed to persist repair flag to params: {e}"),
};
}
args.repair
} else {
match store.get_string_param("repair").await {
Ok(Some(v)) => v == "true",
Ok(None) => false,
Err(e) => {
return RunOutcome::Failed {
message: format!("read `repair` from params: {e}"),
};
}
}
};
// Repair mode is persisted to `params.repair` so the admin
// run-detail view can display it, and restored on resume so a
// paused repair scan stays a repair scan — a mid-scan crash must
// not silently downgrade the remaining rows to discovery-only.
//
// Both happen in `run_or_resume`, for every declared parameter,
// under this same key. This job used to do it itself; that
// duplication is what the parameter declaration removes.
let repair = args.get_bool("repair");
if repair {
tracing::info!(
@@ -97,6 +97,32 @@ impl JobHandler for ConsistencyBatch {
are forwarded to each sub-job."
}
/// The union of what its sub-jobs accept, because it forwards
/// verbatim. A sub-job that does not declare one of these simply
/// never sees it — `run_or_resume` filters each dispatch down to that
/// job's own declaration, so forwarding `deep` to a tenant with no
/// deep mode is inert rather than misrecorded.
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[
JobParam::boolean(
"deep",
false,
"Forwarded to sub-jobs that have a deep mode — currently \
backend_consistency, which re-reads and re-hashes every \
blob. Can take hours.",
),
JobParam::boolean("force", false, "Forwarded to sub-jobs that accept it."),
JobParam::boolean(
"repair",
false,
"Forwarded to every sub-job that can repair, so one call \
fixes both refcount tenants.",
),
];
PARAMS
}
/// Read-only on a plain run because every tenant it dispatches is, but
/// `?repair=true` reaches whichever of them act on it — so the batch
/// inherits the strongest mode any sub-job can be put into.
@@ -198,9 +224,9 @@ impl JobHandler for ConsistencyBatch {
targets.len() as u64,
json!({
"per_check": per_check,
"deep": args.deep,
"force": args.force,
"repair": args.repair,
"deep": args.get_bool("deep"),
"force": args.get_bool("force"),
"repair": args.get_bool("repair"),
"ok": ok_count,
"err": err_count,
}),
+18 -3
View File
@@ -3873,18 +3873,33 @@ impl crate::infrastructure::scheduler::JobHandler for DedupService {
/// the freed disk. GC returning `(0, 0)` is normal — it means trash
/// cleanup already reaped everything.
///
/// `args.force = true` skips the orphan grace window
/// `force = true` skips the orphan grace window
/// (`garbage_collect_force` — grace_secs = 0). Same semantic as
/// `POST /api/admin/jobs/dedup_gc/trigger?force=true`. Unsafe
/// under concurrent uploads: only reachable through the admin
/// endpoint and only intentionally used by tests + operator
/// diagnostic sessions.
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
// A named `const` rather than a bare `&[…]` literal: implicit
// const promotion does not cover `const fn` calls, so the
// literal would be a temporary. Same shape in every job.
const PARAMS: &[JobParam] = &[JobParam::boolean(
"force",
false,
"Skip the orphan grace window. Unsafe under concurrent \
uploads — it reopens the TOCTOU window the grace closes.",
)];
PARAMS
}
async fn run(
&self,
args: &crate::infrastructure::scheduler::JobRunArgs,
) -> crate::infrastructure::scheduler::JobOutcome {
use crate::infrastructure::scheduler::JobOutcome;
let result = if args.force {
let force = args.get_bool("force");
let result = if force {
self.garbage_collect_force().await
} else {
self.garbage_collect().await
@@ -3892,7 +3907,7 @@ impl crate::infrastructure::scheduler::JobHandler for DedupService {
match result {
Ok((items, bytes)) => JobOutcome::ok_with(
items,
serde_json::json!({ "bytes_reclaimed": bytes, "forced": args.force }),
serde_json::json!({ "bytes_reclaimed": bytes, "forced": force }),
),
Err(e) => JobOutcome::err(format!("dedup GC failed: {e}")),
}
@@ -139,19 +139,31 @@ impl JobHandler for GrantCleanupService {
/// `extra.grace_days` records which grace was applied so admin
/// listings can see it without a second lookup.
///
/// `args.force = true` collapses the grace window to zero for
/// this run only — same semantic as
/// `force = true` collapses the grace window to zero for this run
/// only — same semantic as
/// `POST /api/admin/jobs/grant_cleanup/trigger?force=true`. The
/// configured `self.grace_days` is not mutated.
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"force",
false,
"Collapse the expiry grace window to zero for this run. \
The configured grace is not changed.",
)];
PARAMS
}
async fn run(&self, args: &JobRunArgs) -> JobOutcome {
let grace_override = if args.force { Some(0) } else { None };
let force = args.get_bool("force");
let grace_override = if force { Some(0) } else { None };
let effective_grace = grace_override.unwrap_or(self.grace_days);
match self.purge(grace_override).await {
Ok(count) => JobOutcome::ok_with(
count,
serde_json::json!({
"grace_days": effective_grace,
"forced": args.force,
"forced": force,
}),
),
Err(e) => JobOutcome::err(format!("grant cleanup failed: {e}")),
@@ -202,6 +202,17 @@ impl RecoverableJobHandler for ManifestsConsistencyCheck {
)
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"repair",
false,
"Rewrite drifted manifest ref_count values to the recomputed \
truth. Without this the run only reports them.",
)];
PARAMS
}
async fn count_total(&self) -> Option<u64> {
let row: Result<(i64,), sqlx::Error> =
sqlx::query_as("SELECT COUNT(*) FROM storage.chunk_manifests")
@@ -227,8 +238,6 @@ impl RecoverableJobHandler for ManifestsConsistencyCheck {
args: &JobRunArgs,
resume_cursor: Option<Vec<u8>>,
) -> RunOutcome {
let is_fresh = resume_cursor.is_none();
// Cursor: the last `file_hash` as UTF-8. Same convention as
// `blobs_consistency`, which also pages a hash-keyed table.
let mut cursor: Option<String> = match resume_cursor {
@@ -244,33 +253,14 @@ impl RecoverableJobHandler for ManifestsConsistencyCheck {
},
};
// Persist the repair flag into `params.repair` so the admin
// run-detail view can display whether the run was a discovery
// scan or an active repair. Fresh takes it from args; Resume
// reads back so a paused repair scan stays a repair scan (a
// mid-scan crash mustn't silently downgrade the remaining
// rows to discovery-only). Same shape as
// `blobs_consistency_service.rs`'s `deep` handling — see the
// reasoning documented there.
let repair = if is_fresh {
let v = if args.repair { "true" } else { "false" };
if let Err(e) = store.set_string_param("repair", v).await {
return RunOutcome::Failed {
message: format!("failed to persist repair flag to params: {e}"),
};
}
args.repair
} else {
match store.get_string_param("repair").await {
Ok(Some(v)) => v == "true",
Ok(None) => false,
Err(e) => {
return RunOutcome::Failed {
message: format!("read `repair` from params: {e}"),
};
}
}
};
// Persisted into `params.repair` so the admin run-detail view can
// show whether this was a discovery scan or an active repair, and
// restored on resume so a paused repair scan stays one — a
// mid-scan crash must not silently downgrade the remaining rows.
//
// `run_or_resume` does both, for every declared parameter, under
// this same key. This job used to hand-roll it.
let repair = args.get_bool("repair");
if repair {
tracing::info!(
@@ -185,6 +185,17 @@ impl RecoverableJobHandler for ThumbAttachedImport {
)
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"repair",
false,
"Delete each sidecar after its replacement has been read back. \
Without this the job imports and leaves the originals in place.",
)];
PARAMS
}
async fn count_total(&self) -> Option<u64> {
let mut total = 0u64;
for size in ThumbnailSize::all() {
@@ -229,7 +240,7 @@ impl RecoverableJobHandler for ThumbAttachedImport {
// PDF preview has no server-side render path — so it is not
// belt-and-braces, it is the only thing between a migration and
// permanent loss.
let delete_imported = args.repair;
let delete_imported = args.get_bool("repair");
let mut failed = 0u64;
let mut since_checkpoint = 0usize;
@@ -424,6 +424,18 @@ impl RecoverableJobHandler for ThumbDerivedImport {
)
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"repair",
false,
"Delete each sidecar after its replacement has been read back, \
and remove the directory once empty. Without this the job \
imports and leaves the originals in place.",
)];
PARAMS
}
async fn count_total(&self) -> Option<u64> {
let mut total = 0u64;
for size in ThumbnailSize::all() {
@@ -449,7 +461,7 @@ impl RecoverableJobHandler for ThumbDerivedImport {
// makes the migration self-draining: sidecars are LOCAL disk, so no
// release can know whether every instance has finished, whereas each
// instance draining itself needs no coordination at all.
let delete_imported = args.repair;
let delete_imported = args.get_bool("repair");
// Cursor is `{size_dir}/{filename}` — the last file completed. Sizes
// are walked in `ThumbnailSize::all()` order, and names are sorted
// within each, so the pair totally orders the walk.
@@ -214,6 +214,18 @@ impl RecoverableJobHandler for TranscodeImport {
)
}
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"repair",
false,
"Delete each cached transcode after its replacement has been \
read back and compared byte for byte. Without this the job \
imports and leaves the originals in place.",
)];
PARAMS
}
async fn count_total(&self) -> Option<u64> {
Some(Self::entry_names(&self.variant_dir()).await.len() as u64)
}
@@ -240,7 +252,7 @@ impl RecoverableJobHandler for TranscodeImport {
},
};
let delete_imported = args.repair;
let delete_imported = args.get_bool("repair");
let dir = self.variant_dir();
let mut imported = 0u64;