Files
Oxicloud/src/infrastructure/services/grant_cleanup_service.rs
T
Edouard Vanbelle a4101743e0 feat(jobs): jobs declare their own run parameters
`JobRunArgs` was a fixed struct — `force`, `deep`, `storage`, `repair` —
and six places hardcoded that same list: the engine's persist/restore,
the trigger endpoint's query type, the OXICLOUD_STARTUP_JOBS parser, the
frontend API wrapper, the panel's checkboxes, and `StartupTrigger` on
the wire.

Two costs. Adding a parameter meant editing all six, and forgetting one
dropped it silently — most damagingly in persist/restore, where a
resumed run lost it and a `?repair=true` migration came back as
discovery-only after a restart. And the panel offered the same knobs on
every job: only two jobs read `deep`, six read `repair`, so most of
those controls did nothing with no way to tell which.

Now `JobHandler::parameters()` returns `&'static [JobParam]` — name,
type (boolean/string/number), default, and the job's own description of
what it does. `JobRunArgs` holds a map keyed by those names.

Everything reads the declaration:

* `run_or_resume` iterates it to persist and restore, replacing
  `const FLAGS` plus a `storage` special case. `storage` stops being
  special — it was the one Option<String> among three bools.
* `dispatch` normalises every run against it, which is what makes "a
  handler sees its declared parameters with their declared defaults"
  true rather than usual. The periodic tick passes an empty
  `JobRunArgs::default()`, so a `default: true` parameter would
  otherwise read false on every scheduled run.
* The trigger endpoint takes free-form query params and rejects
  undeclared ones with a 400 naming the real set, instead of ignoring
  them.
* OXICLOUD_STARTUP_JOBS keeps raw pairs (config is parsed before the
  registry exists) and validates at dispatch, where the error can name
  the job's actual parameters. Still a boot panic, same as an unknown
  job name — a typo'd `?repare=true` must not leave a migration
  importing forever in discovery mode.
* `JobSummary.parameters` carries it to the panel, whose `supportsDeep`
  was a hardcoded name allowlist (`consistency_batch ||
  backend_consistency`). A job gaining a deep mode needed a frontend
  release; one losing it left a button that silently did nothing. The
  menu now renders from the declaration, so a newly-declared boolean
  appears with no frontend change.

Three consistency tenants were hand-rolling persist-on-fresh /
restore-on-resume for their own flag, under the same `params` key the
engine already used. Deleted — they read `args.get_bool(…)` now.

Fresh runs also filter to the declaration. `consistency_batch` forwards
its args verbatim to sub-jobs, so a tenant's `params` row could grow
`deep` with no deep mode, and the run-detail view would claim a mode the
job never had.

Two things found while wiring it, both worth knowing:

`RecoverableAdapter` bridges the two traits, and `parameters` has to be
forwarded there or the registry sees `&[]`. Both traits have defaults,
so omitting it compiled cleanly — and the trigger endpoint then rejected
`?repair=true` on the very jobs that declare it, with
OXICLOUD_STARTUP_JOBS panicking at boot. Now covered by
`adapter_forwards_job_metadata_from_inner_handler`.

`TriggerJobQuery` was briefly a newtype over the map. `serde_urlencoded`
cannot deserialize a newtype struct at the top level, so axum's `Query`
rejected EVERY trigger with a 400 — even one with no query string —
before the handler ran. It reads exactly like the new validation
rejecting something, which sent the first diagnosis to the wrong layer.
Now covered by `trigger_query_extracts_from_every_url_shape`.

Wire names are a compatibility surface: `params` rows are keyed by them
and the panel switches on them, so a rename breaks existing run history
the same way renaming a `Mutates` variant does. The JSON shape is pinned
in `snapshot_carries_job_metadata`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:23:30 +02:00

173 lines
6.7 KiB
Rust

//! Service that purges expired `storage.role_grants` rows.
//!
//! The AuthZ engine already filters expired grants out of every
//! permission check at read time (`expires_at IS NULL OR
//! expires_at > NOW()` on every `check` / `list_grants_*` path in
//! `PgAclEngine`), so expired rows never leak permission. They just
//! accumulate. This service garbage-collects them, with a grace window
//! past `expires_at` that preserves the audit / support answer to
//! "what happened to my access?" for a few weeks.
//!
//! **Scheduling.** Registered with the periodic-job scheduler
//! (`docs/plan/job-registry.md` Part 1). The retired `start_cleanup_job`
//! used to spawn its own `tokio::interval` loop; the scheduler now
//! dispatches [`GrantCleanupService::purge`] on the configured cadence
//! and handles panic containment + exclusivity + admin trigger routing.
//! Admin trigger with `?force=true` still bypasses the registered
//! job and calls `purge(Some(0))` directly so the grace override reaches
//! the underlying SQL.
use std::sync::Arc;
use std::time::{Duration, Instant};
use tracing::{error, info};
use crate::application::ports::authorization_ports::AuthorizationEngine;
use crate::common::errors::DomainError;
use crate::infrastructure::scheduler::{JobHandler, JobOutcome, JobRegistry, JobRunArgs, Mutates};
use crate::infrastructure::services::pg_acl_engine::PgAclEngine;
use async_trait::async_trait;
pub const GRANT_CLEANUP_JOB_NAME: &str = "grant_cleanup";
/// Service that deletes expired grants.
///
/// Owns an `Arc<PgAclEngine>` (not a `dyn AuthorizationEngine`) to avoid
/// the wrapper allocation on every SQL call — the caller set is small
/// (scheduler tick + admin trigger endpoint), both statically dispatched.
pub struct GrantCleanupService {
authz: Arc<PgAclEngine>,
grace_days: u32,
interval_hours: u64,
}
impl GrantCleanupService {
pub fn new(authz: Arc<PgAclEngine>, grace_days: u32, interval_hours: u64) -> Self {
Self {
authz,
grace_days,
// Minimum 1 hour — matches TrashCleanupService's clamp so
// a mis-set `0` doesn't spin a hot loop.
interval_hours: interval_hours.max(1),
}
}
/// Grace period the service uses on its scheduled runs. Exposed
/// for the admin trigger's default-response field.
pub fn grace_days(&self) -> u32 {
self.grace_days
}
/// Cadence exposed as `Duration`. Internal helper used by
/// [`Self::register`]; kept `pub` for tests.
pub fn interval(&self) -> Duration {
Duration::from_secs(self.interval_hours * 3600)
}
/// Register self with the periodic-job scheduler and return the
/// same `Arc<Self>` for DI-style chaining. Scheduled tenant with
/// interval = `self.interval()`, no timeout. See
/// `docs/plan/job-registry.md` Part 1.
pub async fn register(self: Arc<Self>, registry: &JobRegistry) -> Arc<Self> {
let interval = self.interval();
registry.register(self.clone(), Some(interval), None).await;
self
}
/// Run one purge pass.
///
/// `grace_override`:
/// - `None` → use the configured grace (`self.grace_days`).
/// - `Some(n)` → override with `n`. The admin `?force=true` trigger
/// passes `Some(0)` so Hurl regressions can hit expired grants
/// without waiting the configured grace out.
///
/// Returns `Ok(count)` on success, `Err(_)` on DB error. Audit-log
/// lines fire on both paths (success + failure) — bulk deletion of
/// authorization rows is security-relevant enough to log even a
/// zero-count run, and failures MUST reach the audit channel.
pub async fn purge(&self, grace_override: Option<u32>) -> Result<u64, DomainError> {
let grace = grace_override.unwrap_or(self.grace_days);
let start = Instant::now();
match self.authz.purge_expired_grants(grace).await {
Ok(count) => {
info!(
target: "audit",
event = "grant_cleanup.purged",
count = count,
grace_days = grace,
elapsed_ms = start.elapsed().as_millis() as u64,
"👮🏻‍♂️ Purged {} expired grant(s) older than {} days",
count,
grace,
);
Ok(count)
}
Err(e) => {
error!(
target: "audit",
event = "grant_cleanup.failed",
grace_days = grace,
error = %e,
"Grant cleanup failed"
);
Err(e)
}
}
}
}
#[async_trait]
impl JobHandler for GrantCleanupService {
fn name(&self) -> &str {
GRANT_CLEANUP_JOB_NAME
}
fn description(&self) -> &'static str {
"Deletes expired role grants once they are past the retention \
window. Expired grants never leak permission — every AuthZ check \
filters on expires_at — they just accumulate. The window keeps \
'what happened to my access?' answerable for a few weeks after \
expiry."
}
fn mutates(&self) -> Mutates {
Mutates::Always
}
/// Runs one purge. `count` on the returned `JobOutcome::Ok` is
/// the number of `role_grants` rows physically deleted;
/// `extra.grace_days` records which grace was applied so admin
/// listings can see it without a second lookup.
///
/// `force = true` collapses the grace window to zero for this run
/// only — same semantic as
/// `POST /api/admin/jobs/grant_cleanup/trigger?force=true`. The
/// configured `self.grace_days` is not mutated.
fn parameters(&self) -> &'static [crate::infrastructure::scheduler::JobParam] {
use crate::infrastructure::scheduler::JobParam;
const PARAMS: &[JobParam] = &[JobParam::boolean(
"force",
false,
"Collapse the expiry grace window to zero for this run. \
The configured grace is not changed.",
)];
PARAMS
}
async fn run(&self, args: &JobRunArgs) -> JobOutcome {
let force = args.get_bool("force");
let grace_override = if force { Some(0) } else { None };
let effective_grace = grace_override.unwrap_or(self.grace_days);
match self.purge(grace_override).await {
Ok(count) => JobOutcome::ok_with(
count,
serde_json::json!({
"grace_days": effective_grace,
"forced": force,
}),
),
Err(e) => JobOutcome::err(format!("grant cleanup failed: {e}")),
}
}
}