4d067b3fc694b153d1d9fc11ca76f89620bfac02
12 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ca85ac7307 |
fix(file_attached): doe not increment ref_count if new attachement has same hash
- do not increment ref_count if new attachement to a file with same data - add audit log to help identifying other future issue in ref_count - prevent race condition while attaching a blob |
||
|
|
a4101743e0 |
feat(jobs): jobs declare their own run parameters
`JobRunArgs` was a fixed struct — `force`, `deep`, `storage`, `repair` — and six places hardcoded that same list: the engine's persist/restore, the trigger endpoint's query type, the OXICLOUD_STARTUP_JOBS parser, the frontend API wrapper, the panel's checkboxes, and `StartupTrigger` on the wire. Two costs. Adding a parameter meant editing all six, and forgetting one dropped it silently — most damagingly in persist/restore, where a resumed run lost it and a `?repair=true` migration came back as discovery-only after a restart. And the panel offered the same knobs on every job: only two jobs read `deep`, six read `repair`, so most of those controls did nothing with no way to tell which. Now `JobHandler::parameters()` returns `&'static [JobParam]` — name, type (boolean/string/number), default, and the job's own description of what it does. `JobRunArgs` holds a map keyed by those names. Everything reads the declaration: * `run_or_resume` iterates it to persist and restore, replacing `const FLAGS` plus a `storage` special case. `storage` stops being special — it was the one Option<String> among three bools. * `dispatch` normalises every run against it, which is what makes "a handler sees its declared parameters with their declared defaults" true rather than usual. The periodic tick passes an empty `JobRunArgs::default()`, so a `default: true` parameter would otherwise read false on every scheduled run. * The trigger endpoint takes free-form query params and rejects undeclared ones with a 400 naming the real set, instead of ignoring them. * OXICLOUD_STARTUP_JOBS keeps raw pairs (config is parsed before the registry exists) and validates at dispatch, where the error can name the job's actual parameters. Still a boot panic, same as an unknown job name — a typo'd `?repare=true` must not leave a migration importing forever in discovery mode. * `JobSummary.parameters` carries it to the panel, whose `supportsDeep` was a hardcoded name allowlist (`consistency_batch || backend_consistency`). A job gaining a deep mode needed a frontend release; one losing it left a button that silently did nothing. The menu now renders from the declaration, so a newly-declared boolean appears with no frontend change. Three consistency tenants were hand-rolling persist-on-fresh / restore-on-resume for their own flag, under the same `params` key the engine already used. Deleted — they read `args.get_bool(…)` now. Fresh runs also filter to the declaration. `consistency_batch` forwards its args verbatim to sub-jobs, so a tenant's `params` row could grow `deep` with no deep mode, and the run-detail view would claim a mode the job never had. Two things found while wiring it, both worth knowing: `RecoverableAdapter` bridges the two traits, and `parameters` has to be forwarded there or the registry sees `&[]`. Both traits have defaults, so omitting it compiled cleanly — and the trigger endpoint then rejected `?repair=true` on the very jobs that declare it, with OXICLOUD_STARTUP_JOBS panicking at boot. Now covered by `adapter_forwards_job_metadata_from_inner_handler`. `TriggerJobQuery` was briefly a newtype over the map. `serde_urlencoded` cannot deserialize a newtype struct at the top level, so axum's `Query` rejected EVERY trigger with a 400 — even one with no query string — before the handler ran. It reads exactly like the new validation rejecting something, which sent the first diagnosis to the wrong layer. Now covered by `trigger_query_extracts_from_every_url_shape`. Wire names are a compatibility surface: `params` rows are keyed by them and the panel switches on them, so a rename breaks existing run history the same way renaming a `Mutates` variant does. The JSON shape is pinned in `snapshot_carries_job_metadata`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
10c362a94a |
fix(jobs): flush the checkpoint tail, so progress reflects reality
All three import jobs only checkpointed on a full batch, so the remainder after the last one was never counted. A run shorter than BATCH_SIZE never checkpointed at all: `scanned_count` stayed 0 against a known `total_rows`, and the admin progress bar sat at zero for the whole run and finished there. Seen on a transcode_import run over 20 entries — 13 imported, 5 negatives, 2 already present, progress 0/20 throughout. The thumbnail imports had it too, just less visibly: a 105-file run reported `scanned_count: 100`, losing the tail rather than all of it. Cursor-wise the final checkpoint is a no-op — the walk is finished, so nothing resumes from it — but the scanned delta is what the progress display reads, and it has to include the last partial batch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ce4354f497 |
fix(thumbnails): neither import job may tear down the shared directory
Found on a sandbox restore. `thumb_derived_import` ran first, imported and deleted its own hash-named sidecars, then found `remove_dir` refused because the `ext-*.jpg` previews were still there — those belong to `thumb_attached_import`. The rename fallback fired, moving the tree to `.thumbnails.migrated`; the attached job then looked in `.thumbnails/`, found nothing, and reported zeros. That stranded the user-uploaded previews, which are the one class of file here with no render path to rebuild them. The rename exists for files NEITHER job claims — a `.DS_Store` blocking removal forever — and it fired for the sibling's work in progress instead. Inverting the job order does not help: once the tree is renamed, both jobs look at `.thumbnails/` and find nothing, whatever order they run in. Teardown is now shared and refuses to act while anything remains that either job would claim. Both jobs call it, so whichever finishes last removes the tree in the same boot rather than leaving an empty directory until the next one. The rename survives for its original purpose, and now only fires when the remaining files are genuinely nobody's. Also drops the daily tick on both imports — they are on-demand now. The boot run in repair mode IS the migration: nothing has written a sidecar since step 10d2, so the tail cannot grow afterwards, and a tick could not finish the job anyway because ticks never pass `repair`. Once drained it was a `read_dir` returning nothing, every day, forever. UX: the "at boot" badge moves from beside the job name into the cadence column. It answers WHEN a job runs, which is what that column is for — next to the name it read as a property of the job, and the row could show "on-demand" beside a badge saying otherwise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1ea3826660 |
feat(jobs): jobs describe themselves — description, mutates, repair_description
The admin panel had no repair toggle wired to anything but a hardcoded
name list naming the two refcount tenants, so `thumb_derived_import` and
`thumb_attached_import` could not be run in repair mode from the UI at
all despite supporting it. And nothing in the job list said what any
given job does or whether clicking Run on production writes anything.
Three defaulted methods on `JobHandler` and `RecoverableJobHandler`:
fn description(&self) -> &'static str
fn mutates(&self) -> Mutates // Never | Always | OnRepairOnly
fn repair_description(&self) -> Option<&'static str>
`RecoverableAdapter` forwards them — the registry only holds
`dyn JobHandler`, so a tenant's metadata is invisible otherwise, and
falling back to the defaults would report every recoverable job as
read-only, including the ones that delete files.
Three values rather than a boolean because a job can be read-only by
default and destructive under `?repair=true`; a boolean answers wrongly
for one of its two modes, and `false` on something that unlinks files is
the dangerous direction to be wrong in. `repair_description` returning
`Option` collapses "does it repair" and "what does repair do" into one
method: presence gates the toggle, content is the confirmation text —
which the frontend cannot invent, since correcting a counter and
deleting sidecars are not the same warning.
`OnRepairOnly` with no `repair_description` is rejected at registration:
it claims to mutate only under a flag it does not support.
All 17 registered jobs declare all three. The panel now renders the
description under each name, badges read-only jobs, confirms before a
plain run of a mutating one, and offers the repair variant off the
backend flag instead of the name list.
Descriptions are English in the trait, next to the behaviour: one in
`locales/*.json` rots invisibly the moment a job changes, and a
translator cannot know what `manifests_consistency` reconciles. i18n can
layer on later keyed by job name with these as the fallback.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b485db46fa |
feat(storage): audit every sidecar deletion, and reclaim orphaned uploads
Two changes to the import jobs' destructive path. thumb_attached_import now deletes orphaned sidecars under `repair`, matching the dead-source case on the derived side. An `ext-` file whose owner is gone is unimportable — the FK on file_id would reject the row — so leaving it means it is rediscovered every run, the tail never empties and step 10e's gate never opens. Safe despite these being the non-regenerable bytes: the preview is keyed to a file_id that no longer exists, so nothing can reference it again. Unrecoverable and unreachable are different things, and this is both. And every deletion is now audited. A one-way migration removing user-visible files should leave a trail that outlives the run history: findings are per-run and get purged, whereas target: "audit" is separable and retained. If a preview later turns out to be missing, this is the only record saying the migration removed it and when. `owner` carries the id the file belonged to — source_hash for content-keyed, file_id for uploaded — because that is where an investigation starts, and the raw logs cannot supply it: NEW BLOB names the hash of the STORED BYTES, a different value from the sidecar's own name, which is why grepping one against the other finds nothing. reason is a stable key: `imported` (replaced by a verified blob), `source_gone`, `orphaned`. The first lives inside verify_and_unlink so a verified deletion cannot be logged inconsistently; the other two are explicit, since those paths have nothing to verify against. |
||
|
|
1b68ee093e |
feat(storage): both import jobs tick daily instead of manual-only
Registered with interval None, so they ran only when someone remembered to trigger them — which was your objection to gating anything on operator timing. Now daily. Not boot-time: that would delay readiness for a filesystem walk, and both jobs are idempotent and resumable, so periodic is strictly better. The tick deliberately does NOT delete. `repair` defaults false, so scheduled runs import and stop; unlinking stays a deliberate operator action, per no-silent-auto-repair. That splits the two halves the way their risk differs — the backfill is safe to automate, removing files is not. Cost once drained is a read_dir over three directories returning nothing, and after the directory itself is removed, not even that. |
||
|
|
df619a7ed9 |
feat(storage): thumb_attached_import drains its sidecars too
Completes step 10d. Same `?repair=true` opt-in and the same readback-before-unlink as the derived half, and the check matters more here: these sidecars hold the bytes that CANNOT be regenerated — a client-uploaded PDF preview has no server-side render path — so the read is the only thing between a migration and permanent loss, not belt-and-braces. verify_and_unlink is shared rather than copied. Two versions of "only delete after proving the replacement is readable" would be two chances to weaken one, and it is the rule the whole deletion step rests on. Deletion covers the already-imported branch as well as fresh imports, for the same reason as the derived job: a run without `repair` leaves the sidecar behind, and a later run with it would otherwise see "already imported" and never drain. Import first, enable deletion after, is the expected operator sequence, so that branch is the common path. Orphaned sidecars stay untouched — this job imports, it does not reclaim, and a destructive default on a migration is what no-silent- auto-repair forbids. Unverifiable ones are kept and reported, so the next run retries. With both halves draining, `.thumbnails/` can now actually empty and the directory removal in the derived job can succeed — though not stably until dual-write stops, since any render or upload recreates it. |
||
|
|
12158ccf59 |
fix(storage): thumb_derived_import claims JPEG sidecars too
The filter was strip_suffix(".webp"), but persist_rendered writes
{hash}.{format.ext()} — so any client not advertising WebP leaves
{hash}.jpg on disk. Correct only while the derived tier was WebP-only;
once variant carried the format (20261022000000) a JPEG sidecar became
ordinary content, and leaving it unclaimed would keep .thumbnails/
permanently non-empty — the very signal step 10e gates on. The migration
could never finish.
Both codecs are now claimed and the format comes from the file's own
extension, so a .jpg imports AS JPEG. Deriving the variant and
content_type from it rather than hardcoding WebP is the point: a
mislabelled row would serve the wrong codec to whoever the read path
then matched it for.
ThumbnailFormat::ALL exists so the claim list and the write path cannot
drift — adding a format without teaching the import about it would
strand that codec silently.
The `ext-` rejection now carries real weight. Previously .jpg was
rejected wholesale, so the two jobs could not overlap by construction;
now they share an extension and only the prefix separates them. Both
directions stay under test.
Caught by the cross-job assertion, which counts every real sidecar being
claimed exactly once — the fixture gained a .jpg and the total moved 3
to 4, which is the test noticing rather than a test to update.
|
||
|
|
395296a7e7 |
fix(storage): file_exists misreported every file as missing
`SELECT 1 FROM storage.files WHERE id = $1` decoded as i64. PostgreSQL types a bare `1` as int4, so the decode always failed — and since `.ok().flatten()` turns a decode error into the same None as "no row", file_exists reported false for every file. thumb_attached_import therefore classified every sidecar as an orphan and imported nothing. Caught by thumb_import_check.sh on its first run: the derived import restored its rows, the attached one restored none. Now `SELECT EXISTS(...)`, which yields a real bool and always returns exactly one row, so absence means absence. A query error still degrades to false — the safe direction, leaving the file on disk as a reported orphan rather than importing it against a row that may not exist. The failure mode is the point, and it is the third of this shape in two days: an error converted into an innocuous-looking outcome. So the check script now dumps a job's findings when an assertion fails. The jobs already recorded exactly why they skipped each file — the attached_sidecar_orphan findings naming the cause were sitting in the run while the script reported only "did not restore the row", which is indistinguishable from the job never having run. |
||
|
|
49e7bb15a6 |
test(storage): cover the sidecar walk for both import jobs
Both imports run over ONE directory, where the two legacy shapes sit side by side, so the property worth asserting spans them: together they must claim every real sidecar exactly once, and neither may take the other's. A job that drifted into the other's shape would content-key user-supplied bytes — sharing one user's uploaded preview onto every file with identical content — and no per-job test in isolation would notice. So the fixture is shared. `legacy_tree` builds a directory holding a content-keyed .webp pair, an ext- upload, and a stray README, and both test modules walk it: derived claims exactly the two hashes in sorted order, attached claims exactly the ext- file, the two sets are disjoint, and between them they account for all three real sidecars. `sidecar_names` became an associated function taking the root instead of reading `self`, which is what makes this testable at all — the walk is the half that decides which files a job claims, and it needed no pool to verify. Sorting is asserted rather than assumed, since the cursor resumes by skipping everything at or before it and a stable order is the only thing that makes that correct. A missing size directory is covered too: normal on a fresh install, and it must yield no work rather than abort the walk. Not covered here, and it needs a pooled fixture that does not exist: the round trip itself — store the blob, write the row, and confirm a COPY inherits the preview. That belongs in the API-level harness, where the legacy state can be manufactured through the real write path and then stripped. |
||
|
|
7a2ebe0fdc |
feat(storage): thumb_attached_import — backfill uploaded previews
Twin of thumb_derived_import, for the other sidecar shape:
{thumbnails_root}/{size}/ext-{file_id}.jpg, the previews a user
uploaded — notably the SPA's client-side PDF generator, which has no
server-side render path at all.
Until a row exists, a copy of the file LOSES the preview: the sidecar is
keyed by file_id, no copy path duplicates it, and the server silently
falls back to rendering from the source, or to nothing for a PDF. That
is the bug file_attached_blobs closed for new uploads; this closes it
for everything already on disk.
Separate job rather than an arm of the derived import, because the
keying differs and that difference is the security boundary. These bytes
are not derivable from the file's content, so content-keying them would
share one user's uploaded preview onto every file with identical
content. Each job's name filter rejects the other's shape, and both
directions are under test.
Idempotence needs more care here than in the derived twin.
store_attached_blob is ON CONFLICT DO UPDATE, so calling it for an
existing row releases the previous reference and takes a new one —
harmless once, but a job doing it every run would churn refcounts. The
row is therefore checked first and the store reached only on a genuine
insert.
uploaded_by is the nil sentinel: disk records no uploader, and inventing
one — the file's created_by, say — would fabricate provenance that could
later read as evidence an Editor replaced someone's preview. The column
is NOT NULL with no FK precisely so provenance survives, and a sentinel
says "unknown" honestly.
Orphaned sidecars (no storage.files row) are counted and reported, not
deleted. This job imports; it does not reclaim. Existence is checked
explicitly rather than letting the foreign key reject the insert, so an
orphan is counted as one instead of surfacing as an opaque constraint
error.
|