Commit Graph

17 Commits

Author SHA1 Message Date
Edouard Vanbelle 10c362a94a fix(jobs): flush the checkpoint tail, so progress reflects reality
All three import jobs only checkpointed on a full batch, so the
remainder after the last one was never counted. A run shorter than
BATCH_SIZE never checkpointed at all: `scanned_count` stayed 0 against
a known `total_rows`, and the admin progress bar sat at zero for the
whole run and finished there.

Seen on a transcode_import run over 20 entries — 13 imported, 5
negatives, 2 already present, progress 0/20 throughout. The thumbnail
imports had it too, just less visibly: a 105-file run reported
`scanned_count: 100`, losing the tail rather than all of it.

Cursor-wise the final checkpoint is a no-op — the walk is finished, so
nothing resumes from it — but the scanned delta is what the progress
display reads, and it has to include the last partial batch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:41 +02:00
Edouard Vanbelle fb0925d10a fix(thumbnails): a drained tier is not a teardown failure
Every boot after the migration completes logged
`WARN legacy sidecar directory could not be removed / No such file or
directory`. The directory being absent IS the end state — it is what
success looks like from the second boot onward — so this warned about
the migration having worked, forever, on every restart.

Returns early when the root is gone, which also skips walking three
directories that no longer exist. The remaining `Err` arms keep their
warning for the cases that are genuinely failures: a directory that
exists and cannot be removed or moved aside.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 16:37:40 +02:00
Edouard Vanbelle ce4354f497 fix(thumbnails): neither import job may tear down the shared directory
Found on a sandbox restore. `thumb_derived_import` ran first, imported
and deleted its own hash-named sidecars, then found `remove_dir` refused
because the `ext-*.jpg` previews were still there — those belong to
`thumb_attached_import`. The rename fallback fired, moving the tree to
`.thumbnails.migrated`; the attached job then looked in `.thumbnails/`,
found nothing, and reported zeros.

That stranded the user-uploaded previews, which are the one class of
file here with no render path to rebuild them. The rename exists for
files NEITHER job claims — a `.DS_Store` blocking removal forever — and
it fired for the sibling's work in progress instead. Inverting the job
order does not help: once the tree is renamed, both jobs look at
`.thumbnails/` and find nothing, whatever order they run in.

Teardown is now shared and refuses to act while anything remains that
either job would claim. Both jobs call it, so whichever finishes last
removes the tree in the same boot rather than leaving an empty
directory until the next one. The rename survives for its original
purpose, and now only fires when the remaining files are genuinely
nobody's.

Also drops the daily tick on both imports — they are on-demand now. The
boot run in repair mode IS the migration: nothing has written a sidecar
since step 10d2, so the tail cannot grow afterwards, and a tick could
not finish the job anyway because ticks never pass `repair`. Once
drained it was a `read_dir` returning nothing, every day, forever.

UX: the "at boot" badge moves from beside the job name into the cadence
column. It answers WHEN a job runs, which is what that column is for —
next to the name it read as a property of the job, and the row could
show "on-demand" beside a badge saying otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 16:19:45 +02:00
Edouard Vanbelle 577ecb7cef feat(jobs): run the thumbnail migration at startup, by default
A migration nobody triggers never finishes. Scheduled ticks deliberately
never pass `repair`, so a deployment whose operator never opens the
admin panel re-imported the same sidecars forever and never drained the
directory — and relying on operators to edit `.env` has the same failure
mode one level up.

`OXICLOUD_STARTUP_JOBS` dispatches named jobs once, in the background,
after the scheduler is ready. Entries use the syntax operators already
type at the trigger URL (`name?repair=true`), so the value is literally
the request they would otherwise make by hand. It defaults to both
migration jobs in repair mode, so an untouched deployment migrates and
drains itself.

That is a destructive default and a real exception to
no-silent-auto-repair, so the guard it rests on had to get stronger:
`verify_and_unlink` now compares CONTENT, not length. A blob of the
right size and the wrong bytes used to pass — a key-mapping bug handing
back another file's preview at the same length would have deleted the
original and kept the impostor, and thumbnails cluster tightly enough in
size for that to be a real coincidence. The readback streams from the
backend with no cache in front, so it proves durability rather than that
a write was acknowledged.

Deletion of `.thumbnails/` is attempted first and only falls back to
renaming it `.thumbnails.migrated` when `remove_dir` refuses because a
non-sidecar file is inside (Finder's `.DS_Store`). Either way the
directory stops existing, which lets the read-path probe go back to a
single `stat` on the root instead of walking the size directories.

Validation is fail-fast: an unknown job name or flag panics at boot. A
silently dropped `?repare=true` would leave the job in discovery-only
mode while the operator believed the tier was draining, surfacing months
later as "the migration never finished" with nothing pointing at the
config line.

Interrupted runs resume. Boot recovery flips abandoned rows to Paused
with their cursor, so `run_or_resume` continues rather than rescanning —
a long migration completes across however many restarts it takes. That
is a scoped exception to "we do not auto-resume": here somebody did ask,
in configuration, and not having to ask again is the point.

`StartupJob` holds a `JobRunArgs` rather than re-listing its four
fields, so a fifth flag cannot be added to the scheduler and silently
ignored in configuration.

Jobs named here are ordinary registered jobs — visible in the panel,
triggerable by hand, same runs and findings. Their rows now carry a
`startup` object so an operator can see that a job deletes on every boot
rather than only when someone clicks Run.

Adds docs/config/thumbnail-migration.md: what runs on first boot, how to
snapshot database and storage together beforehand, and how to verify
afterwards with satellites_consistency plus backend_consistency
?deep=true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 03246305f6 feat(thumbnails): the sidecar fallback disables itself
Step 10e was written as a removal release: delete the fallback read
path once the directories are empty. That has the same flaw as gating
deletion on an empty tail, one level up — sidecars are local disk, so
no release can know that every instance has drained.

The only removal that can actually be written is "if the tier is gone,
return". `initialize` now probes the size directories once at boot;
when absent, every fallback read short-circuits on a relaxed atomic
load and touches no filesystem. The code stays, costs nothing, and can
be deleted whenever — or never.

Two things had to change for absence to be reachable at all:

* `initialize` no longer creates the directories. It create_dir_all-ed
  all three at every boot, so the import job removed them and the next
  restart put them back — the absence this gates on was unreachable by
  construction. Found on a sandbox where the job had drained the tier
  and a restart left three empty directories behind. Nothing has
  written a sidecar since step 10d2, so there was nothing to create
  them for.
* The probe tests the size directories, not the root. On macOS Finder
  leaves a .DS_Store in the root, which blocks remove_dir there
  permanently; gating on the root would keep the fallback alive on
  every developer machine for a reason unrelated to thumbnails. No size
  directory means no sidecar.

Every sidecar read and existence check now goes through `read_sidecar`
/ `sidecar_exists`, so the guard exists once rather than at each of the
twelve sites that built a path and read it — the build-then-read pair
was duplicated six times over.

The import job's root removal reports its outcome instead of discarding
it. It is the one result an operator is waiting for, and "directory not
empty" with no sidecars left is a failure worth naming.

Falls open: the flag starts true, so a service constructed without
`initialize` behaves as before. A drain completing mid-process leaves
it stale-true until restart, which costs the same failed opens as
today; it never goes false while sidecars remain.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 1ea3826660 feat(jobs): jobs describe themselves — description, mutates, repair_description
The admin panel had no repair toggle wired to anything but a hardcoded
name list naming the two refcount tenants, so `thumb_derived_import` and
`thumb_attached_import` could not be run in repair mode from the UI at
all despite supporting it. And nothing in the job list said what any
given job does or whether clicking Run on production writes anything.

Three defaulted methods on `JobHandler` and `RecoverableJobHandler`:

    fn description(&self) -> &'static str
    fn mutates(&self) -> Mutates          // Never | Always | OnRepairOnly
    fn repair_description(&self) -> Option<&'static str>

`RecoverableAdapter` forwards them — the registry only holds
`dyn JobHandler`, so a tenant's metadata is invisible otherwise, and
falling back to the defaults would report every recoverable job as
read-only, including the ones that delete files.

Three values rather than a boolean because a job can be read-only by
default and destructive under `?repair=true`; a boolean answers wrongly
for one of its two modes, and `false` on something that unlinks files is
the dangerous direction to be wrong in. `repair_description` returning
`Option` collapses "does it repair" and "what does repair do" into one
method: presence gates the toggle, content is the confirmation text —
which the frontend cannot invent, since correcting a counter and
deleting sidecars are not the same warning.

`OnRepairOnly` with no `repair_description` is rejected at registration:
it claims to mutate only under a flag it does not support.

All 17 registered jobs declare all three. The panel now renders the
description under each name, badges read-only jobs, confirms before a
plain run of a mutating one, and offers the repair variant off the
backend flag instead of the name list.

Descriptions are English in the trait, next to the behaviour: one in
`locales/*.json` rots invisibly the moment a job changes, and a
translator cannot know what `manifests_consistency` reconciles. i18n can
layer on later keyed by job name with these as the fallback.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:41:05 +02:00
Edouard Vanbelle b485db46fa feat(storage): audit every sidecar deletion, and reclaim orphaned uploads
Two changes to the import jobs' destructive path.

thumb_attached_import now deletes orphaned sidecars under `repair`,
matching the dead-source case on the derived side. An `ext-` file whose
owner is gone is unimportable — the FK on file_id would reject the row —
so leaving it means it is rediscovered every run, the tail never empties
and step 10e's gate never opens. Safe despite these being the
non-regenerable bytes: the preview is keyed to a file_id that no longer
exists, so nothing can reference it again. Unrecoverable and unreachable
are different things, and this is both.

And every deletion is now audited. A one-way migration removing
user-visible files should leave a trail that outlives the run history:
findings are per-run and get purged, whereas target: "audit" is
separable and retained. If a preview later turns out to be missing, this
is the only record saying the migration removed it and when.

`owner` carries the id the file belonged to — source_hash for
content-keyed, file_id for uploaded — because that is where an
investigation starts, and the raw logs cannot supply it: NEW BLOB names
the hash of the STORED BYTES, a different value from the sidecar's own
name, which is why grepping one against the other finds nothing.

reason is a stable key: `imported` (replaced by a verified blob),
`source_gone`, `orphaned`. The first lives inside verify_and_unlink so a
verified deletion cannot be logged inconsistently; the other two are
explicit, since those paths have nothing to verify against.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 1a3d7d201a fix(storage): skip sidecars whose source is gone, before writing anything
Running the import on a real install produced a store-then-discard loop:
NEW BLOB (CDC) immediately followed by MANIFEST DELETED, once per
sidecar. store_derived_blob wrote the bytes, the source-exists guard
refused the row, and `inserted == 0` released the reference again.

The refusal is right — `.thumbnails/` outlives years of deleted files,
and importing those would recreate exactly the orphan rows e4c78ae0
eliminated. The mistake was deciding it AFTER the write.

Now checked before the read and the store, via blob_exists (manifest
first, blob as fallback). Two costs it removes: a blob write plus a
manifest delete per dead sidecar on EVERY run, and a tail that never
empties — unimportable files are rediscovered forever, so the job never
reports zero and step 10e's gate never opens.

Reported as `sidecar_source_gone` so the scale is visible before
anything is removed, and deleted under `repair`. That is the one unlink
in this job needing no readback: there is nothing to read back and
nothing to regenerate from.

Counted separately in the completion log, because "skipped, source gone"
and "already present" mean different things to an operator deciding
whether the migration has converged.

Worth noting for anyone reading the raw logs: NEW BLOB names the hash of
the STORED BYTES, while the sidecar filename is the SOURCE hash. They
are different values, so grepping the log hash against .thumbnails finds
nothing. The new finding carries both.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 1b68ee093e feat(storage): both import jobs tick daily instead of manual-only
Registered with interval None, so they ran only when someone remembered
to trigger them — which was your objection to gating anything on
operator timing. Now daily.

Not boot-time: that would delay readiness for a filesystem walk, and
both jobs are idempotent and resumable, so periodic is strictly better.

The tick deliberately does NOT delete. `repair` defaults false, so
scheduled runs import and stop; unlinking stays a deliberate operator
action, per no-silent-auto-repair. That splits the two halves the way
their risk differs — the backfill is safe to automate, removing files is
not.

Cost once drained is a read_dir over three directories returning
nothing, and after the directory itself is removed, not even that.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle df619a7ed9 feat(storage): thumb_attached_import drains its sidecars too
Completes step 10d. Same `?repair=true` opt-in and the same
readback-before-unlink as the derived half, and the check matters more
here: these sidecars hold the bytes that CANNOT be regenerated — a
client-uploaded PDF preview has no server-side render path — so the read
is the only thing between a migration and permanent loss, not
belt-and-braces.

verify_and_unlink is shared rather than copied. Two versions of "only
delete after proving the replacement is readable" would be two chances
to weaken one, and it is the rule the whole deletion step rests on.

Deletion covers the already-imported branch as well as fresh imports,
for the same reason as the derived job: a run without `repair` leaves
the sidecar behind, and a later run with it would otherwise see "already
imported" and never drain. Import first, enable deletion after, is the
expected operator sequence, so that branch is the common path.

Orphaned sidecars stay untouched — this job imports, it does not
reclaim, and a destructive default on a migration is what no-silent-
auto-repair forbids. Unverifiable ones are kept and reported, so the
next run retries.

With both halves draining, `.thumbnails/` can now actually empty and the
directory removal in the derived job can succeed — though not stably
until dual-write stops, since any render or upload recreates it.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle ae9ff0d8b3 feat(storage): thumb_derived_import drains the sidecars it imports
Step 10d, for the content-keyed half. Opt-in via the existing `repair`
flag rather than a new one — the house rule is that a job does not
mutate on its default setting, so early runs import only and an operator
can inspect before committing.

Deleting from the job rather than a later release is what makes the
migration self-draining. Sidecars are LOCAL disk, so no release can know
whether every instance has finished; each instance draining itself needs
no coordination at all.

Verification before unlinking is the load-bearing part.
store_derived_blob reporting success is not proof the bytes are
retrievable — a backend that accepted a write it cannot serve would
otherwise have the last copy deleted on top of it. The blob is read back
and its length compared against the sidecar's; a failure keeps the file,
records a finding, and the next run retries. That read is the difference
between a migration and a data-loss bug.

Deletion applies to already-imported files too, not just fresh ones. A
run without `repair` leaves the sidecar behind, and a later run with it
would otherwise classify the file as "already imported" and never drain
it — and import-then-enable-deletion is the expected operator sequence,
so that is the common path rather than an edge case.

Directories are removed once genuinely empty, because ABSENCE is what
step 10e gates on, not emptiness: empty is momentary and an on-demand
render can repopulate it a second later, while absence is one-way and
cheaper to test (one stat, versus opendir/readdir/closedir). remove_dir
refuses a non-empty directory, so it needs no emptiness check and cannot
race a concurrent write into deleting live files.

thumb_attached_import still needs the same treatment; verify_and_unlink
should move somewhere shared rather than being copied into it.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 12158ccf59 fix(storage): thumb_derived_import claims JPEG sidecars too
The filter was strip_suffix(".webp"), but persist_rendered writes
{hash}.{format.ext()} — so any client not advertising WebP leaves
{hash}.jpg on disk. Correct only while the derived tier was WebP-only;
once variant carried the format (20261022000000) a JPEG sidecar became
ordinary content, and leaving it unclaimed would keep .thumbnails/
permanently non-empty — the very signal step 10e gates on. The migration
could never finish.

Both codecs are now claimed and the format comes from the file's own
extension, so a .jpg imports AS JPEG. Deriving the variant and
content_type from it rather than hardcoding WebP is the point: a
mislabelled row would serve the wrong codec to whoever the read path
then matched it for.

ThumbnailFormat::ALL exists so the claim list and the write path cannot
drift — adding a format without teaching the import about it would
strand that codec silently.

The `ext-` rejection now carries real weight. Previously .jpg was
rejected wholesale, so the two jobs could not overlap by construction;
now they share an extension and only the prefix separates them. Both
directions stay under test.

Caught by the cross-job assertion, which counts every real sidecar being
claimed exactly once — the fixture gained a .jpg and the total moved 3
to 4, which is the test noticing rather than a test to update.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle d202b4b5ca fix(storage): the derived import conflated variant with directory
76590160 changed `variant_of` to return `{size}.{ext}`, but that value
was also being used as the on-disk DIRECTORY. Reads became
`.thumbnails/preview.webp/{hash}.webp`, which does not exist, so every
sidecar counted as unreadable and thumb_derived_import restored nothing.

Caught by thumb_import_check.sh on the run after the migration — the
harness earning its keep twice now, since this is the second defect it
has caught that no unit test could.

They are genuinely two strings and are now named as such: `dir_name` for
the path, `variant` for the row key. The cursor keeps using the
directory, so a run paused before the migration resumes at the same
position rather than restarting.

Also stops podman's compose-provider banner from burying the script's
output. Filtered rather than discarded, so genuine psql errors still
surface — swallowing those would turn a broken query into a silently
wrong assertion. Suppressing it at the source needs
`[engine] compose_warning_logs = false` in containers.conf, which is
per-developer config and cannot be relied on in CI.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 86d0d65583 feat(storage): derived variant encodes the output format
`content_derived_blobs.variant` held the size alone, so one source could
hold exactly one artifact per size regardless of codec. That surfaced
when the read order flipped in 10c: a JPEG request matched the WebP row
and would have been served the wrong codec — hidden previously because
the .jpg sidecar won first. The flip had to be gated to WebP, which
meant JPEG clients could never leave the sidecar, which meant the
sidecar could never be deleted.

It blocks transcodes harder: those are multi-format by nature, so two
output codecs of one source collide on the primary key without a format
term.

The axis goes inside the string rather than into a fourth PK column,
per the column's own rule — "new axes go inside this string, never into
new columns". Shape is {size}.{ext}: preview.webp, icon.jpg, later
720p.webp.

The backfill is deterministic, not a guess: store_derived_blob has only
ever written "image/webp" for thumbnails. content_type is checked anyway
rather than assumed — a row that fails the assumption is left alone and
counted in a warning, because the read path then simply misses it and
falls back to the sidecar, whereas guessing a codec would serve wrong
bytes. Idempotent via NOT LIKE '%.%', so a re-apply cannot produce
preview.webp.webp; verified on a scratch PG by applying it twice.

One helper builds the string, because it is a primary-key component:
a writer and reader that disagree do not fail loudly, they just never
find each other's rows and the derived tier silently looks empty. It
lives on the service's ThumbnailSize, not the port's — they are distinct
types, which the compiler pointed out after I put it on the wrong one.

The WebP gate on the read path is now removed: each codec has its own
row, so JPEG can finally reach the derived tier — the prerequisite for
deleting the sidecar for those clients.

file_attached_blobs keeps a bare size: store_external_thumbnail
re-encodes everything to JPEG, so it is single-format by construction
and a format term would cost a migration for nothing.
2026-08-30 13:41:05 +02:00
Edouard Vanbelle 4ae1531286 docs(plan): job-driven sidecar deletion, and the persist-consolidation blocker
Two revisions from working through step 10.

**Deletion moves into the import jobs, not a release.** Sidecars are
local disk, so a release cannot know whether every instance has drained
— gating on "an empty tail" asks an operator to coordinate a fact
nothing reports, and there is no telling when or whether they trigger
the jobs at all. Each job unlinking what it has imported makes every
instance drain itself. Constrained three ways: verify the derived blob
reads back before unlinking (a store that reported success but landed
unreadable would otherwise take the last copy), only after the
read-order flip (or the derived tier takes its first production traffic
by accident), and opt-in, since a migration that deletes by default is
surprising. Scheduled tick rather than boot trigger — idempotent and
resumable, so periodic is safe, while walking .thumbnails/ at startup
delays readiness for nothing.

**Found while checking the dual-write assumption: it does not hold.**
store_derived_blob has ONE call site; fs::write(&thumb_path, …) has
five. get_thumbnail, generate_and_persist and
generate_all_sizes_background all persist sidecar-only. That breaks the
migration's premise rather than being untidy — on-demand renders keep
producing un-migrated state after the import runs, so the tail never
empties and the deletion gate never opens. One persist_thumbnail owning
sidecar + derived + moka is therefore a prerequisite, and it makes "stop
writing sidecars" a later one-line change instead of four edits. Noted
that ThumbnailService holds no DedupService, so it must be threaded
through.

Also corrects a claim I put in thumb_derived_import's own docs:
transcoding is NOT a later step. ImageTranscodeService exists and caches
.transcoded/{ext}/{file_id}.{ext}, so a third import is needed and it
must re-key file→content — legitimate only because a transcode is
derivable. Its .skip markers remain an open question.
2026-08-30 13:41:04 +02:00
Edouard Vanbelle 49e7bb15a6 test(storage): cover the sidecar walk for both import jobs
Both imports run over ONE directory, where the two legacy shapes sit
side by side, so the property worth asserting spans them: together they
must claim every real sidecar exactly once, and neither may take the
other's. A job that drifted into the other's shape would content-key
user-supplied bytes — sharing one user's uploaded preview onto every
file with identical content — and no per-job test in isolation would
notice.

So the fixture is shared. `legacy_tree` builds a directory holding a
content-keyed .webp pair, an ext- upload, and a stray README, and both
test modules walk it: derived claims exactly the two hashes in sorted
order, attached claims exactly the ext- file, the two sets are disjoint,
and between them they account for all three real sidecars.

`sidecar_names` became an associated function taking the root instead of
reading `self`, which is what makes this testable at all — the walk is
the half that decides which files a job claims, and it needed no pool to
verify. Sorting is asserted rather than assumed, since the cursor
resumes by skipping everything at or before it and a stable order is the
only thing that makes that correct.

A missing size directory is covered too: normal on a fresh install, and
it must yield no work rather than abort the walk.

Not covered here, and it needs a pooled fixture that does not exist: the
round trip itself — store the blob, write the row, and confirm a COPY
inherits the preview. That belongs in the API-level harness, where the
legacy state can be manufactured through the real write path and then
stripped.
2026-08-30 13:41:04 +02:00
Edouard Vanbelle f80a28763e feat(storage): thumb_derived_import — backfill the derived tier from sidecars
First half of step 10. Every server-rendered thumbnail written before
content_derived_blobs existed lives only as
{thumbnails_root}/{size}/{hash}.webp — local-disk state that another
instance cannot see, a backend migration does not carry, and no
consistency job covers. This walks those files into the blob store and
records the mapping, so the derived tier can become authoritative and
the sidecar can be deleted.

A registered JobRegistry tenant rather than a script: the volume is
unbounded, so it needs a cursor, resume, cooperative cancel and run
history, and an operator needs somewhere to watch it. Cursor is
{size_dir}/{filename} over a sorted walk, which totally orders the
traversal.

Idempotent by construction — each file is skipped when a row already
exists, and store_derived_blob is ON CONFLICT DO NOTHING with
release-on-conflict beneath it, so re-runs cannot inflate refcounts.
Re-running is the expected operator behaviour, since Phase 3 (deleting
the sidecars) is gated on a run reporting zero imported.

hash_from_sidecar_name deliberately rejects ext-{file_id}.jpg. Those
bytes are user-supplied and file-keyed; importing them here would
content-key them and share one user's uploaded preview onto every file
with identical content. They belong to thumb_attached_import. Both the
accept and the reject set are under test.

Unreadable files and store failures record a finding and continue: a
sidecar removed by a concurrent GC unlink between listing and read is
expected, not fatal, and the file is left in place for the next run.

Registered unconditionally rather than behind a flag — a migration
nobody can find is a migration nobody runs.
2026-08-30 13:41:04 +02:00