Skip to content

docs: specify M10 — attribution - #28

Merged
davior merged 1 commit into
mainfrom
claude/lucid-faraday-kq5d2l
Sep 20, 2026
Merged

davior merged 1 commit into
mainfrom
claude/lucid-faraday-kq5d2l

Conversation

@davior

@davior davior commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Docs only. No code, no migrations, nothing to deploy — this is the spec for M10, plus two corrections to the outstanding list.

Why M10 jumps ahead of M8 and M9

An asset can currently say how it entered the library and nothing about whose work it is. Asset.source (local_upload | url | ai_generated | gvc_export | clip | sub_video) is routinely mistaken for attribution and is not — it records the pipeline, not the credit.

The cost of deferring is asymmetric. Captured at ingest, attribution is a form field. Reconstructed later it means opening every file by hand and guessing, and for material gathered from the open web the answer is often unrecoverable — the tab is closed and the page may be gone. Neither M8 (needs a fal.ai account) nor M9 (needs GVC to exist) has that property, and both can wait. A library that cannot say where anything came from also cannot be published from: GVC cannot build a credits roll out of prose.

The three decisions

Each was put to the user explicitly. Rejected alternatives are recorded in the doc, because the reasoning is not recoverable from the schema alone.

Eight structured fields, not one free-text credit. source_url, creator, publisher, source_title, published_date, retrieved_at, license, credit_line. A credit line cannot answer "everything from this programme" or "everything published before 2020" without a parser over hand-written prose, which is a worse problem than the one it solves. A leaner four-field option was offered and declined: it defers exactly the columns that cannot be added cheaply later, since a migration does not recover information that was never captured.

Embedded file metadata is written directly; an AI may only suggest. EXIF Artist, ID3 tags, mp4 format.tags and PDF Author are facts about the file, so they are written into empty fields at ingest — and GAM is discarding them today, since ingest/probe.py already calls ffprobe with -show_format and drops everything but duration and codec. What a model reads off a chyron is a claim, so it goes through the existing Suggestion accept/reject flow. The failure modes are not symmetric: a blank field is visibly incomplete and prompts a fix, while a confidently wrong publisher looks finished and ends up crediting somebody else's work to the wrong outlet. Proposals must quote the evidence they were read from, and a field the model cannot ground is omitted rather than guessed.

Clips inherit from the parent at read time, with per-field override. Copying at creation is simpler to query and rots: correcting a publisher six months on would never reach the clips already cut from it, silently. Resolution costs nothing — to_read_model already receives the parent and parents_for_many already batch-loads them, because M7 needed that for file_url.

Two schema notes that break house convention on purpose

Both are commented in the spec so they are not later "fixed":

  • published_date is an ISO partial-date string, not a datetime. Publication dates are routinely partial — a book is from 1994, a magazine piece from March 2019 — and a datetime column forces a fabricated precision indistinguishable from a real one, which in a citation record is the exact falsehood this milestone exists to prevent. YYYY / YYYY-MM / YYYY-MM-DD sort and range-compare correctly as plain strings, so the filters need no parsing.
  • credit_line is an override, not the composed value. A stored composition is stale the moment a component field is corrected — the same rot that made copy-on-create wrong for clips.

field_provenance also gains a third value, "embedded", so a later suggestion can propose over a camera-supplied name without ever proposing over something the user typed.

Risk flagged for implementation

SQLite FTS5 cannot ALTER TABLE ADD COLUMN, so indexing attribution means dropping and recreating asset_fts — which, because it stores its own copy of the text rather than using external-content mode, loses the index for every existing asset unless the migration repopulates it. A recreate that skips the repopulate passes any test that only checks the schema and silently empties keyword search in production. The spec splits this into its own Alembic revision, separate from the eight add_column calls, so a failure in the risky half does not strand the columns. f68af8d exists because 7d4b9c1a6f28 hit the neighbouring version of this.

Corrections to the outstanding list

Two entries were overtaken by code and left listed as gaps. Verified against the source, not assumed:

  • Tag-management screen — built. TagPanel.tsx is mounted at frontend/src/views/SettingsView.tsx:513 and wires rename / remove / recategorise plus the full category CRUD.
  • Search ignoring asset_type and limit — built. SearchView.tsx passes both, with type chips and a load-more toward MAX_LIMIT, persisted in the URL. There is still no offset, so it is a growing page rather than true pagination — deliberate at this scale.

A stale gap is worse than no list, since it invites rebuilding something that works.

Testing

Docs only; nothing to run. Prettier covers frontend/src/**/*.{ts,tsx,css} only, so markdown is outside the CI format gate.

🤖 Generated with Claude Code

https://claude.ai/code/session_019RWM7S6z1UPsZqso4p2HAj


Generated by Claude Code

An asset can say how it entered the library and nothing about whose work
it is. `Asset.source` is routinely mistaken for attribution and is not:
it records the pipeline, not the credit.

The cost of deferring is asymmetric, which is why this jumps ahead of M8
and M9. Captured at ingest it is a form field; reconstructed later it
means opening every file by hand, and for material gathered from the open
web the answer is frequently gone. Neither M8 (needs a fal.ai account)
nor M9 (needs GVC to exist) has that property.

Three decisions taken with the user, recorded with their rejected
alternatives so a later session reads them as decisions rather than
accidents:

- Eight structured fields, not one free-text credit. A credit line cannot
  answer "everything from this programme" without a parser over prose,
  and GVC needs the publisher as a field, not a substring.
- Embedded file metadata is written directly; an AI may only suggest.
  EXIF/ID3/mp4 tags are facts about the file — ffprobe is already
  fetching `format.tags` and discarding them. What a model reads off a
  chyron is a claim, and a fabricated citation is worse than a blank one:
  a blank field prompts a fix, a confidently wrong publisher looks
  finished and ends up crediting the wrong outlet. Proposals must quote
  the evidence they were read from or be omitted.
- Clips inherit from the parent at read time, with per-field override.
  Copying at creation rots: correcting a publisher would never reach the
  clips already cut from it, silently.

Two schema notes that break house convention on purpose, so they are not
later "fixed": `published_date` is an ISO partial-date string because a
datetime cannot represent "1994" without fabricating precision, and
`credit_line` is an override rather than the composed value, because a
stored composition is stale the moment a component field changes.

Also corrected two entries in the outstanding list that the code has
since overtaken — the tag-management screen (`TagPanel.tsx`, mounted at
`SettingsView.tsx:513`) and search's `asset_type`/`limit` handling. Both
were built and left listed as gaps, which is the exact failure this list
exists to prevent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RWM7S6z1UPsZqso4p2HAj
@davior
davior marked this pull request as ready for review September 20, 2026 07:28
@davior
davior merged commit b3e8fc4 into main Sep 20, 2026
3 checks passed
@davior davior mentioned this pull request Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants