diff --git a/docs/m10-attribution.md b/docs/m10-attribution.md new file mode 100644 index 0000000..dd5b0bc --- /dev/null +++ b/docs/m10-attribution.md @@ -0,0 +1,336 @@ +# M10 — Attribution + +**Status:** specified, not built. This document is the spec; no code exists yet. + +Numbered M10 but landing **before** M8 and M9. Out of numeric order on purpose: M8 needs +a fal.ai account and M9 needs GVC to exist, while this needs neither and is the one piece +of work that gets more expensive every day it is deferred. + +--- + +## Why this exists + +GAM records *how* a file entered the library and nothing about **whose work it is**. +`Asset.source` (`local_upload | url | ai_generated | gvc_export | clip | sub_video`) is +routinely mistaken for this and is not: it is a pipeline fact, not a credit. There is no +field for a source URL, an author, a publisher, or the programme a clip was taken from, +no UI for one, and nothing in M8–M9 that adds one. + +The cost of deferring is asymmetric, which is why this jumped the queue. Attribution +captured at ingest costs a form field. Attribution reconstructed later means opening every +file by hand and guessing — and for material gathered from the open web, the guess is +often unrecoverable, because the tab it came from is closed and the page may be gone. A +library that cannot say where anything came from also cannot be published from: GVC cannot +build a credits roll out of prose, and an uncreditable asset is one that has to be left out. + +The user asked for this before adding more content. That is the correct order. + +--- + +## The three decisions + +Each was put to the user explicitly. The rejected alternatives are recorded because the +reasons are not recoverable from the schema alone. + +### 1. Structured fields, not a free-text credit + +**Chosen:** eight columns. **Rejected:** a single `credit_line` string. + +A free-text credit is faster to fill in and impossible to get structurally wrong, and it +fails the moment the library is asked a question. "Everything from this programme", +"everything published before 2020", "everything still missing a source" are all +unanswerable against prose without parsing it, and a parser over hand-written citations is +a worse problem than the one it solves. Downstream, GVC building a credits roll needs the +publisher as a field, not as a substring. + +A middle option — four fields (`source_url`, `creator`, `publisher`, `credit_line`) — was +offered and declined. It is less to type per asset and it defers exactly the columns that +cannot be added cheaply later: adding `published_date` to a populated library means a +migration *and* a manual re-entry pass over every row, because the information was never +captured. + +A ninth field, free-text `rights_notes`, was offered alongside the eight and **not** +taken. Recorded here so a later session reads it as a decision rather than an oversight. +The messy cases it was meant to hold — "cleared for non-commercial use", "permission by +email 2024-03-11", "unknown, found on a forum" — currently go in `license`, which is free +text. If that field starts accumulating sentences rather than licence names, that is the +signal to revisit this, and the migration is additive. + +### 2. Embedded metadata is written directly; AI may only suggest + +**Chosen:** two separate paths, deliberately asymmetric. + +**Embedded metadata is a fact about the file.** A JPEG's EXIF `Artist`, an MP3's ID3 tags, +an MP4's `format.tags`, a PDF's `Author` — these were written by whoever produced the +file. Reading them is not inference, so they are written straight into empty fields at +ingest. GAM is currently discarding this: `ingest/probe.py` already invokes ffprobe with +`-show_format`, which returns `format.tags`, and the result is parsed for duration and +codec and thrown away. The data is already on the wire. + +**What a model reads off a chyron is a claim.** It goes through the existing `Suggestion` +accept/reject flow and never touches the asset until a person says yes. The reasoning is +specific and worth stating plainly: a fabricated citation is worse than a blank one. A +blank field is visibly incomplete and prompts the user to fill it. A confidently wrong +publisher looks finished, gets copied into a credits roll, and ends up attributing +somebody's work to the wrong outlet. The failure modes are not symmetric, so the trust +model is not either. + +This is also why attribution does **not** follow the `describe`/`summarize` pattern, where +pressing Generate overwrites the field. That pattern is right for a description — an +inaccurate description is a poor search result. It is wrong for a citation. + +**The grounding rule.** The proposal prompt requires the model to return a field only +where it can quote the text it read the value from: a visible chyron, a byline, a title +page line, a watermark. That quote is stored on the suggestion as `evidence` and shown in +the accept UI, so accepting is an informed act rather than a reflex. A field the model +cannot ground is omitted rather than guessed, and the test suite asserts that an +ungrounded reply produces no suggestions at all. + +### 3. Clips inherit from the parent, with per-field override + +**Chosen:** resolve through `parent_asset_id` at read time. +**Rejected:** copy the parent's values when the clip is cut. + +Copying is simpler to query and makes each clip self-contained, and it rots. Discovering +six months later that a documentary's publisher was recorded wrong means every clip cut +from it before the correction keeps the wrong value forever, silently, with nothing to +indicate they disagree with their own parent. Citation records that drift apart from the +source they cite are worse than no records, because they are trusted. + +Resolution at read time costs nothing here: `to_read_model` is already passed the parent, +and `parents_for_many` already batch-loads parents for both listings and search results, +because M7 needed exactly that to resolve a clip's `file_url`. The pattern is in place. + +Override is per field, not all-or-nothing, so a clip can carry its own timestamped +citation or a different creator for one interviewee without discarding the parent's +publisher and source title. + +Inheritance is **one level only**, and that is sufficient rather than a simplification: +`create_clip` refuses to clip a clip, and `promote_clip` leaves `parent_asset_id` pointing +at the original. There is no chain to walk, and the code should say so, so nobody adds a +recursive resolver for a depth that cannot occur. + +--- + +## Schema + +Eight columns on `Asset`, in their own block after the transcript header. + +| Column | Type | Notes | +|---|---|---| +| `source_url` | `str \| None` | Where it came from. Not validated as reachable — a dead link is still a record, and often the only one. | +| `creator` | `str \| None` | Author, photographer, speaker, director. | +| `publisher` | `str \| None` | Outlet, channel, studio, imprint. | +| `source_title` | `str \| None` | The programme, film, article or book this is part of. | +| `published_date` | `str \| None` | ISO 8601 partial date — see below. | +| `retrieved_at` | `datetime \| None` | When it was downloaded or captured. Naive UTC from `app.clock.utcnow`. | +| `license` | `str \| None` | Rights status, free text. | +| `credit_line` | `str \| None` | Override only — see below. | + +### Why `published_date` is a string + +This breaks the house rule that timestamps are naive UTC `datetime`s, and the model +comment must say why so it does not get "fixed" later. + +Publication dates are routinely partial. A book is from 1994. A magazine piece is from +March 2019. A broadcast has a full date. A `datetime` column cannot represent the first +two at all — it forces a fabricated precision (1994-01-01) that is then indistinguishable +from a real one, which in a citation record is exactly the kind of quiet falsehood this +milestone exists to prevent. + +The column stores `YYYY`, `YYYY-MM` or `YYYY-MM-DD`, regex-validated on write. ISO partial +dates sort and range-compare correctly as plain strings — `"2018-12-31" < "2019" < +"2019-03-01"` — so the date filters are lexicographic comparisons with no parsing and no +special cases. + +`retrieved_at` is a real `datetime`, because a download happened at an instant and there +is no partial-precision case to serve. + +### Why `credit_line` is an override, not the composed value + +`credit_line` is `NULL` until somebody types one. The displayed credit is composed on read +from the other fields. + +Storing the composed string would make it stale the instant any component field is +corrected — the same rot that made copy-on-create the wrong answer for clips, one level +down. The API therefore exposes both: `credit_line` (the raw override, for the edit form) +and `credit` (the resolved display string, composed unless overridden). + +### Why a third provenance value + +`field_provenance` gains `"embedded"` beside `"human"` and `"ai"`. + +An embedded tag is a fact about the file but not a claim a person has verified — an EXIF +`Artist` is often the camera owner's name, or a studio's default, or blank-but-not-empty. +Keeping it distinct from `"human"` is what allows a later AI suggestion or a bulk pass to +propose over an embedded value without ever proposing over something the user typed. Fold +it into `"human"` and that distinction is gone permanently, because nothing else records it. + +--- + +## Migrations — two revisions, deliberately + +`f68af8d` exists because `7d4b9c1a6f28` failed against a populated database. The second +half of this change carries the same risk, so it is split rather than stranding the +columns if the risky half fails: + +1. **`add_attribution_columns`** — eight `add_column` calls on `asset` inside + `op.batch_alter_table`, following the pattern `7d4b9c1a6f28` had to learn. + +2. **`rebuild_asset_fts_with_attribution`** — `asset_fts` gains an indexed + `attribution_text` column. + + **SQLite FTS5 does not support `ALTER TABLE ... ADD COLUMN`.** The virtual table must + be dropped and recreated, and because `asset_fts` stores its own copy of the text + rather than using external-content mode (see the reasoning at the top of + `search/fts.py`), dropping it **loses the index for every existing asset**. The + migration must therefore repopulate, in raw SQL, from `asset` left-joined through + `assettag`/`tag` for the tag text. A recreate that skips the repopulate passes every + test that only checks the schema, and silently empties keyword search in production. + + The migration carries its own literal copy of the new DDL, per the convention + `search/fts.py` already documents. `test_migrations.py`'s drift check + (`test_migrations.py:80`) extends to the new column set, and a new test asserts the + rebuild preserves previously indexed rows. + +--- + +## Backend shape + +### `app/attribution.py` (new) + +- `ATTRIBUTION_FIELDS` — the field-name tuple. The harvester, the suggestion path, the + FTS text builder and the inheritance resolver all read it, so they cannot drift apart. +- `compose_credit(...) -> str` — pure, session-free, directly unit-testable. Order is + creator — *source_title* — publisher — published_date — license, skipping blanks and + emitting no orphaned punctuation for a single-field asset or an empty string for a blank + one. +- `resolve(asset, parent) -> ResolvedAttribution` — per-field coalesce, clip's own value + wins. One level, with the comment explaining why that is complete. + +### `app/ingest/embedded_metadata.py` (new) + +Best-effort throughout and never raises. It runs inside `_describe`, which is already +post-commit precisely so that a failure here cannot lose an upload. + +- **video / audio** — `ProbeResult` gains `tags: Mapping[str, str]`; ffprobe already + returns them. `artist`/`album_artist` → creator, `publisher`/`copyright` → + publisher/license, `date`/`creation_time` → published_date, `title` → source_title, + `comment`/`purl` → source_url. +- **images** — Pillow EXIF `Artist` (0x013B), `Copyright` (0x8298), `DateTimeOriginal` + (0x9003), plus IPTC `By-line`/`Credit`/`Source` via `IptcImagePlugin`. No new dependency. +- **PDF** — pypdfium2 document metadata: `Author`, `Title`, `CreationDate`. +- **docx / pptx / xlsx** — `core_properties.author` / `.title` / `.created`. All three + libraries are already dependencies, added for `extract_text`. + +Writes only into empty fields, stamped `"embedded"`. + +### `services/assets.py` + +- `apply_metadata` and `apply_ai_metadata` share a private + `_apply(session, asset, changes, provenance)`; a third wrapper + `apply_embedded_metadata` joins them. Both existing functions keep their signatures and + behaviour exactly. +- `_reindex` composes `attribution_text` from the asset itself — no extra query, and for + the same reason it already reads tags rather than accepting them from callers: one + function every write path calls is one that nobody can forget. +- `to_read_model` resolves inheritance from the `parent` it already receives. + +### API + +- `AssetRead` gains the eight fields, plus read-only `credit` (resolved) and + `attribution_inherited: list[str]` naming which fields came from the parent, so the UI + can mark them without a second request. +- `AssetUpdate` gains the eight as editable, with the `published_date` regex validator. +- `list_assets` gains `creator`, `publisher`, `source_title`, `published_after`, + `published_before`, and **`unattributed: bool`** — the last is how a backlog actually + gets worked through, and is the most operationally useful filter in this milestone. +- `POST /api/assets/harvest-attribution` — a library-wide job re-reading embedded metadata + for assets already uploaded. `997f000122a7` made `EnrichmentJob.asset_id` nullable for + exactly this shape; it follows `routers/embeddings.py`'s backfill endpoint. + +### Suggestions + +`KIND_ATTRIBUTION = "attribution"`, one row per proposed field, `value` holding +JSON-as-TEXT `{"field": ..., "value": ..., "evidence": ...}`. One kind rather than eight, +and still per-field accept/reject. `enrichment/attribute.py` reuses `enrichment/source.py` +to choose the source material rather than rebuilding that decision. Accept dispatches to +`apply_metadata` — the *human* path, stamping `"human"` — exactly as `KIND_TITLE` does. + +--- + +## Frontend + +- `components/AttributionPanel.tsx` (new) — a tab in `DetailDock`'s `Tabs` beside + transcript / document / clips. Eight inputs, a live preview of the composed credit, and + a **Copy credit** button, which is the reason the composed field exists at all. + Inherited values render greyed with an "inherited from *parent*" note and an override + control. +- `components/SuggestionPanel.tsx` — renders the new kind as field, proposed value and the + quoted evidence. +- `components/FilterBar.tsx` — creator and publisher inputs, plus an **Unattributed** chip. +- `components/AssetDetail.tsx` — the credit line in the facts block. + +The panel's inputs must call `stopPropagation` on Escape. The existing bug — `TagInput` +and the transcript editor let Escape reach `DetailDock`'s window listener and close the +whole panel — is **not** in this milestone's scope, but this must not become a ninth +instance of it. + +--- + +## What this deliberately does not do + +- **No fetching metadata from `source_url`.** OpenGraph/oEmbed lookup against a + user-supplied URL is an outbound request to an arbitrary host; it needs `safe_url.py`, + a job, and a rate-limit story. Worth doing, and it belongs with URL import, which is its + own outstanding item. +- **No rights-clearance workflow.** `license` records what is known. Whether something may + be used is a judgement, not a column. +- **No credits-roll export.** That is GVC's job, and this milestone exists to give it + something to read. +- **No per-clip timestamped citation format.** A clip can override `credit_line` freehand; + a structured "at 04:12" convention can wait until GVC shows what it needs. + +## Open, and genuinely undecided + +Whether a bulk attribution pass over a selection is worth building, as M6 did for +enrichment. It probably is for a batch imported from one source in one sitting — twenty +screenshots from the same programme share every field. Left out of the first pass because +the `unattributed` filter plus the existing `SelectionBar` may make it a small addition +rather than a design. + +--- + +## Verification + +Backend (`cd backend && pytest -q`): +- `compose_credit` over each field subset, including a single-field asset (no orphaned + punctuation) and an all-blank one (empty string, not a bare separator). +- Inheritance: a clip with no attribution resolves the parent's; a clip overriding + `publisher` keeps the parent's `creator`; correcting the parent changes the clip's + resolved value on next read. +- Harvest per format against the existing `backend/tests/fixtures/` media, asserting the + `"embedded"` stamp and that a non-empty field is never overwritten. +- Migrations: both run against a **populated** database; the FTS rebuild preserves every + previously indexed asset; the drift check covers the new column. +- Search and filters: findable by publisher and by `source_title`; `unattributed=true`; + `published_after=2019` excludes `2018-12-31` and includes `2019-03-01`. +- Suggestions: a grounded reply creates one row per field; an ungrounded one creates none; + accept writes through `apply_metadata` and stamps `"human"`; reject is remembered. + +Frontend (`cd frontend && npm test && npm run build`) — panel renders, edits and saves; +inherited fields show their origin; the composed preview updates live; Copy credit writes +the resolved string. + +End to end (`uvicorn --port 8001` + `npm run dev` on 5174, after `alembic upgrade head` — +migrations do not run outside Docker and this adds two): + +1. Upload a JPEG carrying EXIF `Artist` and an MP3 with ID3 tags; both land attributed and + marked as having come from the file. +2. Upload a file with no embedded metadata, fill it in by hand, confirm the credit composes + and the asset is findable by publisher in `/search`. +3. Cut a clip of an attributed video; confirm the inherited credit, then correct the + parent's publisher and confirm the clip follows. +4. Run an attribution suggestion over a document with a visible byline; confirm the + evidence quote is shown and that accepting writes the field. +5. `unattributed=true` returns exactly the assets still missing a source. diff --git a/docs/plan-of-attack.md b/docs/plan-of-attack.md index 085b96c..b9a7701 100644 --- a/docs/plan-of-attack.md +++ b/docs/plan-of-attack.md @@ -308,6 +308,7 @@ Each is demoable on its own and ends in a draft PR on `claude/loving-mendel-x3xh | **M7** | **Clips & sub-videos** | In/out point editor; non-destructive clips; ffmpeg extraction; parent-delete guard with promote-to-sub-video. | | **M8** | **AI generation** | fal.ai image→image and image→video with multi-image bases; admin model catalog (`ModelCatalogEntry` pattern); generated assets through the standard ingest pipeline; regenerate-from-stored-prompt. | | **M9** | **GVC surface** | Documented read API; embeddable `/picker` route with `postMessage`; GN-4 links from Notes. | +| **M10** | **Attribution** | Eight attribution fields on `Asset`; embedded metadata harvested at ingest (EXIF/ID3/mp4 tags/PDF Author); AI proposals through the existing suggestion flow, grounded in quoted evidence; clips inherit from the parent with per-field override; attribution in the keyword index and in the library filters. Specified in [`m10-attribution.md`](m10-attribution.md). Numbered after M9 but scheduled **before** M8 — see Status. | Phase 2 (out of scope, named so it is not accidentally designed out): cloud/R2 storage, checksum dedup (the `checksum_sha256` column is already there for it), asset versioning, @@ -328,6 +329,7 @@ it was written in does not survive the session. | M3 Tagging | Merged — [#7](https://github.com/davior/gam/pull/7) (fast-forwarded, so no merge commit) | | M6 AI enrichment | **Complete, all 8 steps** — [#14](https://github.com/davior/gam/pull/14) `AIProvider`, its migration, `/api/providers` CRUD and the settings panel; [#15](https://github.com/davior/gam/pull/15) the three protocol clients and the retry/backoff layer; [#16](https://github.com/davior/gam/pull/16) the source-material spec and `summarize`; [#17](https://github.com/davior/gam/pull/17) `autotag`, the suggestion model, and generated titles; [#18](https://github.com/davior/gam/pull/18) `describe`; [#19](https://github.com/davior/gam/pull/19) `UsageEvent`, the pricing table and the cost readout; [#20](https://github.com/davior/gam/pull/20) bulk enrichment over a selection, which also picked up the `SelectionBar` embed deferred from M5. Two gaps M6 did **not** close are recorded in [`m6-ai-enrichment.md`](m6-ai-enrichment.md) and below: `extract_text`, since built, and attribution, still undesigned. | | M7 Clips & sub-videos | **Complete** — [#26](https://github.com/davior/gam/pull/26). Non-destructive clips, ffmpeg sub-video extraction (fast stream-copy with an automatic re-encode fallback), and a parent-delete guard with a one-click promote path. Two bugs caught before shipping — a promoted clip almost kept its parent-relative `in_point`/`out_point`, and the delete guard's own state was briefly getting wiped by a store-driven remount — are recorded in [`m7-clips-and-subvideos.md`](m7-clips-and-subvideos.md), along with why the milestone's actual file layout diverges from this document's own architecture sketch. | +| M10 Attribution | **Specified, not built** — [`m10-attribution.md`](m10-attribution.md). Scheduled ahead of M8 and M9 at the user's request, and the ordering is the point: M8 needs a fal.ai account and M9 needs GVC to exist, while this needs neither and gets more expensive every day it waits. Attribution captured at ingest costs a form field; attribution reconstructed later means opening every file by hand, and for material gathered from the open web the answer is often gone. Three decisions were taken with the user and are recorded with their rejected alternatives in that document: structured fields over a free-text credit, embedded metadata written directly while AI may only suggest, and clips inheriting from the parent with per-field override. | | **M8–M9** | **Not started.** Both need an external dependency M7 did not (fal.ai for M8, GVC itself for M9). | Non-milestone PRs, so a `git log` that does not match the table above still makes sense: @@ -371,24 +373,22 @@ from a decision, which is the distinction PR bodies do not preserve. `.odt`, `.ods`, `.odp` are in `DOCUMENT_EXTENSIONS` and nothing reads them. The refusal says which format and what to do about it, rather than failing vaguely. - A scan is refused too, and always will be without OCR, which is its own project. -- **Attribution is not modelled.** Raised by the user, recorded here so it is not - mistaken for a decision: an asset needs to carry where its content *came from* — a - website URL, the film or programme a clip is taken from, the news outlet and date of a - report — so it can be credited when it is used downstream in GVC. `Asset.source` - (`local_upload | url | ai_generated | gvc_export`) is **not** this: it records how the - file entered the library, not whose work it is. There is no field for a citation, no - UI for one, and nothing in M6-M9 that adds one. Deliberately undesigned so far; the - open questions are whether it is one free-text credit or structured fields, whether AI - enrichment may propose it from the content, and whether it belongs on the asset or on - each clip taken from it. -- **No tag-management screen.** `tagsApi.create`/`recategorise`/`updateCategory` exist, and - `rename`/`remove`/`createCategory`/`removeCategory` are wired into `stores/tags.ts` — all - reachable from no component. A user can create a tag by typing it and can then never - rename, delete or categorise it. The recursive-CTE category tree M3 built is, from the UI, - read-only. -- **Search ignores `asset_type` and `limit`.** The backend accepts both - (`routers/search.py:51-53`) and `api/search.ts` types them; `SearchView.tsx` passes - neither. Results cap at the server default of 30 with no way to page or filter by type. +- **~~Attribution is not modelled.~~** Specified as **M10** — see + [`m10-attribution.md`](m10-attribution.md). The three questions this entry used to + leave open are now answered, with the rejected alternatives recorded: structured + fields rather than one free-text credit, embedded file metadata written directly + while an AI may only *suggest* (grounded in a quote it must supply), and clips + inheriting from the parent at read time with per-field override. Not yet built. +- **~~No tag-management screen.~~** Built. `TagPanel.tsx` is mounted at + `views/SettingsView.tsx:513` and wires `rename`/`remove`/`recategorise` and the full + category CRUD. This entry outlived the component by several PRs, which is the failure + mode this list exists to prevent — check a claim here against the code before acting on + it. +- **~~Search ignores `asset_type` and `limit`.~~** Built. `SearchView.tsx` passes both, + with `TypeFilterChips` and a load-more that raises `limit` toward `MAX_LIMIT`, and both + persist in the URL so a result set can be linked to. There is still no `offset`, so + this is a growing page rather than true pagination — deliberate at this scale, and the + module comment says so. - **~~No `/a/{id}` deep link.~~** Built in `cdcabbf`, so GN-4's Notes→GAM reference now resolves. It rendered as a floating dialog over an empty shell until the panel chrome moved into `DetailDock`; it is a full-width page now.