Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ three commits past it), and a bug report can name a release instead of a sha nob
Sections dated before 2026-09-19 predate the cycle and stay as they are.

## Unreleased
- fix(project-bin/page-fidelity.js, check-design-reaches-app.sh): **the page score no longer reports a confident number it never measured, and it says plainly that it is not a LOOK.** A wireframe template whose opening HTML comment mentions `<main>` made the scorer take the comment as the page, so headings, actions, content and classes all compared 0 of 0 and only bindings were scored: two field builds on one template read "fidelity 100%" and "50%", called it done and never opened a screenshot. Now: HTML comments are stripped before parsing; a wireframe that yields nothing to compare prints `UNMEASURED` with the reason and exits 3 without logging; a page whose model still carries a `Stub:` caption, scored without `--stub`, prints `STUB IN MODEL`, logs 0% and exits 4; `snippetcall`s are followed into snippets found in the input (depth 4), and a page whose snippets are missing is scored `(partial)` with the `describe snippet` command to fetch them; the printed label is `text-match` (file and TSV names unchanged), and every run ends with "text-match compares identifiers, not pixels — it is not a LOOK". `check-design-reaches-app.sh` also reads the Docker build's `.docker/build/app/web/theme.compiled.css` when it is the newest built theme. Skills: `ui-loop.md` step 4 and `iterative-build-loop.md` (mark a forward-reference stub with a `Stub:` caption). Fixtures: `test-page-fidelity-mocks.sh` OUT6–OUT11 over a genericized reduction of the field template (`CAPTURE.md` records the before/after). — field report from two builds on one wireframe template (#188)
- fix(bin/context-audit.sh): **a Read `offset` or `limit` stored as a string no longer crashes the audit, and the script now exits 0 as its header promises.** Older transcripts store these as strings, sometimes as junk like `'30, 90'`, and `offset - 1` raised a TypeError in the embedded reader, which stopped the whole run. They are now read as numbers when they parse and fall back to the Read tool's defaults (offset 1, limit 2000) when they do not; if the reader ever dies on an unseen transcript shape, the script says the numbers are partial and still exits 0. Field run: 1,300 sessions and 724 subagent runs on a Mac (2026-08-27 to 2026-09-30), which crashed on the old version. — MendixMau
- fix(routing): **`learned-mcp-patterns.md` is no longer always-on in the build stage; it loads before the first MCP write in a session.** It sat in the Stage 5 baseline pack and in `mdl-agent`'s always-read rows, so every build session and every MDL helper agent carried ~4,900 tokens of MCP save/handoff rules and JSON payloads, including sessions that never open Studio Pro and cloud containers where MCP does not exist. Choosing the write mode is already Step 0 of `learned-mdl-preflight.md`, which stays always-on, so nothing is lost at the moment of choice; the MCP skill's trigger now names the moment it is needed (`mxcli --mcp` exec or a `pg_*`/`ped_*` call). Stage 5 pack: 74,055 → 71,443 words, 23 → 22 files; baseline 79,752 → 77,140 words. Found by the context report (`bin/context-audit.sh`, `bin/render-routing.sh --check`). — MendixMau
- new(bin/context-audit.sh): **what fills the context window, per file, from the real Claude Code transcripts on this machine.** `token-burn.sh` says how many tokens a project burned; this says which files burned them, so decisions about splitting, trimming or un-routing a skill rest on measured runs instead of `wc` on the skill files. It reports: the context size before any work (first call, input + cache, for sessions and for subagents separately); the instruction files loaded every run (CLAUDE.md, CLAUDE.local.md) and their size; every file read (Read tool and simple shell reads like `cat`, `sed -n`, `git show REV:path`, following `cd` and `VAR=`) with reads, sessions, total and per-read size, and re-reads within a session (paging through a file is not a re-read; asking for the same part again is); other tool output by tool; and each compaction with the files read before it. Project names are masked by default (`project-1/architecture/modules/*.md`), so the output is safe to paste; `--names` shows them locally. First field numbers, from a captured pipeline-start subagent: it starts at 52,503 tokens before reading anything, then reads the runbook in 6 pages with 2 repeats. Fixture: `tests/wave2/test-context-audit.sh` over a scrubbed real capture. — MendixMau
Expand Down
13 changes: 11 additions & 2 deletions project-bin/check-design-reaches-app.sh
Original file line number Diff line number Diff line change
Expand Up @@ -126,10 +126,18 @@ if [ -z "$DESIGN" ]; then
[ -f "$c" ] && { DESIGN="$(rel "$c")"; break; }
done
fi
# A `mxcli docker build` writes its compiled theme under .docker/build/, not deployment/, so a
# project that only ever ran the docker stack had no BUILT stylesheet here and this check
# exited 2 on an app that was up and serving the theme (field report, 2026-10-02). When both
# trees exist, the NEWER file wins — it is the one the running app was built from.
BUILT_ABS=""
if [ -z "$BUILT" ]; then
for c in "$MODEL_DIR/deployment/web/theme.compiled.css" "$MODEL_DIR/deployment/web/theme.css"; do
[ -f "$c" ] && { BUILT="$(rel "$c")"; break; }
[ -f "$c" ] && { BUILT_ABS="$c"; break; }
done
c="$MODEL_DIR/.docker/build/app/web/theme.compiled.css"
if [ -f "$c" ] && { [ -z "$BUILT_ABS" ] || [ "$c" -nt "$BUILT_ABS" ]; }; then BUILT_ABS="$c"; fi
[ -n "$BUILT_ABS" ] && BUILT="$(rel "$BUILT_ABS")"
fi

# The framework's own customization surface. Atlas has called this file the same thing since
Expand All @@ -145,7 +153,8 @@ if [ -z "$DESIGN" ] || [ ! -f "$DESIGN" ]; then
fi
if [ -z "$BUILT" ] || [ ! -f "$BUILT" ]; then
printf 'check-design-reaches-app: no BUILT stylesheet found.\n' >&2
printf ' Searched: deployment/web/theme.compiled.css, deployment/web/theme.css under %s\n' "$MODEL_DIR" >&2
printf ' Searched: deployment/web/theme.compiled.css, deployment/web/theme.css,\n' >&2
printf ' .docker/build/app/web/theme.compiled.css under %s\n' "$MODEL_DIR" >&2
printf ' Run a build first. "Did the design system reach the app" cannot be answered from\n' >&2
printf ' source, and answering it from source is how this defect shipped.\n' >&2
printf ' This is NOT a pass.\n' >&2
Expand Down
122 changes: 112 additions & 10 deletions project-bin/page-fidelity.js
Original file line number Diff line number Diff line change
Expand Up @@ -18,8 +18,20 @@
//
// Scoring: headings 25%, action labels 30%, content blocks 25%, structural classes 20%,
// bindings 25% (normalized over the dimensions the wireframe actually uses — one it
// doesn't use is dropped from the denominator). Prints per-dimension hits and every miss;
// exits 0 always — it is an instrument, not a gate. The gate is check-page-shell.sh.
// doesn't use is dropped from the denominator). Prints per-dimension hits and every miss.
// It is an instrument, not a gate (the gate is check-page-shell.sh), but it does not
// print a number it did not measure. Exit codes:
// 0 scored (a low score is still 0 — the number is the verdict)
// 2 usage / no such wireframe / no declaration of the page in the input
// 3 UNMEASURED — nothing to compare; printed with the reason, NOT logged
// 4 STUB IN MODEL — a `Stub:` caption scored without --stub; logged as 0%
//
// THE LABEL IS `text-match`, NOT `fidelity` (renamed 2026-10-02). Two field builds on one
// template read "fidelity 100%" as "the page looks like the wireframe" and never opened a
// screenshot. The number is identifier overlap between two texts; it cannot see placement,
// nesting, size or colour. The LOOK (skills/ui-loop.md) is the fidelity check; this is its
// cheap pre-filter. The file and TSV keep their names so existing projects and the
// fidelity obligation keep working.
//
// Third field run (ToeicBuddy Reading_Part, 2026-08-27) added bind-table awareness:
// on a data-heavy page nearly all visible copy is BOUND (passages, stems, options),
Expand Down Expand Up @@ -294,6 +306,13 @@ function bindRows(html) {
}

function wfFacts(html) {
// HTML COMMENTS ARE NOT PAGE. contentOf() picks its boundary with a regex, and a template
// whose header comment says "put the screen inside <main> … the bind table goes AFTER
// </main>" handed it the comment text as the page. Field case, 2026-10-02, two builds
// on one template: every page scored headings/actions/content/classes 0/0, so the only
// dimension left was bindings — "100%" on a page with 1 of 1 bindings, "0%" on another,
// neither describing the page. Strip comments before anything reads the document.
html = html.replace(/<!--[\s\S]*?-->/g, ' ');
const mock = localMockClasses(html);
let main = contentOf(html);
const mockUsed = [];
Expand Down Expand Up @@ -378,10 +397,10 @@ let MODULE = null;
// script fixed, so it does not; instead the binding script is passed as another input
// (`… - path/to/alter.mdl < describe.mdl`) and its ALTER body counts. An ALTER-only input
// is not a page: without a CREATE the run still exits 2.
const SRCS = MDLS.map(f => f === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(f, 'utf8'));
function pageMdl() {
let out = '', alters = '';
for (const f of MDLS) {
const src = f === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(f, 'utf8');
for (const src of SRCS) {
const re = new RegExp(
'CREATE(\\s+OR\\s+(MODIFY|REPLACE))?\\s+PAGE\\s+"?([A-Za-z0-9_]+)"?\\."?' + PAGE + '"?\\b', 'gi');
for (const m of src.matchAll(re)) {
Expand All @@ -394,6 +413,50 @@ function pageMdl() {
return out && out + alters;
}

// SNIPPETS ARE PAGE. A page assembled from snippet calls carries its headings, buttons and
// classes in the snippet bodies, so scoring the page body alone marks all of it missing —
// and the cheapest way to raise that number is to inline the text into the page, which
// undoes the reuse the snippets were for. Snippet bodies found in ANY input (the build
// script, or `mxcli describe snippet` output passed as another file) are added to the
// corpus, nested calls included. A call whose body is not in the input is NAMED, and the
// run is marked partial — the number is then a floor, and says so.
const SNIP_CALL = /\bsnippetcall\s+"?[A-Za-z0-9_]+"?\s*\(\s*snippet\s*:\s*"?([A-Za-z0-9_]+)"?\."?([A-Za-z0-9_]+)"?/gi;
function expandSnippets(body) {
let extra = '';
const seen = new Set(), missing = [];
let frontier = body;
for (let depth = 0; depth < 4 && frontier; depth++) {
let next = '';
for (const m of frontier.matchAll(SNIP_CALL)) {
const q = m[1] + '.' + m[2];
if (seen.has(q)) continue;
seen.add(q);
const re = new RegExp('CREATE(\\s+OR\\s+(MODIFY|REPLACE))?\\s+SNIPPET\\s+"?' + m[1] + '"?\\."?' + m[2] + '"?\\b', 'i');
let found = null;
for (const src of SRCS) { const d = src.match(re); if (d) { found = pageBody(src, d.index); break; } }
if (found) { extra += found + '\n'; next += found + '\n'; } else missing.push(q);
}
frontier = next;
}
return { extra, missing };
}

// The page's OWN widgets, with its declaration header and snippet calls removed: if nothing
// text-bearing is left and a called snippet is missing, there is nothing to score at all.
function ownsContent(body) {
const open = body.indexOf('{');
const inner = (open < 0 ? '' : body.slice(open)).replace(new RegExp(SNIP_CALL.source + '[^)]*\\)', 'gi'), ' ');
return /'[^']+'|\b(attribute|datasource)\s*:/i.test(inner);
}

// A FORWARD-REFERENCE STUB IN THE MODEL. iterative-build-loop.md § Forward references marks a
// stub page with a caption that starts `Stub:` or `Stub -`. Scored WITHOUT --stub, such a
// page is being reported as the build — field case, 2026-10-02: a stub answer page scored
// 100% on its one binding and was counted as done. Scored as what it is: 0%, logged
// `stub-in-model`, exit 4. With --stub it is the declared forward reference and scores
// normally, exempt from the target as before.
const STUB_MARK = /'\s*stub\s*(:|-|\u2013|\u2014)/i;

// Known limit: the counter does not skip quoted strings, so an UNBALANCED brace inside a
// string literal truncates the body early (balanced placeholders like '{1} of {2}' are fine).
// The page body, from its declaration to the brace that MATCHES the body's opening
Expand Down Expand Up @@ -466,12 +529,43 @@ function score(wf, mdl) {
}

const wf = wfFacts(readWireframe());
const mdl = pageMdl();
if (!mdl.trim()) { console.error('page-fidelity: no declaration of page "' + PAGE + '" found in input'); process.exit(2); }
const ownMdl = pageMdl();
if (!ownMdl.trim()) { console.error('page-fidelity: no declaration of page "' + PAGE + '" found in input'); process.exit(2); }
const snips = expandSnippets(ownMdl);
const mdl = ownMdl + snips.extra;
const s = score(wf, mdl);
console.log(PAGE + ' fidelity ' + s.pct + '% headings ' + s.h.ok + '/' + s.h.n +
' actions ' + s.b.ok + '/' + s.b.n + ' content ' + s.k.ok + '/' + s.k.n + ' classes ' + s.c.ok + '/' + s.c.n +
(s.bd.n ? ' bindings ' + s.bd.ok + '/' + s.bd.n : ''));
const dims = ' headings ' + s.h.ok + '/' + s.h.n + ' actions ' + s.b.ok + '/' + s.b.n +
' content ' + s.k.ok + '/' + s.k.n + ' classes ' + s.c.ok + '/' + s.c.n +
(s.bd.n ? ' bindings ' + s.bd.ok + '/' + s.bd.n : '');
const snipNote = () => console.log(' snippets not in input (their content is not scored): ' + snips.missing.join(' ') +
'\n pass their bodies as more inputs: ./mxcli -p <app>.mpr -c "describe snippet <Mod.Name>" > snip.mdl');

// UNMEASURED IS NOT A SCORE. Two ways to have nothing to compare, both reported and neither
// logged, because a logged row is what the fidelity obligation counts as the page measured:
// * the wireframe yields no heading, action, content block or class — the content
// boundary landed on the wrong element (the comment case above, a missing <main>), so a
// number off bindings alone would be a number about the annotation table;
// * the page's own body has no text-bearing widget and the snippets it calls are not in
// the input — the page is somewhere this run cannot see.
// Exit 3, so a loop that scores pages cannot read either case as a pass.
const unmeasured = !(s.h.n + s.b.n + s.k.n + s.c.n)
? 'the wireframe gave no headings, actions, content or classes to compare' +
(s.bd.n ? ' (only ' + s.bd.n + ' binding row(s))' : '') +
' — its content boundary is likely the wrong element; open the wireframe and check where <main> / .wf-screen is'
: (snips.missing.length && !snips.extra && !ownsContent(ownMdl))
? 'the page is built from snippet calls whose bodies are not in the input' : null;
if (unmeasured) {
console.log(PAGE + ' text-match UNMEASURED — ' + unmeasured + dims);
if (snips.missing.length) snipNote();
console.log(' not logged: an unmeasured page has no score of record');
process.exit(3);
}
const stubHit = !STUB && (ownMdl.match(STUB_MARK) || [])[0];
if (stubHit) s.pct = 0;
console.log(PAGE + ' text-match ' + s.pct + '%' + dims + (snips.missing.length ? ' (partial)' : ''));
if (stubHit)
console.log(' STUB IN MODEL: a caption starts ' + stubHit.trim() + "' — this page is a forward-reference stub," +
' not the build. Scored 0%. Build the page, or score it with --stub while it is still a stub.');
const miss = [...s.h.miss.map(x => 'heading: ' + x), ...s.b.miss.map(x => 'action: ' + x),
...s.k.miss.map(x => 'content: ' + x.slice(0, 78)),
...(s.c.miss.length ? ['classes: ' + s.c.miss.join(' ')] : []),
Expand All @@ -488,6 +582,11 @@ if (wf.mockUsed.length) console.log(' bound-data mocks (wireframe-local, text n
if (wf.structural.length) console.log(' wireframe structure (kept as page content, not a mock): ' + wf.structural.join(' '));
if (wf.mockCls.length) console.log(' wireframe-local classes (not scored): ' + wf.mockCls.join(' '));
if (STUB) console.log(' stub — forward-reference target, exempt from the 80% target; the first non-stub row is the score of record');
if (snips.missing.length) snipNote();
// The number above is identifier overlap between two texts. It cannot see placement,
// nesting, size or colour, so it is not a LOOK and never stands in for one.
console.log(' text-match compares identifiers, not pixels — it is not a LOOK. Screenshot the page and' +
' check it against the wireframe (skills/ui-loop.md).');

// ---- record the run (see header: EVERY RUN IS RECORDED) -------------------------------
if (!NOLOG) {
Expand Down Expand Up @@ -564,7 +663,9 @@ if (!NOLOG) {
new Date().toISOString().slice(0, 16).replace('T', ' '),
PAGE, MODULE || '-', s.pct === null ? '-' : s.pct + '%',
frac(s.h), frac(s.b), frac(s.k), frac(s.c), frac(s.bd),
STUB ? 'stub' : MDLS.includes('-') ? (MDLS.length > 1 ? 'describe+script' : 'describe') : 'draft',
(STUB ? 'stub' : stubHit ? 'stub-in-model'
: MDLS.includes('-') ? (MDLS.length > 1 ? 'describe+script' : 'describe') : 'draft') +
(snips.missing.length ? '+partial' : ''),
path.relative(root, path.resolve(WF_FILE)) + (WF_REF ? '#/' + WF_REF.route : ''),
].join('\t');
fs.appendFileSync(tsv, row + '\n');
Expand All @@ -574,3 +675,4 @@ if (!NOLOG) {
console.error('page-fidelity: score NOT logged (' + e.message + ')');
}
}
if (stubHit) process.exit(4);
7 changes: 7 additions & 0 deletions skills/iterative-build-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -609,6 +609,13 @@ module close, any page whose *only* rows are `stub` rows is an unfinished page w
stub label — the same finding as no row at all (`ui-preflight-pages.md` Step 5,
`module-review.md` rubric row 6).

**Mark the stub in the model, not only in the script name.** Give the stub page one text widget
whose caption starts `Stub:` (e.g. `Stub: built in 20-order-inbox.mdl`). `page-fidelity.js`
reads that marker: scored **without** `--stub`, a page carrying it is reported as
`STUB IN MODEL`, logged at 0% and exits 4 — so a stub that was never replaced cannot pass as
the build (field case, 2026-10-02: a stub answer page printed 100% off its one binding). The
real page script drops the caption with the rest of the stub.

---

## CE Error Triage
Expand Down
6 changes: 5 additions & 1 deletion skills/ui-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,11 @@ After a script that creates or changes a page:
4. **Score it when a wireframe exists** — `node project-bin/page-fidelity.js` for the page. The
row it appends to `docs/PAGE-FIDELITY.tsv` is the score of record (target ≥80%) and the
`fidelity` obligation reads it; a page nobody scored is a page nobody checked, however the
screenshot looked. No wireframe: waive it explicitly (`--waive fidelity/<Module> --reason
screenshot looked. The number prints as `text-match`: identifier overlap, blind to placement,
nesting, size and colour. **It never replaces step 3** — a field build that ran it 38 times
and never opened a screenshot shipped pages whose text matched and whose layout did not.
`UNMEASURED` (exit 3) means there was nothing to compare — fix the input, it is not a score;
`STUB IN MODEL` (exit 4) means a `Stub:` page is being scored as the build. No wireframe: waive it explicitly (`--waive fidelity/<Module> --reason
"no wireframe"`) — there is no automatic discharge, and a module the obligation never hears
about stays PENDING forever.
5. **Fix it now, or write it down now.** A defect that survives into the next script costs more to
Expand Down
Loading
Loading