From 9c5738bd1e4abc01f548c1888f266d78a6057cd6 Mon Sep 17 00:00:00 2001 From: MendixMau Date: Fri, 2 Oct 2026 12:53:45 +0800 Subject: [PATCH 1/2] new(design-artifacts, checkpoint-design): ideation library becomes an explicit opt-in question (D1) Step 0b had a paragraph pointing at the public ui-ux-pro-max-skill dataset for projects with no brand. In two full pipeline runs on the same procurement source neither run used it; one carried the source app's look and produced generic, dated wireframes, and the user picked the other run, which they had pushed to "more 2027". A passive paragraph is not an interview step. Now, with no client brand, Step 0b asks D1 as a choice: (a) fetch the dataset read-only at a pinned commit into a git-ignored .ideation/ and propose three directions, or (b) Atlas defaults; recommend (a) for demo/presales/POC and (b) for a like-for-like rebuild. Only a yes fetches. The brand question stays user-only, and CAC-4 no longer offers "use Atlas defaults" as its default. The step names the commands (the repo's own search.py, no toolkit wrapper per skills-over-scripts.md), what to keep and drop from its output, a contrast check, what never ports, and the Decisions row that records the direction. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + skills/checkpoints/checkpoint-design.md | 16 +++- skills/conversion-runbook.md | 2 +- skills/design-artifacts.md | 102 +++++++++++++++++++++--- 4 files changed, 106 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 30824c6c..7be57a74 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,6 +16,7 @@ three commits past it), and a bug report can name a release instead of a sha nob Sections dated before 2026-09-19 predate the cycle and stay as they are. ## Unreleased +- new(design-artifacts, checkpoint-design): **the ui-ux-pro-max ideation library is now an explicit, opt-in interview question (D1) instead of a paragraph nobody acted on.** With no client brand, Step 0b now asks "where does the look come from?" as a `choice`: (a) fetch the public MIT dataset read-only at a pinned commit (`09170ee`) into a git-ignored `.ideation/` and propose three directions, safe to bold, or (b) Atlas defaults; recommendation (a) for demo/presales/POC, (b) for a like-for-like rebuild. Only a yes fetches; on no there is no network call. The brand question itself stays `user-only` and CAC-4 no longer offers "use Atlas defaults" as its default answer. The step names the commands (the repo's own `search.py --design-system -f markdown`, no toolkit wrapper per `skills-over-scripts.md`), what to keep and drop from its output, a contrast check before presenting, what never ports (`--persist`, CSS/Tailwind/React, landing patterns, chart colours, spacing), and the Decisions row that records direction + commit + source rows. Evidence: two full pipeline runs on the same procurement source used neither the library nor a deliberate direction, and the user picked the run they had pushed to "more 2027"; a test on that app produced three directions from three queries in 0.10–0.23 s each, ~3.1 KB (~800 tokens) per direction, and one of them ported to a `:root` token block + class-only px CSS and rendered. "Atlas defaults are usually right" becomes "Atlas defaults for a like-for-like rebuild; a showcase direction for a demo". — MendixMau, from two cook-off pipeline runs - fix(bin/context-audit.sh): **a Read `offset` or `limit` stored as a string no longer crashes the audit, and the script now exits 0 as its header promises.** Older transcripts store these as strings, sometimes as junk like `'30, 90'`, and `offset - 1` raised a TypeError in the embedded reader, which stopped the whole run. They are now read as numbers when they parse and fall back to the Read tool's defaults (offset 1, limit 2000) when they do not; if the reader ever dies on an unseen transcript shape, the script says the numbers are partial and still exits 0. Field run: 1,300 sessions and 724 subagent runs on a Mac (2026-08-27 to 2026-09-30), which crashed on the old version. — MendixMau - fix(routing): **`learned-mcp-patterns.md` is no longer always-on in the build stage; it loads before the first MCP write in a session.** It sat in the Stage 5 baseline pack and in `mdl-agent`'s always-read rows, so every build session and every MDL helper agent carried ~4,900 tokens of MCP save/handoff rules and JSON payloads, including sessions that never open Studio Pro and cloud containers where MCP does not exist. Choosing the write mode is already Step 0 of `learned-mdl-preflight.md`, which stays always-on, so nothing is lost at the moment of choice; the MCP skill's trigger now names the moment it is needed (`mxcli --mcp` exec or a `pg_*`/`ped_*` call). Stage 5 pack: 74,055 → 71,443 words, 23 → 22 files; baseline 79,752 → 77,140 words. Found by the context report (`bin/context-audit.sh`, `bin/render-routing.sh --check`). — MendixMau - new(bin/context-audit.sh): **what fills the context window, per file, from the real Claude Code transcripts on this machine.** `token-burn.sh` says how many tokens a project burned; this says which files burned them, so decisions about splitting, trimming or un-routing a skill rest on measured runs instead of `wc` on the skill files. It reports: the context size before any work (first call, input + cache, for sessions and for subagents separately); the instruction files loaded every run (CLAUDE.md, CLAUDE.local.md) and their size; every file read (Read tool and simple shell reads like `cat`, `sed -n`, `git show REV:path`, following `cd` and `VAR=`) with reads, sessions, total and per-read size, and re-reads within a session (paging through a file is not a re-read; asking for the same part again is); other tool output by tool; and each compaction with the files read before it. Project names are masked by default (`project-1/architecture/modules/*.md`), so the output is safe to paste; `--names` shows them locally. First field numbers, from a captured pipeline-start subagent: it starts at 52,503 tokens before reading anything, then reads the runbook in 6 pages with 2 repeats. Fixture: `tests/wave2/test-context-audit.sh` over a scrubbed real capture. — MendixMau diff --git a/skills/checkpoints/checkpoint-design.md b/skills/checkpoints/checkpoint-design.md index f403b4f2..7e2b0852 100644 --- a/skills/checkpoints/checkpoint-design.md +++ b/skills/checkpoints/checkpoint-design.md @@ -26,7 +26,9 @@ generic Atlas output that will need rework. Design direction is the checkpoint users most often want to *talk about*, not pick from a list. Open with a divergent conversation: show references (the source app's look, the client's brand, -2–3 mood directions described concretely or as quick HTML swatches) and discuss — what should +2–3 mood directions described concretely or as quick HTML swatches — with no client brand, +take them from the ideation library when the user says yes to it, `design-artifacts.md` Step 0b +question D1) and discuss — what should this app *feel* like, what must carry over from the brand, what should deliberately change? If the user says anything like "let's ideate on the design" at any point, this conversation IS the response — do not generate the design system or launch agents until it converges and the @@ -76,7 +78,12 @@ or desktop-first. Use that to set the recommended option. > "Do you have branding guidelines, a Figma file, or an existing design system? > > Drop a link, paste the key details (colors, fonts, logo), or describe the visual direction. -> If none — say 'use Atlas defaults' and we'll proceed with out-of-the-box Atlas styling." +> If there is none, say so — the next question is where the look comes from." + +This question is `user-only`: no default answer is offered in it. A "none" leads to +`design-artifacts.md` Step 0b question D1 (ideation-library directions vs Atlas defaults), a +`choice` with its own recommendation — Atlas defaults are one answer to D1, not the default +answer to this question. **What to do with the answer:** @@ -85,7 +92,7 @@ or desktop-first. Use that to set the recommended option. | Figma link | Add to `PROJECT.md` → `## Decisions` → `Design assets:`. Note which flows are designed vs wireframe-only. | | Brand doc / PDF | Same as above. Extract: primary color, font family, logo usage rules. | | Verbal description | Record key decisions (primary color, font, tone) in `design-artifacts.md` inputs. | -| "Atlas defaults" | Note it. No custom tokens needed. Skip Atlas customization in MDL layer. | +| "None" | Ask D1 (`design-artifacts.md` Step 0b). On "Atlas defaults": note it, no custom tokens, skip Atlas customization in the MDL layer. On "ideation library": fetch at the pinned commit, propose three directions, record the chosen one. | | Existing Mendix design system | Identify the theme module. Use its layout + widget naming conventions throughout. | --- @@ -96,5 +103,6 @@ or desktop-first. Use that to set the recommended option. PROJECT.md → ## Decisions: Atlas layout: [chosen layout] Platform target: [web / responsive / native / both] - Design assets: [Figma URL / brand doc link / 'Atlas defaults' / description] + Design assets: [Figma URL / brand doc link / 'none' / description] + Design direction: [brand / Atlas defaults / — ui-ux-pro-max @ + source rows] ``` diff --git a/skills/conversion-runbook.md b/skills/conversion-runbook.md index 9eb993eb..768861fe 100644 --- a/skills/conversion-runbook.md +++ b/skills/conversion-runbook.md @@ -666,7 +666,7 @@ The biggest gap before this runbook existed. Module boundaries, wiring diagrams | | | |---|---| -| **User defines** | ① One Mendix app or several (if flagged at Stage 0). ② **Module boundaries** (agent proposes with `modularize-domain.md` criteria). ③ **Buy vs build vs stub, per fit-gap item** — the confirming step `brd-to-build-plan.md` assumed already happened. ④ **Target security / role model** — not just whether auth existed in the source, but what the target should be. ⑤ **Data volumes, concurrency, NFRs** — these decide indexing, pagination, datagrid-vs-paged-gallery, loop batch sizes. ⑥ **Integration contracts** — real or stub, endpoint, credentials, owner, test environment. ⑦ **Branding inputs** — logo, palette, type, spacing, per `design-artifacts.md`. | +| **User defines** | ① One Mendix app or several (if flagged at Stage 0). ② **Module boundaries** (agent proposes with `modularize-domain.md` criteria). ③ **Buy vs build vs stub, per fit-gap item** — the confirming step `brd-to-build-plan.md` assumed already happened. ④ **Target security / role model** — not just whether auth existed in the source, but what the target should be. ⑤ **Data volumes, concurrency, NFRs** — these decide indexing, pagination, datagrid-vs-paged-gallery, loop batch sizes. ⑥ **Integration contracts** — real or stub, endpoint, credentials, owner, test environment. ⑦ **Branding inputs** — logo, palette, type, spacing, per `design-artifacts.md`; with no brand, the design direction (Step 0b question D1: ideation-library directions or Atlas defaults). | | **Agent produces** | `.mx-brd.json`, `architecture/` (module defs, layer diagram, wiring diagram, `fit-gap.md`, `blueprint.html` checkpoint render — plus a workflow diagram and/or agent-wiring diagram in `blueprint.md`/`blueprint.html` when CAC-3's Q3 flags real scope for either, plus cross-persona journey diagrams + journey list — `architecture-blueprint.md` Step 3d — whenever the BRDs carry more than one persona; single-persona skip recorded as a one-line note, never silent), `design/` per `design-artifacts.md`'s full output list: `ds.css` + `design-system.html` + **`wireframes/*.html`, one annotated wireframe per screen** — the design system without the wireframes is half the deliverable and fails the gate. | | **Surface** | `module-design.html` · `architecture/blueprint.html` (generated render of `blueprint.md` — architecture-blueprint.md Step 7, never hand-edited) · `design-system.html` + `wireframes/*.html` | | **Gate ✋** | Boundaries approved. Marketplace calls made. Role model, volumes, integrations and branding **each asked and answered**: `CONFIRMED`, or explicitly delegated by the user ("you decide" → `ASSUMED` with risk). Never `ASSUMED` without the question having reached the user. Close-out block pasted before the ✋ decision is asked (§1b rule 7) — this stage has no closing checkpoint, so the gate is where the Stage-4 open (its approach, skills and optional artifacts) is shown. **No architecture/design artifact is produced before its checkpoint ran.** | diff --git a/skills/design-artifacts.md b/skills/design-artifacts.md index ac2c9deb..8952c7f8 100644 --- a/skills/design-artifacts.md +++ b/skills/design-artifacts.md @@ -66,19 +66,101 @@ Branding is an input, not an afterthought — and not a checkbox to tick silentl | Basis | When | Effort | |---|---|---| | **Client branding guidelines** (logo, palette, type, spacing) | Real project — request them as an analysis deliverable | depends | -| **Atlas defaults** | POC / no brand yet — matches the actual build target 1:1 | zero | +| **Atlas defaults** | No brand, like-for-like rebuild — matches the actual build target 1:1 | zero | +| **Ideation-library direction** (ui-ux-pro-max, question D1 below) | No brand, demo / presales / POC — a deliberate, current look | low | | **Neutral placeholder palette** (the `dataviz` reference palette) | Design-forward but brand-agnostic, swap later | low | -For a faithful rebuild POC, **Atlas defaults** are usually right — the wireframes then look like what Mendix will actually render, so coverage comparison is like-for-like. Record the choice; it cascades into every token below. +**Atlas defaults fit a like-for-like rebuild, not a demo.** When coverage comparison against the +source is the point, Atlas defaults are right: the wireframes look like what Mendix will render. +For a demo, presales or POC build, prefer a deliberate **showcase direction** instead — fidelity +covers behaviour, not looks. Measured on two runs of this pipeline over the same procurement +source (2026-10): the run that carried the source app's own look produced clean, generic, +dated wireframes; the run where the user demanded "more 2027, slick, modern" is the one they +picked side by side. Neither run used the ideation library below, which was then a passive +paragraph here. Record the choice; it cascades into every token below. + +### Two questions, asked in order — never merged into one + +1. **"Is there a brand guideline, a style guide or an existing design system?"** — `user-only` + (`interview-protocol.md`): asked on its own, with **no recommendation**, never as a line in a + batch. A yes ends here: the brand is the basis. +2. **If there is no brand: "Where does the look come from?"** — a `choice`, asked in the Step 0b + batch next to the navigation-layout question below, in the two-options-plus-recommendation + shape: + +> **D1 — Design direction (no client brand)** +> There is no brand guideline, so the look is ours to choose. Purpose on record: * Stage P purpose row from `PROJECT.md`>*. +> - **(a) Three directions from the ui-ux-pro-max ideation library.** I fetch a public, MIT-licensed +> design dataset (88 styles, 192 palettes, 74 font pairings, 119 UX rules) read-only, at a +> pinned commit, into `.ideation/` (git-ignored, ~30 MB, ~2 s). I search it for this domain and +> come back with three directions, safe to bold, as the next question. One more question; a +> deliberate, current look. +> - **(b) Atlas defaults.** No fetch, zero effort, and the wireframes look exactly like what +> Mendix renders, so coverage comparison is like-for-like. It also looks like every other +> Mendix app. +> +> **I recommend (a)** for a demo, presales or POC build; **(b)** when the purpose is a faithful +> rebuild judged on coverage. + +"Something else" is always open: a direction the user describes, or the neutral placeholder +palette. **Only a yes to (a) triggers the fetch. On (b) or anything else, nothing is fetched and +no network call is made.** + +### On (a): fetch, search, propose three + +The dataset is the fact source and its own search CLI is the instrument (Python 3 standard +library only, BM25 over the CSVs) — there is no toolkit wrapper. Fetch exactly the pinned commit, +from the project root: + +```bash +PIN=09170eec67eefd46a7ae85de61b40c194020f997 # ui-ux-pro-max-skill master, 2026-09-27 +git init -q .ideation/ui-ux-pro-max +git -C .ideation/ui-ux-pro-max fetch -q --depth 1 https://github.com/nextlevelbuilder/ui-ux-pro-max-skill "$PIN" +git -C .ideation/ui-ux-pro-max -c advice.detachedHead=false checkout -q FETCH_HEAD +grep -qxF '.ideation/' .gitignore 2>/dev/null || echo '.ideation/' >> .gitignore +cd .ideation/ui-ux-pro-max/src/ui-ux-pro-max/scripts +python3 search.py "" --design-system -p "" -f markdown # one bundle per seed +python3 search.py "" --domain ux -n 4 # UX rules for the key screens +``` -**No brand and no opinion? Borrow the ideation database.** When the client has neither guidelines -nor a preference, the public [ui-ux-pro-max-skill](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill) -repo (MIT) is a good input for this interview: clone it read-only and search its -`src/ui-ux-pro-max/data/` CSVs (styles, color palettes, font pairings, UX guidelines) for 2–3 -directions that fit the client's domain, then present those as the interview options. **Reuse -stops at the data.** Its implementation guidance targets CSS frameworks (React, Tailwind, etc.), -not Mendix — the chosen direction lands as `ds.css` tokens, the Atlas mapping table, and -StyleGallery choices per `learned-stylegallery.md`, never as its CSS or component code. +Moving the pin is a toolkit PR that changes this one line, never a per-project `git pull`. + +Then the judgement, which is yours: + +1. **Map the domain to rows yourself.** The dataset has no row for most enterprise back-office + domains (procurement, ERP, claims, case handling): a bare domain query returns noise (a + procurement query ranked "Food Delivery" second). Pick three seed queries that span safe to + bold: the nearest product row (e.g. `invoice billing finance back-office enterprise`), the + AI/agent row when the app has an agent (`AI agent copilot automation platform SaaS`), and a + bolder option (`financial dashboard dark data-dense analytics`, or a bento/soft-UI style). +2. **Run `--design-system -f markdown` once per seed.** Each is ~3 KB (~800 tokens) and returns in + under 0.3 s, so three directions cost ~2.5k tokens — cheap enough to run inside the interview. + Keep the Style name, the Colors table, the Typography pair and the Avoid list. Drop the + Pattern section (landing-page CTA advice such as "Contact Sales", wrong for a back-office + app), the font `@import` URLs (the font goes into the theme) and the pre-delivery checklist + (web/Tailwind items). +3. **Check every on-colour pair before you present it.** The palettes are mostly AA-adjusted + already, but a dark row can make the primary vanish: `#0F172A` on `#020617` is 1.13:1, so that + direction must use its accent as the action colour. Say so in the option, do not fix it + silently. +4. **Present the three as the follow-up question** (D1a), each with a name, a one-line feel, + five hex swatches, the font pair, one signature idea (what the agent or decision screen does + differently), and the source rows. Recommend one, with the reason. If the user wants to see + them, render quick HTML swatches — that is the CAC-4 brainstorm, not a design system yet. + +**Reuse stops at the data.** The chosen direction lands as `ds.css` tokens (three tiers per +`learned-stylegallery.md`), the Atlas mapping table and StyleGallery choices — class-only CSS in +px. What never ports: `search.py --persist` (it writes `design-system//MASTER.md`, a second +design spec that would compete with `ds.css`); any CSS, Tailwind, React or `--stack` guidance; +the landing-page patterns; its chart colours (chart series come from the `dataviz` palette); +its spacing scale (the toolkit's is `design-spacing.md`, 8/16/24/32/48). + +**Record it** as a Stage 3 row in the `PROJECT.md` Decisions table, with the commit and rows so +the direction can be reproduced: +`| 3 | Design direction: — ui-ux-pro-max @ 09170ee; colors.csv "", styles.csv "