Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions .agents/skills/capture-ui-evidence/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
---
name: capture-ui-evidence
description: Use the standard Jones Code UI verification and evidence workflow for UI changes and screenshots. Route interactive shared-renderer checks to T3 Browser, native Linux desktop scenarios to the Electron harness after exact-revision qualification, and mobile checks to test-t3-mobile.
---

# Capture Jones Code UI evidence

This is the default workflow for verifying user-visible UI changes and recording evidence. Choose the affected client:

- Shared web/desktop renderer: [test-t3-app](../test-t3-app/SKILL.md) and T3's Browser panel for live inspection, clicks, typing, screenshots and iteration.
- Linux native desktop shell: the Electron runner below for repeatable screenshots and scripted interactions after exact-revision qualification.
- Native mobile: [test-t3-mobile](../test-t3-mobile/SKILL.md).

Exercise the changed flow, assert its observable result, and check backend readback or reload persistence when the change depends on them. Use the same evidence format on every route: full candidate/comparison revisions (plus patch identity for dirty source), build correspondence, client/route, actual viewport/scale/theme, meaningful fixture, action and readback, captures, and coverage limits. The source/build distinction in [runner output and evidence limits](references/runner-output-and-limits.md) applies to Browser evidence too; a dev source OID alone is not an attestation of the served bundle.

T3 Browser remains the shared-renderer route. An unavailable Preview does not authorize switching to a standalone browser or using Electron as a Browser fallback.

The Linux Electron runner is a separate route for behavior that depends on the desktop shell. Require a recorded passing qualification for the exact harness revision; exploratory output alone leaves Electron-specific behavior unverified. This nightly source port does not qualify a host or adopt an installed wrapper. A separately installed host command must remain pinned to its qualified revision, with readback in `~/.local/state/jones-code-ui-evidence/qualification.json`; an older source qualification does not qualify this port.

The runner launches existing built outputs; it does not build the app. On a host with the installed command, first run `jones-code-ui-evidence doctor`, then use `jones-code-ui-evidence run --source /absolute/candidate --build-receipt /absolute/candidate-receipt.json` with a scenario that exercises the changed flow. The default `sidebar-rename` is a qualification/smoke scenario; it does not verify an unrelated UI change. A custom scenario receives a Playwright page and Electron app for inspecting and interacting with the candidate inside the isolated run.

From the Jones Code repository root, use:

```sh
node apps/desktop/scripts/ui-evidence.mjs doctor [--source DIR]
node apps/desktop/scripts/ui-evidence.mjs setup [--source DIR]
node apps/desktop/scripts/ui-evidence.mjs run [--source DIR] [--scenario sidebar-rename|/absolute/path/scenario.mjs] [--size 1280x800] [--scale 1] [--theme system|light|dark] [--comparison OID] [--build-receipt FILE] [--out DIR]
node apps/desktop/scripts/ui-evidence.mjs cleanup --run <run-id>
```

The repository also exposes the CLI and its bounded test group as package scripts `ui:evidence` and `test:ui-evidence`, invoked through the existing `vp run <script>` pattern. The test group does not qualify the live Electron route.

`doctor` is read-only and reports prerequisites or exact missing system package names; it installs nothing. `setup` downloads the lockfile-matched Electron runtime into its owned user cache and checks the official checksum. No command elevates privileges. `run` requires an existing built app, defaults to the harness source and `sidebar-rename`, and writes to `<source>/.t3/ui-evidence/runs/<run-id>/` unless `--out` names a path that does not already exist. A full commit OID may be supplied to `--comparison`; `--build-receipt FILE` records a receipt for the exact source and output hashes. Every command emits JSON with an `outcome` of `complete`, `no_change`, `rejected`, `partial`, or `unknown`. A run returns 0 for a pass, 1 for a scenario failure, 2 for a rejected precondition, or 130/143 for cancellation.

Read [runner output and evidence limits](references/runner-output-and-limits.md) before interpreting captures, comparing revisions, or writing a PR claim. Do not weaken isolation or retry a rejected probe by disabling Chromium's sandbox.
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# Runner output and evidence limits

## Run lifecycle

The runner uses Xvfb and launches the candidate's built Electron app inside a rootless Bubblewrap sandbox with private namespaces. The sandbox receives the staged app/build inputs and an allowlisted environment, not the real home, `.git`, dotenv files, provider credentials, or host/live network access; its backend uses a private loopback endpoint. It verifies its isolation before app launch and records the probe and resolved backend topology in the manifest. A failed probe rejects the run before launch; report that result and stop. Chromium receives a recorded CDP navigator-online marker so the real renderer attempts its isolated loopback connection; this does not add network interfaces or permit external access.

Runs use separate scratch and artifact directories. The runner registers the exact scratch root before launch, forwards interruption to the owned process, and removes registered scratch on success, failure, or cancellation. Captures and the manifest remain in the artifact directory. Use `cleanup --run <run-id>` only to recover a registered leftover; it rejects unknown IDs, active recorded processes, and paths outside the registered scratch root. Do not remove paths by pattern or clean another run's artifacts.

## Scenarios and captures

A custom scenario is an absolute-path ESM module with a default async function. Its content hash is recorded. Keep it self-contained (Node built-ins are available); only the scenario file is staged, not sibling imports. The context provides the Electron app and Playwright page, `step`, `capture`, `dispatch`, `readSnapshot`, `reload`, `completeOnboarding`, `setTheme`, `setWindowSize`, a disposable workspace, and non-secret logging. Fixture dispatch is limited to `project.create`, `thread.create`, and `thread.metadata.update`. On protocol V2, projects use the authenticated project mutation HTTP endpoint, thread commands use ticket-authenticated WebSocket RPC, and shell readback sends the required protocol header. Tokens and tickets stay inside the private renderer and are not returned as evidence. Create fixture state through the app dispatch or UI; do not write the database directly or send a provider turn.

`setTheme` controls a recorded native-theme and CSS-media fixture and reads the rendered appearance back. It does not prove user theme-preference persistence. Use `ctx.reload()` to restore the recorded navigator-online fixture after navigation.

The built-in `sidebar-rename` scenario completes the normal provider-free first-run UI, then creates a synthetic project and metadata-only thread, captures `before`, renames through the Sidebar, checks the rendered title and backend snapshot, captures `after`, reloads and verifies persistence, then switches to dark theme with readback and captures `after-reload` and `after-reload-dark`.

**Those four capture names describe interaction states in one candidate build. They are not source-code before/after evidence.** The `--comparison OID` option records a declared comparison commit; it does not build or run that revision. To compare a code change, run the same scenario against separately built, exact base and candidate revisions, and retain both run manifests and capture sets.

## Reading the manifest

Each run writes `manifest.json` using schema `jones-code-ui-evidence/v1`. It records the run, harness and source identities, build digests, runtime and window readbacks, fixture and scenario steps, assertions, captures, isolation probe and topology, resource measurements, coverage claims, comparison declaration, status, cleanup outcome, and any error. Capture entries include the PNG hash and actual dimensions, scale, rendered theme, window bounds, and viewport.

Source provenance includes the source HEAD, dirty status, a digest of the tracked diff and untracked-file content, and the declared comparison OID. A dirty-tree capture must be described as that base plus its recorded patch identity, not as a commit-only result. A source OID alone does not prove which code produced existing build outputs: `build.sourceCorrespondence` is `receipt-matched` only when the exact source and build-output receipt match; otherwise describe it as unproved. The unpackaged runtime identity may be `null`.

Use the manifest to state what the run proves and what it leaves open. The built-in scenario covers project/thread metadata and a Sidebar rename with backend readback and reload persistence. It does not cover message content, provider turns, or external services. The Linux harness does not establish behavior for macOS, packaged/fused builds, signing, native dialogs, provider/network flows, or GitHub attachment upload.

## Build receipts

Build the candidate through the repository's desktop build command (`vp run build:desktop`) with development and packaged-executable overrides unset. Keep the source revision and dirty identity unchanged from immediately before the build until the final output digest collection. The harness does not attest an arbitrary receipt's honesty: the operator owns that build sequence.

For automated build orchestration, import `collectProvenance` from `apps/desktop/scripts/ui-evidence/provenance.mjs`. Supply `source`, the pinned `harness`, a new artifact directory, and `buildPaths` with `desktop: <source>/apps/desktop/dist-electron`, `server: <source>/apps/server/dist`, and **`web: <source>/apps/server/dist/client`** (the served client). Collect before and after the build into separate directories, and reject any changed source identity. Write this receipt from the post-build result:

```js
{
schema: "jones-code-ui-evidence-build/v1",
source: {
head: result.source.head,
statusSha256: result.source.statusSha256,
patchSha256: result.source.patch.sha256,
untrackedSha256: result.source.untracked.sha256
},
build: Object.fromEntries(
["desktop", "server", "web", "boot"].map(name => [name, result.build[name].sha256])
),
invocation: "vp run build:desktop (development/packaged overrides unset)",
startedAt: /* ISO timestamp immediately before build */,
finishedAt: /* ISO timestamp after successful build */
}
```

Pass its absolute path through `--build-receipt`. The command rejects changed source or outputs. Without a receipt, exploratory runs remain usable, but source-to-build correspondence is explicitly unproved. An installed command's pin identifies the harness, not an independently supplied candidate source; use `--source` for a different candidate.

## Sharing evidence

The default `.t3/ui-evidence/runs/` output is task-local and untracked. Attach screenshots or video through the repository's authorized GitHub workflow and verify that each attachment can be retrieved and renders in the PR. The runner does not upload evidence or commit generated captures. Keep the manifest with the PR evidence record so reviewers can identify the exact build, fixture, route, and coverage.

For every evidence set, identify the comparison and candidate revisions, build identity when available, route and platform, actual viewport, scale and theme, fixture and scenario, observable action/readback, and coverage limits. A combined preview must name its integration revision and constituent heads. Refresh any exact-current-head claim after the PR head changes.
15 changes: 12 additions & 3 deletions .agents/skills/test-t3-app/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,10 @@ description: Test T3 Code's web and desktop UI through its built-in Browser pane
Use T3's built-in Browser panel for verification. If its tools are absent or
the panel reports unavailable, explain the blocker and stop verification.
Do not install or switch to another automation system. For native mobile
testing, use [test-t3-mobile](../test-t3-mobile/SKILL.md).
testing, use [test-t3-mobile](../test-t3-mobile/SKILL.md). For separately
qualified Linux Electron app-window evidence, see
[capture-ui-evidence](../capture-ui-evidence/SKILL.md); it does not replace the
Browser-panel route for shared-renderer work.

## Start the app

Expand All @@ -31,8 +34,14 @@ for a fresh one. Keep using the same tab.

## Verify and retain

Exercise the affected flow and capture the state that proves it works. Keep
the server, state, and panel available while the user inspects or iterates.
Follow [capture-ui-evidence](../capture-ui-evidence/SKILL.md) for the common
evidence format. Exercise the changed flow with snapshots and focused locators,
assert its visible result, and check backend readback or reload persistence when
relevant. Record the revision, served build/source correspondence, actual
viewport/scale/theme, fixture, and coverage limits. Distinguish interaction
states within one build from two separately built source revisions.

Keep the server, state, and panel available while the user inspects or iterates.
An assistant turn ending is not teardown. Stop only processes you started,
using retained terminal sessions or captured PIDs.

Expand Down
7 changes: 6 additions & 1 deletion .agents/skills/test-t3-mobile/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,11 @@ through AgentDevice. For a backend on the device host, use
on Android. For a remote backend, use its reachable origin.

Confirm the intended projects appear, exercise the affected flow, and capture
evidence. Retain the app and environment while iterating. At teardown, remove
evidence using the common format in
[capture-ui-evidence](../capture-ui-evidence/SKILL.md). Record the candidate
revision, native-client and served-backend build correspondence, device and
platform, actual viewport/scale/theme, fixture, action/readback, and coverage
limits. Distinguish interaction states in one build from two source builds.
Retain the app and environment while iterating. At teardown, remove
the disposable connection, close the AgentDevice session, call `device_close`,
and stop only your backend and Metro processes.
8 changes: 7 additions & 1 deletion .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,12 @@ is insufficient; broad repo-wide checks are not required.

For UI changes, include clear before/after screenshots. Include a short recording
when motion, timing, transitions, or interaction details are needed to demonstrate
the change. Upload evidence to GitHub and embed or link it here. Never commit PR-only assets. -->
the change. Upload evidence to GitHub and embed or link it here. Never commit PR-only assets.
For each evidence set, identify the full comparison and candidate revisions (plus
base and patch identity for dirty source), build correspondence, route/platform,
actual viewport/scale/theme, fixture/scenario, observable action/readback, and
coverage limits. State whether before/after describes one interaction in one
build or two exact source builds. Name the integration revision and constituent
heads for combined previews; refresh exact-head claims after each PR-head change. -->

<!-- If you used an agent, end with the model and harness that did the work. -->
Loading
Loading