Skip to content

ci(e2e): pull the current backend images on every run - #549

Merged
MrYuion merged 11 commits into
developfrom
e2e/pull-images-each-run
Oct 8, 2026
Merged

MrYuion merged 11 commits into
developfrom
e2e/pull-images-each-run

Conversation

@camreeves

Copy link
Copy Markdown
Contributor

What was wrong

The e2e stack tracks ${PLACEOS_TAG:-latest}, but docker compose up never re-pulls a tag that is already present. On the persistent self-hosted runner that meant latest was whatever had been pulled the day the machine was set up: the nightly was testing the July release while 2.2608 and 2.2609 shipped, and the "Backend inputs" summary was the only sign.

Change

  • e2e/stack/up.sh --pull pulls the current images before starting the stack, prunes dangling layers after a successful pull, and on a failed pull warns and runs with the images already present (the recorded backend inputs show which ones ran).
  • The advisory workflow calls up.sh --pull.
  • A line in the runner doc, since the Docker Hub hosts now have to stay reachable rather than work once.

Verified

Locally, up.sh --pull pulled and started the stack. On the runner, the push of this branch ran the advisory job (https://github.com/PlaceOS/user-interfaces/actions/runs/37578866715): the pull step fetched new layers for every PlaceOS image, which is the release channel's current placeos-2.2609.6, bring-up took about nine minutes longer for that one run, and the suite passed (14 of 14 on this develop-based branch).

The stack tracks `${PLACEOS_TAG:-latest}`, but on the persistent runner
`latest` was whatever had been pulled when the machine was set up, so
the nightly tested a backend from July while releases moved on.
`up.sh --pull` fetches the current images before starting; a failed
pull warns and runs with what is present, and the recorded backend
inputs show which images ran.
@vercel

vercel Bot commented Oct 7, 2026

Copy link
Copy Markdown

Deployment failed for project frontend-templates with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/placeos?upgradeToPro=build-rate-limit

@MrYuion

MrYuion commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

The change does what it says for the normal case. I found two medium issues and four low issues in the failure paths and the docs.

1. One failed image cancels all other pulls (medium)

e2e/stack/up.sh:93, dc pull --quiet

Without --ignore-pull-failures, docker compose pull fails fast. When one image fails, compose cancels the errgroup context, so the pulls that are in progress stop and the queued pulls do not start (see runServicePull in compose pkg/compose/pull.go: "fail fast if image can't be pulled nor built"). One timeout or one Docker Hub rate-limit response can therefore leave most of the stack on the old images. The stack can also start on a mix of old and new images. The warning says "running with the images already present", which hides how much of the stack did not update.

This also occurs when E2E_AUTH_IMAGE points to a local-only build. Its pull always fails, so --pull updates almost nothing.

Fix: add --ignore-pull-failures so that each image pulls independently. With that flag, compose exits 0 and prints each failure in the step log. Thus, the else branch does not run. Change the warning text, or remove the branch.

dc pull --quiet --ignore-pull-failures

2. Secrets are made with the old placeos/init image (medium)

e2e/stack/up.sh:50-60 runs before the pull block at :89-98.

actions/checkout runs git clean -ffdx, and e2e/.gitignore ignores .secrets/. Thus CI makes new secrets on every run. It uses docker run placeos/init:latest generate-secrets, which runs before the pull, so it uses the image from the previous run. Later, dc run --rm init start uses the new image. If a release changes the secret format or adds a key, the new services get secrets from the old generator.

Fix: move the pull block to after the touch at line 52 and before generate-secrets. The touch is necessary first, because compose does not load the project when an env_file is missing. (I did a check: docker compose config --images fails without .secrets/.)

3. docker image prune -f affects the full host (low)

e2e/stack/up.sh:94

The prune removes all dangling images on the Docker host, not only images from the placeos-e2e project. The script header (line 6) and e2e/README.md:44 say that the stack "cannot disturb" the developer's own stack. On a developer machine, up.sh --pull now removes dangling images from other projects.

Fix: move the prune into the workflow (for example, into the "Reclaim the machine" step or the teardown step), where the runner is dedicated. This keeps up.sh limited to its project.

4. --pull also moves postgres:18-alpine (low)

e2e/stack/docker-compose.yml:68

dc pull pulls every service. Before this change, postgres:18-alpine was frozen on the runner. Now it moves to each new 18.x patch and each new Alpine base. The keydb comment at line 161 gives the rule for infrastructure images: "a moving tag here can only ever add noise". Elastic (exact tag) and keydb (digest) obey that rule. Postgres does not.

Fix: pin postgres to an exact version (for example postgres:18.x-alpine or a digest). Alternatively, pull only the PlaceOS services.

5. A slow pull has no limit and shows no output (low)

e2e/stack/up.sh:93 and .github/workflows/e2e-advisory.yml:165

The fallback works only when the pull fails. When the pull stalls on a slow link, the job continues until the 45-minute job timeout (line 96), and the fallback does not run. Also, --quiet gives no output for the full pull, so the log does not show which image is slow. The first run after a release already added about 9 minutes. The script bounds the other waits ("Bounded so a stuck container fails with a clear message"), but it does not bound the pull.

Fix: in CI, run the pull as a separate step with timeout-minutes and continue-on-error: true, and then run up.sh without --pull. Alternatively, remove --quiet so that the log shows progress.

6. Docs do not show --pull (low)

  • e2e/README.md:47-48 lists the up.sh options without --pull.
  • e2e/README.md:240, in the CI table, says that bring-up needs GitHub "once". It does not say that every run now pulls from Docker Hub. SELF_HOSTED_RUNNER.md has this information, but the README table does not.

Fix: add e2e/stack/up.sh --pull to the option list, and add a Docker Hub row to the CI table.

Checks done

  • bash -n e2e/stack/up.sh: OK.
  • The workflow YAML parses.
  • docker compose config --services gives the 10 expected services. init has no profile, so dc pull includes it.
  • I did not run the stack, because no Docker daemon was available.

Review by Claude Opus 5.5 in Claude Code (T3 Code).

… farm

The e2e stack had no core, so nothing in it could run a driver and the
specs that bind module state (the home availability panel, room check-in)
could only be guarded. core is now part of the stack, started after init
has created the tables it subscribes to, with its binaries coming from the
PlaceOS build farm the way every deployment's do.

The seed creates a repository row and driver rows for Place::Bookings and
the demo calendar (PlaceOS/drivers#639), pinned to a commit so the farm
builds each CPU architecture once, and waits until core holds the
binaries before bring-up returns. Helpers for logic modules and module
state go with it, for the specs that attach the drivers to rooms.

The runner needs build.placeos.run and the drivers S3 bucket reachable.
camreeves and others added 2 commits October 8, 2026 10:50
…65 sandbox tenant

The tenant row carried placeholder credentials, so every calendar-backed
route (the concierge day view, the staff directory, the rooms and
attendance reports) died at Microsoft and their specs could only be
guarded.

With E2E_O365_TENANT, E2E_O365_CLIENT_ID and E2E_O365_CLIENT_SECRET set,
the seed writes app-only credentials to the row, updating one left over
from a placeholder run, and creates a local admin whose address is a
mailbox in the tenant so staff-api can create events on its behalf.
Without them nothing changes and the specs that need the tenant skip.

CI takes the values from the repository variables and secret of the same
names; they belong to the sandbox app "PlaceOS Bookings Visualiser".
- Move the drivers repository row to the configured uri and branch.
- Replace a room's module from an earlier driver instead of adding a second one.
- Retry /compiled on network errors, and ignore build output left by an earlier attempt.
- Re-check the search-backed driver listing before creating a row.
- Log a warning when drivers do not load, so a build farm outage does not fail every spec.
- Share the listing helpers through e2e/support/api.ts and type the rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
frontend-templates Ignored Ignored Preview Oct 8, 2026 2:41am UTC

… step

Review round on #549.

- `docker compose pull --ignore-pull-failures`: one image that cannot be
  pulled no longer cancels the others, so the stack cannot end up half
  updated because of one timeout or one rate-limit response. The warning
  branch goes with it; compose prints each failure itself.
- The pull now runs before `generate-secrets`, so the secrets come from the
  init image the run is about to use. The env files are touched first
  because compose will not load the project without them.
- `--pull-only` pulls and stops. CI runs it as its own step with a 20 minute
  limit and `continue-on-error`, then brings the stack up without a pull, so
  a stalled registry costs at most that step. The pull is no longer quiet,
  so the log shows which image is slow.
- `docker image prune -f` moves out of `up.sh` into the workflow's reclaim
  step; on a developer machine the script now touches only its own project.
- postgres pinned to `18.6-alpine`, the version the runner already has, so
  the pull cannot move an infrastructure image.
- README: the two new options and a Docker Hub row in the CI table; the
  runner doc names the step.
@camreeves

Copy link
Copy Markdown
Contributor Author

Thanks, all six are in as 61b960d.

  1. dc pull --ignore-pull-failures, and the warning branch is gone with it: compose reports each failure itself and exits 0, so the else could never run.
  2. The pull now sits straight after the touch of the two env files and before generate-secrets, so the secrets are made by the image the run is about to use.
  3. The prune is out of up.sh and in the workflow's reclaim step, where the host is ours.
  4. postgres is 18.6-alpine, which is the version the runner already resolved 18-alpine to, so nothing moves on the next run.
  5. Done as a separate step: up.sh --pull-only pulls and stops, the workflow runs it with timeout-minutes: 20 and continue-on-error: true, then runs up.sh without a pull. --quiet is gone too, so the log shows which image is slow.
  6. README has both options and a Docker Hub row in the CI table; SELF_HOSTED_RUNNER.md names the step.

Checked locally: up.sh --pull-only takes 9 s with the images present and prints one Pulling and one Pulled line per service; with E2E_AUTH_IMAGE pointed at an image that does not exist, compose prints auth Error pull access denied ..., pulls the rest and exits 0; a full up.sh on the restructured script reaches stack ready.

camreeves and others added 5 commits October 8, 2026 13:42
# Conflicts:
#	.github/workflows/e2e-advisory.yml
#	e2e/README.md
#	e2e/stack/SELF_HOSTED_RUNNER.md
# Conflicts:
#	e2e/stack/SELF_HOSTED_RUNNER.md
- The seed names the tenant row after what it holds, and specs read it with
  calendarBacked(), so the test step needs no O365 secret.
- The O365 vars move from the job env to the bring-up step only.
- ensureTenant always writes the row back, so a reused stack goes back to
  placeholders.
- ensureUser upserts users for staff and calendar roles and updates the
  password and sys_admin of an existing user.
- calendar.env.ts loads e2e/.env itself and is the one source for the
  calendar user.
- Fix the directory route in the docs and list the new vars in .env.example.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ci(e2e): run core in the stack and load pinned drivers from the build farm
ci(e2e): back the stack's calendar and directory with the Microsoft 365 sandbox tenant
@MrYuion
MrYuion merged commit 73af7f6 into develop Oct 8, 2026
1 of 2 checks passed
@MrYuion
MrYuion deleted the e2e/pull-images-each-run branch October 8, 2026 02:57
camreeves added a commit that referenced this pull request Oct 8, 2026
#549 moved the Microsoft 365 variables to the bring-up step and added
`calendarBacked(api)`, which reads the tenant row the seed wrote. The test
step no longer sees the variables, so a skip keyed on them would skip the
calendar rows on a stack that is backed.

The concierge fixtures gain a worker fixture, `conciergeCalendarBacked`,
answered once per worker with the admin token, and the calendar identity's
state and token fixtures key off it. The three tenant-only groups skip from
a `beforeEach` on the same fixture.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants