A headful, collaborative browser MCP server for the VEGA agent harness. VEGA drives a single visible Chromium window that the user can watch and take over at any time — co-browsing, not headless scraping.
The same server runs headless under Hermes, where the user is reached through chat rather than at a window — see docs/HERMES_INTEGRATION.md.
Renamed from vega-browser. The old
VEGA_BROWSER_*environment variables and an existing~/.vega-browserdata directory are still honoured; theLYRA_BROWSER_*names win when both are set.
Built on Playwright (headful, persistent profile) and exposed over FastMCP, so VEGA picks it up through its existing MCP integration with near-zero glue.
VEGA is local-first and model-agnostic. This plugin gives it a browser the user shares: the agent reads the page, clicks, and types, while the user can step in for the things an agent must not do alone (passwords, CAPTCHAs, 2FA, payments) and grab full control whenever they want. Every action is written to an append-only audit trail.
VEGA agent loop ──(MCP: stdio/http)──> lyra-browser server ──> Playwright ──> visible Chromium window
│ ▲
├─ audit.jsonl (every action) │ user watches / takes over
└─ approval + takeover gating ──────────┘
- Standalone MCP server — registered in VEGA's
mcp.json(no VEGA code change). - Python + Playwright — same runtime as VEGA; pins track the harness
(
fastmcp>=3.2,playwright>=1.59— one minor above the 1.58.0 VEGA bundles, because element refs need it; see Addressing an element). - Dedicated headful window —
launch_persistent_context(headless=False)with a persistent profile, so logins survive across sessions. One server at a time holds that profile (anflockon<data_dir>/profile.lock, taken when the browser opens, not when the server starts). A second server on the same data dir opens a private, empty profile at<data_dir>-instances/<pid>/profileinstead of failing, andopen_browsersays so (profile: "instance", "no saved logins in this instance"). - No setup tax — reuses the user's installed Chrome/Edge by default
(
channel="chrome"→"msedge"→ bundled Chromium), so an end user with only VEGA.app needs no pip and noplaywright install. If no browser is found, tools return abrowser_unavailableenvelope for VEGA's UI to prompt an install. - Full preset — human-in-the-loop (highlight, ask-the-user, approval gate), takeover/handoff, audit trail, CI + lint + pre-commit.
| Group | Tools |
|---|---|
| Navigation | open_browser, navigate, go_back, reload_page, close_browser |
| Tabs | tabs |
| Waiting | wait_for |
| Interaction | click, type_text, press_key, hover, scroll, handle_dialog |
| Downloads | list_downloads |
| Forms | read_form, select_option, set_editor, upload_file, read_draft, save_draft, publish |
| Reading | get_url, read_page, screenshot, read_image |
| Collaboration | highlight_element, ask_user_to_do, request_takeover, resume_after_takeover |
click and type_text cover what every page has, but a form on a real site needs
more. Each of these was measured on KVR's developer forms before it was
written:
read_form— a form'sname,type,labeland<select>options.read_pagereturnsinnerText, so a 188-field product form arrives as a wall of labels with nothing to address. Hidden inputs are counted, not dumped. Every field carriesusable— false means it is in the DOM but hidden or disabled right now, which is how a form keeps a section it has not revealed (measured on KVR:event_countrysits in the Deal/Offer block). Checkusableinstead of finding out through a 30s locator timeout.select_option— a<select>is not a click target. Chooses byvalue,labelorindex, and fireschangethe way a page listens for it. Passsubmits=truewhen the choice posts. Answershiddenordisabledby name when the control cannot be used as it stands.set_editor— CKEditor/TinyMCE keep the editable body in aniframeand leave the<textarea>the form posts hidden and empty, sotype_texton that textarea does nothing.html=truewrites markup and syncs the editor's own data. Leaveselectorempty to use the first editor on the page.upload_file— attaches local files to aninput[type=file]. A styled picker usually keeps it hidden, which is fine. Missing paths are reported before anything is attached.save_draft— sends the form while keeping the item private. It names the control that sends (submit, fromread_form'ssubmitText), first confirms the publish control is still on its private setting and refuses otherwise, and asks forSUBMITbut neverPUBLISH— so it cannot release an item, and a page left armed by anything else will not slip out through it.read_draft— everything a human needs to approve an item, in one answer: the filled fields, each editor's real content (a separate documentread_pagenever shows), the attachment slots, the submit control, and whether the item is a draft or already public. A read, so it needs no approval.publish— the one action that sends an item to an audience. Kept apart from every other tool so that writing a draft can never release it: it asks forPUBLISH, which nothing implies, and it is the only tool that does. A site splits the act across two controls and means neither alone — measured on KVR, the publish radio only arms the form and the item travels on the form's own Submit — so it takes bothselector(the live control) andsubmit(the one that sends). Giving only the first is refused rather than left half-done: arming a form with no way to send it would leave the page primed to publish on someone else's next click.
Writing an item and releasing it are different acts, and the tool surface keeps
them apart. Every writing tool can be driven to completion with the item still
private; publish is the only way out, and it is PUBLISH-gated and single-use.
The intended flow is: fill the form → read_draft → show the human → only then,
on an explicit yes, publish.
On KVR this matches the page: a new item opens on Draft and stays there until
the publish radio is switched, so a form that is merely filled is not visible to
anyone. A form sent without the publish control set saves as a draft — which also
means every writing tool must stay away from Submit, since KVR has no separate
"save draft" button. That is why publish sends the form itself rather than
leaving the send to click.
UPLOAD and PUBLISH are single-use: a file handed to a page, and content shown
to an audience, cannot be recalled.
read_page has two modes, both read-only (no approval, works during a takeover):
-
mode="text"(default) — the rendered text (innerText), not HTML.links=trueadds the visible links as{text, href}(absolute, de-duplicated, at most 100, withlinks_truncated). -
mode="tree"— one line per thing that can be acted on (links, buttons, inputs, selects, checkboxes, tabs, menu items), each with a ref:aria-ref=e12 link "Pricing" -> /pricing aria-ref=e15 button "Save" aria-ref=e16 checkbox "Remember me" [checked]Input values are never shown. The reply carries
elementsandrefs, plusin_viewport(on-screen elements come first) on Playwright 1.60+.
selector limits either mode to one region (first match; not_found at once
when nothing matches). offset and max_chars page through long output:
total_chars is the full length, truncated says more follows and next_offset
is where to continue. Tree pages break on whole lines, so a ref is never cut in
half.
read_page(mode="tree") → pass the token aria-ref=eN, exactly as written, as
the selector of click or type_text. A ref reaches what a CSS or text
selector cannot name: icon-only links, several links with the same text, elements
inside open shadow roots and iframes. A ref is a handle into the latest tree
read, not a query over the live page: a newer read (a region read included), a
navigation or a removed frame ends it, and so does the element being removed. A
stale ref answers not_found (or hidden for an element that is gone but still
counted) with a hint to read the tree again; refs get no waiting window, because
they cannot appear later.
Requires playwright>=1.59. Refs come from aria_snapshot(mode="ai"), which
1.59 introduced; pyproject.toml still floors at 1.58 because that is what a
VEGA runtime ships. On 1.58 the tree still works but answers refs: false and its
lines carry no ref — address elements with role=link[name="Pricing"] or text=
selectors instead.
Both fail fast and explain, instead of holding the call for Playwright's 30s:
status |
Meaning |
|---|---|
not_found |
Nothing matches the selector, even after up to ~3s for a page that builds its controls late |
hidden / disabled |
The match is in the page but stayed invisible / disabled for that window (present-but-not-ready controls are re-checked until ready, so a button enabled 0.8s after load still works) |
timeout |
The action outran timeout_ms. The click may have landed (a navigation it started can still be loading) — look at get_url/read_page before repeating it |
element_not_actionable |
Covered by another element (the hint names it), detached, read-only, not a text field, unstable or outside the viewport |
page_closed |
The tab died under the call (a popup that closed itself, a closed window) |
Each carries selector, url and a hint with the next move. timeout_ms
(default 10000, at most 30000; 0 is raised to 1ms, never "no limit") bounds the
action itself. The usability check runs before permission is asked, so an
unusable control neither prompts nor leaves a single-use grant behind; the same
driver failures during the action are converted into these envelopes (and
audited) rather than raised. A successful click also reports matches (how
many elements the selector hit — the first is clicked) and clicked
(tag, role, name of what was hit), with a hint when matches > 1.
click, type_text, press_key and hover answer when the driver reports the action
done, and two things can still be true of that moment:
- The guard refused a navigation the action caused (a link to a site no approval
covers, a form sent without
submits=true, a page that redirected itself). The 204 leaves the tab where it was, so the reply used to readokwith the sameurland the agent had no idea why nothing happened — the audit trail was the only account. It is nowblocked_by_policywithurl(where the tab stands), ahintand, for a refused redirect,redirected_to; the fields the reply already had (matches,clicked,dialogs) are kept, and the audit row of the action carries the same status. The hint is the way out:navigateto the destination (that asks the user), or repeat the action withsubmits=true(submit=trueontype_text) when it was meant to send a form. A file the call saved staysokwith itsdownload, and a download the browser started and the call cancelled keepsdownload_blocked. Inobservemode nothing is refused, so the reply staysokand the audit sayswould_deny. - The action opened a tab (
target=_blank,window.open). The session follows onto it whileurlis still the page the action was made on, so the reply carriesnew_tab: trueandtab_count, and a hint to usetabs.get_urlandopen_browsercarrytab_counttoo. A popup whose first navigation the guard refused opens no tab at all, and the reply isblocked_by_policyinstead.
Neither costs a click a wait of its own. The driver holds a click until the navigation it
started has been judged, so that verdict is in before the call returns (0.5 to 48 ms
before, measured over playwright and patchright, both guard backends, headless and
headful). What it does not hold for lands 3 to 22 ms after — a popup's tab, a form posted
from a frame, a page's own setTimeout(0) — inside the download_settle_s window every
click and key press already listens for. hover and type_text have no such window, so
they linger up to 40 ms after the driver returns, and only while nothing has shown (their
navigation reaches the guard 1 to 8 ms after they return); a page timer that fires after
that is only on the audit trail. Numbers in scripts/verify_click_refusal_e2e.py
(INFO lines).
wait_for holds until the page is ready instead of a sleep — or a click used to
pass time. Give exactly one condition: text (in the visible text read_page
shows; state="hidden" waits for it to vanish), selector (reaches state:
visible, hidden, attached, detached), url (a glob over the whole URL,
e.g. **/checkout**) or load_state (load, domcontentloaded,
networkidle). It waits up to timeout_ms (default 10000, capped at 30000) and
returns ok with waited_ms. Running out of time is an answer, not an error:
{"status": "timeout", "waited_ms", "last_seen"} with a short page excerpt (the
URL for a url wait). Bad arguments return error. It only observes: no
approval, and it keeps working while the user holds a takeover.
The session follows onto any tab a page opens (target=_blank, window.open),
so the agent is never left reading the page that opened it. A click or
press_key that opened one says new_tab: true and tab_count, and get_url and
open_browser carry tab_count. tabs covers the rest:
action="list"(default) —tabs: [{index, url, title, active}]plustab_countandactive_index. A read; a tab stuck in a script loop is listed without a title rather than hanging the list.action="switch",index=N— make that tab the one to read and act on, and bring it to the front so the user sees what the agent sees.action="close",index=N— close it. Closing the active tab returns to the tab that opened it, else the newest one; closing the last tab leaves a blank one (closing headful Chrome's last window would quit the browser).
Indexes shift whenever a tab opens or closes, so list again before using one; a
bad index answers not_found with tab_count. Switching and closing are
mutations: they stand down for a takeover, are audited, and ask for INTERACT
on the tab's own site, not the active page's — being approved for one site is
no licence to read or close another's. A tab that closes itself (a sign-in popup)
sends the session back to its opener; if the active page is ever closed, the next
tool moves onto the opener, else the newest tab, else a fresh blank one, instead
of failing on a dead page.
ok means the navigation happened, not that the page is good. navigate,
go_back and reload_page report http_status (404, 500 …) and content_type
(media type only, lower-cased, e.g. application/pdf) of the document that
answered; both are null when no HTTP response was involved (about:blank,
data:, a page the browser restored from history). wait_until chooses how much
of the load to wait for: domcontentloaded (default), load, commit or
networkidle (which never settles on a page that keeps polling); anything else is
an error before any prompt. A page that draws itself after loading needs
wait_for. A URL that turns out to be a file download does not move the tab: see
Downloads for download=true and download_blocked. Only when the
driver announced a download the browser never reported does it answer
{"status": "download_started", "url", "requested_url", "hint"}; this server then wrote
no file to the download dir.
hover— move the pointer over an element, for menus and tooltips that open only while it is on their trigger:hover, thenclickthe item that appears. Nothing is pressed, so no form is sent. Fails fast likeclick(not_found,hidden,disabled,timeout,element_not_actionable,page_closed,timeout_ms) and reportsmatchesandhovered(tag,role,name). NeedsINTERACT.scroll— give exactly one ofto(top/bottom),by_y(pixels, negative up, at most 20000 per call, done with the mouse wheel) orselector(bring it into view). The reply isscroll_y,scroll_heightandat_bottomonce the position has settled. It moves the window only: a feed or panel that scrolls inside its own container needsscroll(selector=...)on an element inside it, and a scroll that moved nothing says so. Lazy lists grow after a scroll, so repeatscroll(to='bottom')untilat_bottomis true.
A file a page hands to the browser is kept only when the call declared it. Chromium accepts a download the moment a page offers one, and the navigation guard judges requests, not responses, so this is a separate gate on the file as it arrives:
download=trueonclick,press_keyornavigatedeclares it and buys a single-useDOWNLOADgrant (asked like any other, never implied by a trusted origin) inside that call. The grant pays for exactly one file and is retired when the call stops waiting. The reply carriesdownload:{filename, path, bytes, url_origin}.- An undeclared download is cancelled and deleted before anything reaches the
download dir, and the call answers
download_blocked(blocked_download:filename,url_origin) — repeat it withdownload=true. A declared call that produced no file answersdownload_not_started; one whose file was not kept answersdownload_failedwith areason(too_large,timeout,error).go_backandreload_pagecannot declare one; their hint points atnavigate. Chrome reports a reload or history step that turns into a download asnet::ERR_ABORTED(not "Download is starting"); an abort that a download follows within 2 s is answered as that download, and any other abort is still an error. - With
LYRA_BROWSER_ENFORCEMENT=observe("record, do not block") nothing is cancelled either: an undeclared download is saved like a declared one — same name rules, caps and ledger — and answered asdownloadplusobserved: true, audited aswould_block. Observe never spends a grant (like the navigation guard), so it under-reports a second file in one call that enforce would have cancelled. A download aDOWNLOADgrant covered is auditedsavedand is not flagged. - The site picks the file name, so it is treated as hostile: only the last path
component survives, control and right-to-left-override characters, characters
Windows forbids, leading dots and reserved device names are removed, long names
are cut keeping their extension, and a taken name becomes
name (1).ext. Nothing is ever written outsideLYRA_BROWSER_DOWNLOAD_DIR. - Size and time are capped (
LYRA_BROWSER_DOWNLOAD_MAX_BYTES,LYRA_BROWSER_DOWNLOAD_TIMEOUT), and the dir is pruned oldest-first toLYRA_BROWSER_DOWNLOAD_KEEPusing a ledger, so files that are not this server's are never deleted. - While the user holds a takeover the gate stands aside and their own
downloads are saved, recorded as
user_driven, as with navigation. list_downloads— the files this session saved (filename,path,bytes,url_origin), oldest first, plusdownload_dir. A read: no approval, works during a takeover, does not open the browser. Blocked and failed downloads are not listed.
The listener is per tab (popups included), not BrowserContext's download
event: that only exists from Playwright 1.60 while this project supports 1.58,
where a context listener would never fire and every download would go unjudged.
alert, confirm, prompt and beforeunload stop a page until answered, and a
listener that leaves one open freezes the tab. So one context-level listener
always answers the moment a dialog appears. The default is what a browser with
no listener does: alerts and beforeunload are accepted, confirm is dismissed
(false) and prompt is dismissed (null).
- What was raised is reported as
dialogs(type,message, how it was answered) in the reply of theclick,type_textorpress_keythat raised it — only present when there were any. The message is the page's text, not the user's. The form tools (save_draft,publish) use the same handler and fold dialog text into theirerrors. handle_dialog(accept, text)chooses the answer for the dialog the next browser action raises: call it immediately before that action.accept=falseanswers no;textis what apromptreceives (empty accepts the default). Like a one-shot grant it belongs to that one call —click,type_text,press_key,hover,scroll,navigate,go_back,reload_page,select_option,set_editor,upload_file,save_draft,publishor atabsswitch/close — and ends with it, used or not, even when the call was refused (arm again before retrying). A dialog raised after it, by a page's own timer or on a page reached later, gets the default, so a yes armed for one click cannot answer a laterDelete this?. Reading (read_page,screenshot,get_url…) does not use it up, and about 60s is the most it ever waits. It asks nobody (the click that raises the dialog was already asked for), stands down for a takeover, is audited, and the prompt text is never written to the audit log.- Nothing here judges the answer: accepting a
confirmthat goes on to post a form is still stopped where the request leaves unless it was declared (submits=true).
screenshot (the viewport or the full page) and read_image (one element —
an img, a canvas, an svg, a chart) write a PNG to disk and return its
path:
{"status": "ok", "image_path": "…/uploads/browser/page-20260917T123557Z-fd1a0721.png",
"mime_type": "image/png", "width": 1280, "height": 800, "bytes": 115889}The file is the handoff. Measured on VEGA: a tool result is shown to the model
as its text blocks, so a base64 PNG inside the JSON arrives as tens of
thousands of tokens of prose and no picture. A path costs nothing, and the
client attaches the file to the model's next turn as an image. Because VEGA
attaches only files under its own uploads root, captures default to
$VEGA_DATA_DIR/uploads/browser whenever VEGA_DATA_DIR is set — no
configuration for the model to be able to look. inline=true adds the base64
for a client with no access to this filesystem.
Captures are written atomically and pruned oldest-first
(LYRA_BROWSER_CAPTURE_KEEP, default 200), so an agent that looks every turn
does not fill the disk. The audit records the path and size, never the pixels.
read_image fetches nothing — it captures what the element renders, so no
request leaves the browser and nothing is gated.
The agent declares what it intends; the browser is judged on what it actually
does. Those are separate layers because a tool cannot know what a page will do
with a click — onclick="form.submit()" submits from a control that looks inert,
an input listener submits while the agent only typed, and an element's type can
change between being checked and being clicked.
Grants are (origin × capability) and short-lived:
| Capability | Meaning |
|---|---|
NAVIGATE |
Load a document from this origin. Implies INTERACT on it |
INTERACT |
Click and type while staying on this origin |
SUBMIT |
Send a non-idempotent request. Single-use |
UPLOAD |
Give a file to a page (a picker, a drop target). Single-use — requested by upload_file |
PUBLISH |
Make content publicly visible — an item an audience can see. Single-use — requested by publish, implied by nothing |
DOWNLOAD |
Write a file to disk. Single-use — requested by click, press_key or navigate with download=true; the grant pays for one file and ends with the call |
Arriving on a site carries permission to use it, so ordinary clicking is not
re-asked. Sending a form and leaving for another site are separate decisions and
are asked — declare them with submits=true (submit=true on type_text).
Enforcement watches outgoing navigations at the request, so it catches a
submission or an escape whatever produced it. A refusal answers with HTTP 204
rather than aborting, which leaves the page standing so the agent can recover.
URLs with no host — file:, data:, javascript: — are never "the same site"
and are asked every time. Where the guard stands — Playwright's route, or a second
DevTools connection that also sees redirect hops — is a choice, see
Guard backends.
Where the line falls, measured rather than asserted:
- A same-origin GET navigation is interaction, not a submission. A search box
puts what you typed in the query string; gating that would gate every search
box and every link with parameters.
SUBMITmeans a body or a mutating method. - A submission is authorised by the origin that sends it, not by wherever the
form points. Cross-origin form actions are ordinary — payment handlers, SSO,
third-party endpoints — and no tool can read a form's
actionbefore the click without the TOCTOU this layer exists to avoid. The destination is recorded in the audit under the labelcross_origin_send— an audit key, not a capability, and never purchased. - A page that moves itself to another host is leaving, and is asked. A site
bouncing to its own subdomain lands here too; the guard cannot tell that from a
hostile page escaping. The refusal is recoverable — the original document is
still there and the audit names the destination — so the agent can read where
it was being sent and ask for it.
navigate,go_backandreload_pageanswer it asblocked_by_policy, with theurlthe tab still stands on, about a second after the refusal — also when the page did it before it finished loading. The 204 commits nothing, so the driver would otherwise wait out its own 30 s timeout for a navigation that is gone (scripts/verify_scriptredirect_e2e.py; neverssl.com does exactly this).click,type_text,press_keyandhoveranswer it too when the navigation was theirs (see What an action set off); they used to sayok(scripts/verify_click_refusal_e2e.py). - A redirect hop is judged like the navigation it leads to. A 301/302/303/307/308
to another origin is classified exactly as a first request there would be: leaving
the site is
navigateand is asked, a 307/308 that repeats a POST body issubmit, and a hop that stays on the origin that sent it (or upgrades it fromhttptohttps) is not judged at all. The audit row carriesredirect_from, and an approved site that answers 302 to an unapproved one is refused like the same destination asked for directly:navigate,go_backandreload_pageanswerblocked_by_policyand the tab stays where it was. The agent asked for one URL and was turned away from another, so that answer also carriesredirected_to, the origin it was sent to (never the path or the query), for the agent to ask for. Playwright never shows a hop to a route handler, so the guard hears of it from therequestevent, after the request was sent, and cancels the load there (scripts/verify_redirect_hop_e2e.py; what that leaves open is listed below).
Verified against the shipped defaults: four undeclared POSTs (space key, JS
form.submit() from a type="button", an oninput listener, and an attribute
flipped after the check) are all stopped, while the declared submit goes through.
What this does not cover, stated rather than discovered later:
- With the default
routebackend, an HTTP redirect hop is judged after it has left, not before. The request to the first hop that no grant covers is sent — with that site's cookies, and the body of a 307/308 POST — and only the rest of the chain is cut. A server that answers before the cancellation lands commits the page first, and loopback or a LAN, where the targets most worth refusing live, always does: the guard then replaces the tab withabout:blankso the agent does not read it, the audit recordsnot_stopped, and the call still answersblocked_by_policy. Seeing a hop before it leaves is not possible fromcontext.route: Playwright continues every hop itself, also one that follows aroute.fulfill(302), and answering each document from Node (route.fetch) to see the redirect first puts a Node TLS and HTTP/1.1 fingerprint on every page while still leaving the later hops unseen (measured).LYRA_BROWSER_GUARD=cdpjudges every hop before it leaves, so a refused hop's destination sees nothing; see Guard backends. - A page can ask the browser to prefetch or prerender another page
(
<script type="speculationrules">). Chrome sends those requests, and serves the navigation that later activates one, without an interception point: neither backend is shown the request, nor the click that follows (measured, Chrome 154: a prefetched cross-site link was followed with the guard seeing nothing, and the cross-site prefetch itself reached its server). It is a GET, so it is the same channel as<img src="https://elsewhere/?data">; what it adds is a navigation nobody judged. - With
route, a navigation that a service worker answers is never shown to the guard either. Playwright'sservice_workers="block"is an init script that a page defeats by calling the prototype'sregister, and a worker already in the profile is untouched by it (measured: an origin's pre-registered worker served that origin's page to a link click with no judgement). Thecdpbackend turns workers off for every page it guards. - A single-page app does its damage over
fetch(), which is not a navigation and is not judged. Deleting mail in a webmail client is an XHR, not a form post. - An approved page choosing where to send its form. Judging a submission on its
sender is what makes payments and SSO work, and it is also an exfiltration
channel: the audit records it (under the key
cross_origin_send, which is only a label on the audit row, not a capability), it does not stop it. - A single-use scope is handed back about two seconds after the action that
bought it, not instantly, because a click's navigation can leave just after
the call returns. A page that submits inside that window rides an approval
meant for the click. Two seconds rather than the ten minutes it used to be
(
LYRA_BROWSER_RELEASE_GRACE). - The approval for a
file:URL names that URL and covers only it, once — everyfile:URL is otherwise the same origin, which would make one yes a yes to the whole disk. - The content of a download. It is gated, not inspected (see Downloads): an approved download is bytes from that site, written as-is under a sanitised name.
Enforcement narrows what a compromised turn can reach; it does not make an approved origin safe.
One browser, one session. Reading checks who is driving; only a mutating
tool becomes the driver, so a passive client cannot take the browser by asking
for a screenshot. Reading still counts as using it, so a holder part-way
through a read-only pass is not idle and does not lose the window. Grants are kept per session, but every session drives the
same window, and the route handler has no way back to a request
context — so whichever session called a tool last would decide how the next
request is judged, and one client's single-use approval would be spent by
another's traffic. Rather than offer an isolation it cannot deliver, the server
serves one session at a time and answers the rest with session_conflict.
Ownership lasts as long as the window: close_browser ends a turn, drops the
grants, and lets the next session claim it. A holder that goes quiet for
LYRA_BROWSER_OWNER_IDLE_TIMEOUT (default 900s) also hands it on — closing is
owner-only, so a client that drops with its window open would otherwise lock the
server for the life of the process. The handoff closes the window: passing on
a running browser would leave the previous holder's document loaded and still
able to start navigations, which the guard would then judge, and charge, against
whoever now holds the session.
Who is answering. auto asks the user through MCP elicitation, where the
model cannot forge the answer. Only a client that cannot be asked at all — no
elicitation handler, no live request — falls back to honouring confirm=true,
and the audit records consent_channel=legacy with the reason. A decline, a
cancel and a timeout are answers: the model asserting confirm=true never
overrides one. VEGA does not pass an elicitation handler yet, so it takes the
fallback today; wiring one is what turns the gate from a convention into a
boundary, and needs no change here. Hermes does pass one: its approval prompt
answers, and its value-less accept is read as the yes it is.
Trusted origins. Every new site costs one prompt, and a prompt nobody
answers holds its tool call for LYRA_BROWSER_CONSENT_TIMEOUT (default 300s)
before it becomes a denial. For sites the operator always uses, that answer can
be given in advance: LYRA_BROWSER_TRUSTED_ORIGINS is a comma- or
whitespace-separated list of sites that are pre-approved for NAVIGATE — and the
INTERACT it implies — so they are not asked about again. The operator writes it;
nothing the model or a page says can add to it.
| Entry | Trusts |
|---|---|
https://www.kvraudio.com, http://localhost:3000 |
exactly that scheme, host and port |
kvraudio.com |
https on the default port, that host only — not www.kvraudio.com, not http:// |
*.example.com |
https on the default port, any subdomain at any depth — not example.com itself |
Hosts are compared as parsed (scheme, host, port) fields, never as text, so
https://kvraudio.com.evil.io, https://kvraudio.com@evil.io and
https://kvraudio.com:8443 are not kvraudio.com. A site that redirects between
example.com and www.example.com needs both (example.com and
*.example.com). An entry that cannot be read is dropped and the rest still load:
empty and malformed entries, a lone *, single-label wildcards (*.com),
wildcards on an IP address, IPv6 literals, paths, userinfo, a port on a bare host
(write http://localhost:3000), and any scheme but http(s) — file:, data:,
about: and blob: name no site. The list does not know public suffixes:
*.co.uk is two labels and would be accepted, so do not write it.
What it covers is deliberately small. It never answers SUBMIT, UPLOAD,
PUBLISH or DOWNLOAD: those stay single-use and are still
asked on a listed site, every time. (To pre-approve sending from your own product for
end-to-end runs, LYRA_BROWSER_TRUSTED_SEND_ORIGINS is a separate list that answers
SUBMIT and UPLOAD and nothing else.) The list also holds where the browser arrives without
a navigate call -- a redirect hop or a followed link to a listed site -- because the guard
applies it to NAVIGATE too. It does not apply to an opaque origin, while
the user holds a takeover, or when approval is off (that stays
consent_channel=off). A listed site gets a real grant with the same lifetime
and initiator binding an approved one would, so enforcement judges it identically,
and the audit records consent_channel=trusted with the entry that matched.
trusted is a way a decision was reached, not a value of
LYRA_BROWSER_CONSENT_CHANNEL. It applies under auto, elicit and legacy.
Takeover. While the user holds the session, every mutating tool returns
takeover_active and the agent waits — including highlight_element and
ask_user_to_do, which write to the page. Reading tools keep working.
Enforcement steps aside too. Grants record what the agent was allowed to do;
holding a person to them would refuse them their own browser, and takeover exists
for the passwords, CAPTCHAs, 2FA and payments the agent must not do alone — every
one of which is a navigation. Those are recorded as user_driven, so the trail
distinguishes what a person did from what was approved for the agent.
Enforcement has to see every document request before it leaves the browser. There
are two ways to stand there, chosen with LYRA_BROWSER_GUARD (route unless set);
both ask the same judgement (NavigationGuard.decide), so grants, single-use
scopes, takeover, observe mode and the audit trail behave identically. They differ in
what they can see and what they cost.
route (default) |
cdp |
|
|---|---|---|
| How | Playwright's context.route("**/*") |
A second DevTools connection: Fetch.enable for Document requests on every page and every out-of-process iframe |
| Redirect hops | Judged after they have left: a context.on("request") listener cancels the load (the request to the first refused hop still goes out; on loopback the page commits first and is blanked) |
Every hop judged before it leaves: the destination of a refused hop sees nothing |
| HTTP cache | Off: Playwright sends Network.setCacheDisabled(true) with any route |
On |
| etsy.com (DataDome), headful, fresh profile, N=5 | refused 5/5 | let in 5/5 |
| Service workers | service_workers="block" is an init script a page can step around; a navigation a worker answers is never seen |
Bypassed on every guarded page (Network.setBypassServiceWorker) |
| Open port | none | a debugging port on 127.0.0.1, see below |
| Guard lost | cannot happen | the browser is killed and every call answers guard_lost |
| Launch | Playwright's defaults | adds --remote-debugging-port=0, drops --enable-unsafe-swiftshader |
| Dependencies | none | none: a stdlib WebSocket client (cdp_socket.py, ~250 lines), not the websockets package |
How cdp guards. Chrome is started with --remote-debugging-port=0 next to the
pipe Playwright uses, and cdp_guard.py opens a second connection to it. It
auto-attaches to every page and out-of-process iframe at browser level with
waitForDebuggerOnStart, so a popup or a cross-site iframe is held until
Fetch.enable is in place on it (a session made through Playwright after the tab
exists cannot promise that: popups' first requests and cross-site iframe requests
escaped in the earlier probe). Each paused request becomes the two objects
classify reads, goes through decide, and is answered Fetch.continueRequest or a
204 Fulfill — never an abort. Nothing is asked of the page while a request is
paused (its renderer does not answer then).
Redirect hops. Each hop is its own Fetch.requestPaused, judged as the
navigation it is: a 307/308 keeps method and body, a 301/302/303 turns a POST into a
body-less GET. A site you approved that answers 302 to an unapproved origin is
refused before the request leaves — the destination sees nothing — and the audit row
names redirect_from, and navigate answers blocked_by_policy with redirected_to. A hop that stays with the origin that issued it (the same
origin, or http→https of the same host on default ports) is not judged again: it is
the request carrying on, and asking again would strand every trailing-slash redirect on
a form.
Uploads. Chrome puts a navigation's whole body into the event (as text and as base64: a 40 MB form is a 98 MB message). The transport reads such a message through without keeping it and the guard judges the event without the body (it only asks whether one exists). A 3, 12 and 40 MB multipart POST were refused when ungranted, delivered when granted, and the guard stayed up.
If the guard is lost. Chrome releases every request it holds, and every target
waiting for the debugger, the moment the connection ends, so a lost connection is a
browser running unguarded. The session therefore kills the browser, records
guard_lost in the audit trail and answers every call guard_lost (also to readers,
and to open_browser) until close_browser acknowledges it; the next browser is a
fresh, guarded one. A socket that ends because the browser quit (a crash, the user
closing the last window) is told apart without waiting: the process is already gone, or it
is running with no open tab and leaves within 300 ms. With a tab open nothing explains the
socket, so it is killed at once. Measured (scripts/verify_cdp_guard_e2e.py): socket
aborted, process gone in 1–2 ms, and a hostile page posting a form every 4 ms got 0 of
them through to its server in 3 of 3 runs (waiting 300 ms first let 72 through). Requests
already in flight at that instant were released by Chrome and are not judged. A sidecar
that cannot start, or a target it cannot guard, is the same event: the window is closed
rather than left open.
Debugging port. --remote-debugging-port is an open, unauthenticated service, and
this is what cdp costs.
- What it exposes. Any process on the machine that finds the port can attach to the
browser and do what the agent can and more: list tabs, read cookies (HttpOnly ones
too), run script in any tab of the logged-in profile. Finding it needs no secret:
GET /json/versionhands out the WebSocket path. Withroutethere is no port; Playwright drives Chrome over a pipe only it holds. Reading the profile directly would also give a same-user process the cookies, so the difference is mostly other users of the host and sandboxed code that has loopback but not the profile directory; on a one-user workstation it is small, on a shared host it is the whole profile. - What is done. Random port, chosen by Chrome, bound to
127.0.0.1only (::1refuses; measured). Chrome refuses a handshake that carries anOriginheader (a web page cannot connect, 403) or a foreignHost(no DNS rebinding, 500). The profile directory is made0700before launch;DevToolsActivePort(written0664, and left behind when Chrome exits) is removed before launch and as soon as it is read. The port and the path are never put in an exception, a log line, an audit row or an envelope. - What remains. The port exists as long as the browser does; Chrome has no way to close it or to require a token. A second consumer cannot use the pipe (Chrome serves one, and Playwright holds it); putting a proxy in front of a pipe-launched Chrome means launching Chrome ourselves, which this does not do. Chrome 136 and later ignore the flag on their default profile directory because the port is a cookie-theft vector [Chrome's own announcement, not measured here]; this server never uses that directory.
- Measured by the
exposuresection ofscripts/verify_cdp_guard_e2e.py, which attaches from a separate process the way a scanner would.
What neither backend sees. fetch()/XHR (not navigations), WebSockets, GETs with the
data in the URL, speculation-rules prefetch/prerender and the navigation that activates
one (above), and data:/about:/blob:/javascript: URLs (not requests). Workers never
issue a document request and are not attached; a page's own back_forward_cache is off
under Playwright's switches, and history navigations are judged with the cache on.
Choosing. cdp for sites behind DataDome, where redirect chains matter, or where
a page's service worker must not answer navigations; route where an open loopback
port is unacceptable. Every scripts/verify_*_e2e.py gate runs on either:
LYRA_BROWSER_GUARD=cdp python scripts/verify_tabs_e2e.py.
End users (VEGA.app): they do not run pip or playwright install. VEGA
bundles the playwright Python package in its own runtime, and this server drives
the user's already-installed Chrome or Edge — zero downloads. If they have
neither, a browser_unavailable envelope tells VEGA's UI to prompt a one-click
Chrome install. See docs/VEGA_INTEGRATION.md.
Developers (working on this repo):
pip install -e ".[dev]"
python -m playwright install chromium # only needed if you have no system Chrome
lyra-browser # serve over stdio (what VEGA spawns)
# or: lyra-browser --http --port 8765 # serve over HTTP for dev| Var | Default | Meaning |
|---|---|---|
LYRA_BROWSER_CLIENT |
detected | vega, hermes or generic (also --client). Picks default paths and window mode only, never permissions. Detected from VEGA_DATA_DIR → vega, HERMES_HOME → hermes. |
LYRA_BROWSER_DATA_DIR |
— | Base data dir (profile + audit). Overrides everything. |
VEGA_DATA_DIR |
— | vega: data goes to $VEGA_DATA_DIR/browser. |
HERMES_HOME |
~/.hermes |
hermes: data goes to $HERMES_HOME/browser. |
LYRA_BROWSER_HEADLESS |
auto | Run without a visible window (see Headless mode). Unset: headless for hermes, and on Linux with no DISPLAY/WAYLAND_DISPLAY; otherwise a window. |
LYRA_BROWSER_VIEWPORT |
1280x800 |
WIDTHxHEIGHT of the page (390x844 for a phone layout). Headless uses it as the emulated viewport; headful as the window size. Anything unreadable falls back to the default. |
LYRA_BROWSER_PROXY |
unset | Route the browser through a proxy, e.g. socks5://100.x.y.z:1080. For UAT runs that must not share the operator's egress IP, since per-IP rate limits, quotas and IP-based analytics exclusion all key on it. Unset means a direct connection. |
LYRA_BROWSER_REQUIRE_APPROVAL |
true |
Ask before risky actions. false sets the consent channel to off. |
LYRA_BROWSER_CONSENT_CHANNEL |
auto |
auto asks the user over MCP and falls back to confirm=true only on a client that cannot be asked; elicit pins asking and denies otherwise; legacy always takes the model's word; off does not ask. |
LYRA_BROWSER_ENFORCEMENT |
enforce |
observe records what it would have blocked without blocking: navigations go through, and an undeclared download is saved (observed: true, audited would_block) instead of cancelled. |
LYRA_BROWSER_GRANT_TTL |
600 |
Seconds a grant stays usable. Single-use ones ignore this. |
LYRA_BROWSER_TRUSTED_ORIGINS |
— | Sites pre-approved for NAVIGATE/INTERACT, comma- or whitespace-separated: https://host[:port], a bare host (https only), or *.example.com (https subdomains, not the apex). Never covers SUBMIT/UPLOAD/PUBLISH/DOWNLOAD. See Trusted origins. |
LYRA_BROWSER_TRUSTED_SEND_ORIGINS |
— | Sites where sending is also pre-approved — SUBMIT and UPLOAD — for acceptance runs against your own product. Same entry forms, separate list: a site in TRUSTED_ORIGINS is not send-trusted by being there. Never PUBLISH or DOWNLOAD; each approval stays single-use, and it steps aside during a takeover. Audited as consent_channel=trusted_send. |
LYRA_BROWSER_CONSENT_TIMEOUT |
300 |
Seconds a prompt waits for an answer before it counts as a denial (timed out waiting for the user). The tool call is held that long, so for an unattended Hermes gateway set it lower — e.g. 60. |
LYRA_BROWSER_RELEASE_GRACE |
2 |
Seconds before a finished action's unused single-use scope is reclaimed. |
LYRA_BROWSER_OWNER_IDLE_TIMEOUT |
900 |
Seconds a session may hold the browser without using it. Handing it on closes the window. |
LYRA_BROWSER_CAPTURE_DIR |
— | Where screenshot/read_image write PNGs. Default: vega $VEGA_DATA_DIR/uploads/browser (where VEGA attaches images), hermes $HERMES_HOME/cache/browser (where Hermes' vision may read), else <data_dir>/captures. |
LYRA_BROWSER_CAPTURE_KEEP |
200 |
Captures kept on disk; oldest are deleted first. 0 keeps all. |
LYRA_BROWSER_DOWNLOAD_DIR |
— | Where declared downloads are saved. Default <data_dir>/downloads. |
LYRA_BROWSER_DOWNLOAD_KEEP |
200 |
Saved downloads kept; oldest are deleted first, and only files this server saved (tracked in a ledger in the dir). 0 keeps all. |
LYRA_BROWSER_DOWNLOAD_MAX_BYTES |
209715200 |
Largest file a download may be (200 MiB); a bigger one is cancelled and answered download_failed (too_large). |
LYRA_BROWSER_DOWNLOAD_TIMEOUT |
120 |
Seconds a declared download may take to arrive once it has started before it is cancelled (download_failed, timeout). |
LYRA_BROWSER_CHANNEL |
— | Force one channel (chrome/msedge). Unset = try chrome→msedge→bundled. |
LYRA_BROWSER_ALLOW_BUNDLED |
true |
Allow bundled-Chromium fallback (only exists after playwright install). |
LYRA_BROWSER_DRIVER |
auto |
Which Playwright to launch with: auto takes patchright when installed and falls back to playwright; either name pins it. open_browser reports the one in use as driver. |
LYRA_BROWSER_GUARD |
route |
Where the navigation guard stands: route (Playwright's route) or cdp (a second DevTools connection: every redirect hop judged, cache on, DataDome lets the browser in, a loopback debugging port, guard_lost if the guard fails). See Guard backends. |
LYRA_UAT_CLAUDE_BIN |
claude on PATH |
The Claude Code executable the claude-code UAT brain runs. See UAT mode. |
LYRA_UAT_CODEX_BIN |
codex on PATH |
The Codex executable the codex UAT brain runs. See UAT mode. |
lyra-uat plays a persona — who is visiting, from where, in what language, with what goal and how many steps — against a site, through the tools above, and writes report.json. The model that plays it is replaceable: Claude Code, Codex, the Anthropic API, OpenRouter, Ollama Cloud, or any OpenAI-compatible endpoint. What makes a run trustworthy is not: the server, not the model, keeps the trace. Whichever brain drives, every tool call is recorded as it happens, the step budget is enforced, a screenshot is taken after each action, a type_text value that looks like a payment-card number is refused, an upload outside the persona's directory is refused, and leaving the sites under test cannot be approved by the model's own confirm=true.
pip install -e ".[uat]"
lyra-uat demo-site --port 8787 & # a small site with things to find
lyra-uat run examples/uat/demo/run.yaml # one persona, Claude Code plays it
lyra-uat batch examples/uat/demo/batch.yaml # two personas in parallel, one process eachA run leaves a directory: report.json (and report.md), trace.jsonl, events.jsonl, captures/, network.jsonl, console.jsonl, and the browser's own audit.jsonl. Persona, target, brain, hooks and limits are YAML or JSON; lyra-uat schema run|batch|report prints the JSON Schemas. Details, the report format and each brain's requirements are in docs/UAT.md.
lyra-browser --headless # or LYRA_BROWSER_HEADLESS=true; the flag wins
lyra-browser --no-headless # force a window even if the env says otherwiseUnset, the mode follows the client and the host: --client hermes is headless
(Hermes reaches the user over chat and gives its servers no display), and so is
any Linux process with neither DISPLAY nor WAYLAND_DISPLAY — launching a
window there would crash, headless is a working session that says so.
Headless does not merely hide the window — it means nobody is watching. The
collaboration tools have no human to reach, so they refuse with an unattended
envelope instead of reporting success for something no one will see:
| Tool | Headless behaviour |
|---|---|
highlight_element |
unattended — an outline nobody sees is not a signal |
ask_user_to_do |
unattended — there is no user at this window to ask |
request_takeover |
unattended, and control is not handed over — granting a takeover no human can release would block every later mutation |
Headless Chrome also announces itself: its UA says HeadlessChrome, and
Cloudflare's managed challenge refuses on that token alone (13 of 34 login
pages in the earlier survey). The session therefore sends the UA the same
binary would send with a window — the real version, the real platform, only
that token dropped — so navigator.userAgentData and the Sec-CH-UA-* headers
stay the browser's own. The version is read from the running Chrome on the
first headless launch and remembered in <data_dir>/browser-ua.json; that
first launch, and the one after each Chrome update, relaunch once (about 0.7s).
Navigation, interaction, and reading are unaffected. open_browser reports which
mode you are in via its attended field so the agent never assumes an audience.
Sites that sort agents from people check the launch, not the behaviour. The
session is launched the way Playwright's own MCP server launches — with
--disable-blink-features=AutomationControlled, so navigator.webdriver is
false as in any Chrome a person opens, and without an emulated viewport in
headful mode, so the page sees the real screen instead of a window larger than
the screen it is on. Measured on 42 login pages behind Cloudflare, Akamai,
DataDome, PerimeterX and Kasada: of the 34 plain Chrome passes,
the Playwright defaults failed 6 headful and 19 headless; with these settings
the count is in the survey comment, per site.
What is deliberately not done: the --enable-automation infobar stays (it is
how a person at the shared window is told what is happening), and nothing is
forged — no WebGL or canvas noise, no patched CDP. Route interception and the
service-worker block contribute nothing to the Cloudflare and Akamai refusals; for
DataDome the route does matter, which is why the cdp guard backend exists.
Two vendors still refuse, and the cause of each is measured rather than guessed:
- Kasada (hyatt.com) notices one CDP call,
Runtime.enable. Plain Chrome driven over raw CDP passes; the same Chrome withRuntime.enableon is refused; every other call Playwright makes on attach is fine. Playwright enables it on every page and has no switch not to. patchright is a drop-in fork that does not send it and evaluates in an isolated world instead; withpip install "lyra-browser[patchright]"the session launches through it (LYRA_BROWSER_DRIVER=auto), and hyatt passes in both modes. What patchright changes for this server:page.evaluateruns outside the page's own JavaScript world (our evaluate calls only touch the DOM, which is shared) and console messages are not delivered (we read none). - DataDome (etsy.com; tripadvisor.com does not discriminate in the
measurements) refuses on either of two causes, each sufficient on its own
(N=5–6 per cell, headful, one IP):
context.routeswitches the HTTP cache off, so every document request carriesCache-Control/Pragma: no-cache(plain Chrome with onlyNetwork.setCacheDisabledis refused 0/6, while raw CDPFetch.enableon documents or on all requests is fine), and on a GPU-less host--enable-unsafe-swiftshadergives the page a WebGL that plain Chrome lacks. Thecdpguard backend removes both: no route, documents only over a second connection, and the switch dropped. Measured on that backend with the product itself (scripts/measure_compat.py, headful, fresh profile per visit, 12 s apart, N=5): etsy.com let in 5/5 against 0/5 for theroutelaunch; hyatt.com (patchright) 5/5 against 4/5 (the one miss was a 429). Dropping the switch removes WebGL on a host with no GPU, as in plain Chrome there; on a host with a GPU it is inert [INFERENCE: no GPU host here]. Headless is undecidable on the measuring host.
ruff check . # lint
pre-commit install # enable hooksReal-browser gates live in scripts/verify_*_e2e.py, for
example python scripts/verify_browser_e2e.py --headless-only, or headful under
xvfb-run -a.
-
CDP guard gate —
python scripts/verify_cdp_guard_e2e.py(both modes,--driver) re-runs the browser, script-redirect and tabs gates on thecdpbackend, then checks what only it can do: redirect hops (301/302/303/307/308, with and without a body), popups at N=10, out-of-process iframes, service workers, large uploads, kill-the-socket fail-closed, a sidecar that cannot start, and what another local process can do with the debugging port.scripts/measure_compat.pyis the etsy/hyatt measurement above. -
Real-sites gate —
python scripts/verify_realsites.py(needs the internet) visits ten live sites and a loopback control, reads each as a tree and as text, clicks its first internal link through thearia-ref=the tree printed, and compares judged and blocked navigations and page sizes withscripts/realsites_baseline.json(±10% plus a small slack; success must match exactly). A site that cannot be reached isSKIP, and the gate fails only when fewer than eight ran. It runs against live sites, so a site editing its header can turn it red with no code change: treat it as a report to read, not a merge gate (the loopbackverify_*_e2e.pygates are the blocking ones).--update-baselinere-records it,--sites a,band--retries Nnarrow and steady a run,--driver patchrightand--headfulpick the driver and the window.
The server, session manager, tools, permission layer, enforcement and audit are
implemented, unit-tested (no browser needed) and exercised end-to-end
against real sites with a real Chrome — ten rounds over eleven sites, 187 judged
navigations. Not yet driven by a live VEGA session; that is the next milestone,
and it is what would let consent_channel=elicit replace the confirm flag.
MIT