A minimal, customizable and auditable chat interface for provider-native LLM
APIs. Claude-first, built on LLM and
llm-anthropic. Independent and
unofficial.
Customize every call. Inspect every response. Keep every conversation.
"Native" here means the provider's own API, not a lowest-common-denominator
compatibility layer: the request carries the provider's real parameters, and
the right pane shows the actual SDK call the run was about to make. The first
provider is Anthropic through llm-anthropic, which uses Anthropic's official
Python SDK.
Status: first conversation slice. The chat pane now runs a real multi-turn conversation through
llm.Conversationandllm-anthropic, and the right pane renders the parameters that path actually built. No test in this repository contacts a provider: the Anthropic transport is faked, so no API key is needed and no paid call is made. Running the app against the real API requires your own key, configured the usual LLM way.
llm native-chatstarts a local ASGI server on127.0.0.1.- The left pane sends a message into an
llm.Conversation; replies stream back over/api/chat/streamas server-sent events. - The right pane has two views of the same turn. Request renders the
anthropicPython SDK call frommodel.build_kwargs()— the dictionaryllm-anthropichands to the SDK. There is no second copy of the request: the renderer takes that dictionary and nothing else. Response shows what came back, read off the Message the SDK finished with: content blocks, thinking, citations, server tool blocks, stop reason, input/output/cache tokens and web search requests, as formatted JSON plus a one-line usage summary. The left bubble and the Response pane are projections of the same record, so they cannot disagree: clicking the user's bubble opens the Request that went out, and clicking the assistant's opens the Response that came back. - Each bubble shows the time of its turn in your computer's own zone, rendered through the browser's locale. The stored stamp is UTC, so the page converts it; a turn with no stamp shows no time rather than "now".
- Conversations are saved and can be reopened. Storage goes through LLM's own
SQLite schema (
llm.logs.LogStore), so these conversations are not a second chat database and they show up inllm logsunchanged. Titles come from the first message; no model call is made to name a conversation.llm-anthropicalways opensclient.messages.stream(), so that is what the pane shows; a non-streamingmessages.create()is never rendered as if it happened. - The request's settings are eight pills in the topbar — Model, Thinking, Effort, Max tokens, System, Web Search, Caching, Streaming — each opening a small menu that writes back into one hidden form, so what Send builds always matches what the pills show. A pill the model or the runtime cannot honour is dashed and grey with the reason attached, never a fake control.
- Every form control reports a status, so a greyed-out value says why:
Editable,API supported · runtime fixed,Unsupported by selected model,Unsupported by current tool version,Provider default, orFallback capability data. - Model facts come from the Anthropic Models API when a key is present (cached on disk, refreshed in the background), and from a versioned fallback profile when it is not. The UI always says which one it is showing.
- If no Anthropic key is available, the turn fails locally and nothing is sent.
- Offline tests drive the real LLM Python API with a faked transport: they assert history, streaming, the four Anthropic behaviours, and that the code matches the request that was sent.
Not implemented: any provider other than Anthropic, any paid or authenticated call in the test suite, hand-written conversation titles, and history management beyond opening a saved conversation (no search, delete or export).
Every completed turn is written to LLM's own database (the one llm logs
reads) through llm.logs.LogStore:
| Endpoint | Purpose |
|---|---|
GET /api/conversations |
saved conversations, newest first |
GET /api/conversations/{id} |
one conversation, with each turn's request and response |
Resuming passes llm's stored messages back in, so history keeps its original content blocks instead of being rebuilt from chat text. The request this app built and the code it rendered are not part of llm's schema, so they go in one sidecar table keyed by turn id — written in the same transaction, so a failed save leaves nothing behind and is reported as a failure rather than as success.
The first Anthropic integration is expected to prove these four behaviours:
web_search_20260318- Dynamic filtering, without forcing
allowed_callers=["direct"] response_inclusion="excluded"- Prompt caching
Their current status against llm-anthropic is recorded in
docs/upstream-contributions.md. Items 1, 2
and 4 are already properties of the released plugin and are asserted by offline
tests. Item 3 needs a plugin that accepts response_inclusion; the released
llm-anthropic does not, so CI on this branch installs a pinned commit of the
contribution branch and the assertion is skipped when that capability is
absent.
- Python 3.10 or later
llm>=0.36llm-anthropic>=0.29
Verified on Python 3.13.12 with llm==0.36 and llm-anthropic==0.29.
With uv:
uv sync --extra test
source .venv/bin/activateOr with pip:
python -m venv .venv
source .venv/bin/activate
pip install -e '.[test]'llm native-chatThen open http://127.0.0.1:8000. The server binds to 127.0.0.1 by default
and warns if told otherwise.
Without activating the environment, use uv run --extra test llm native-chat.
The same application can be started directly as a module:
python -m native_api_chatjustOr:
pytest
ruff check .Part of the suite drives the page in a real browser, because reading the source cannot see a clipped menu or a click the browser never dispatched. Those checks skip themselves unless the extra is installed; Playwright uses an installed Google Chrome, so there is usually nothing to download:
pip install -e '.[test,browser]'
pytest -m browser # only the real-page checks
pytest -m "not browser" # skip themOn a machine with no Chrome, playwright install chromium provides one.
This is an independent, unofficial project. It is not maintained, endorsed, affiliated with, or sponsored by Simon Willison, the LLM project, or Anthropic. It is a plugin and a local UI built on top of their published, open-source Python packages. "LLM", "llm-anthropic" and "Anthropic" are the marks of their respective owners and are used here only to describe what this project depends on.
LLM already provides the provider plugin model, conversations, streaming
abstractions, SQLite logging, key management, structured messages, and usage
tracking. llm-anthropic uses Anthropic's official Python SDK. This project
should contribute generally useful provider capabilities upstream instead of
maintaining a broad private fork.
Checked against the installed llm==0.36 and llm-anthropic==0.29 source,
three of the four target behaviours already work upstream and one does not:
web_search_20260318— already emitted for adaptive-thinking models;- dynamic filtering — already preserved,
allowed_callersis never sent; - prompt caching — already applied to the final message with
options.cache; response_inclusion— not exposed, and the only upstream change needed.
The first upstream contribution is therefore narrowly scoped to a single optional argument:
WebSearch(response_inclusion="full" | "excluded")
It has been submitted as
simonw/llm-anthropic PR #95
and this project waits for it rather than carrying a private fork. Status and
follow-up steps are in
docs/upstream-contributions.md.
native_api_chat/ Python package and LLM plugin
app.py Local ASGI application
capabilities.py Model capability matrix, from three ranked sources
chat.py The conversation path and the form schema
cli.py `llm native-chat` command
codegen.py Provider SDK code generation from build_kwargs()
model_api.py Optional Anthropic Models API layer with a disk cache
models.json Versioned fallback capability profile
static/index.html Minimal two-pane UI
tests/ Focused unit tests
docs/ Architecture, contribution and upstream notes
- No API key is required to install, test, or run this project today.
- When model execution is added, keys must stay in LLM's key store or in
environment variables. Generated code reads
ANTHROPIC_API_KEYfrom the environment and never embeds a secret. - The development server binds to
127.0.0.1by default. - Model execution will not be enabled until request construction can be inspected and tested without accidental paid calls.
- To report a vulnerability, see
SECURITY.md.
Apache License 2.0, matching LLM and llm-anthropic. See LICENSE.