Skip to content

Run agents on a local model - #36

Merged
arjunvlama merged 9 commits into
mainfrom
ollama-local-models
Sep 16, 2026
Merged

arjunvlama merged 9 commits into
mainfrom
ollama-local-models

Conversation

@arjunvlama

@arjunvlama arjunvlama commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Charter couldn't run on Ollama. Two call arguments that Ollama's client rejects stopped every run, both fixed in BoundFlow 0.7.4 (#114); this floors the pin there and covers the path from Charter's side.

When Error
Output cap every call TypeError: AsyncClient.chat() got an unexpected keyword argument 'max_tokens'
Forced finalizer the last permitted call, once max_llm_calls is spent TypeError: ... unexpected keyword argument 'tool_choice'

tests/e2e/test_ollama.py

The agent is charter init's own scaffold, so this is the README's quickstart with the model line changed. Nothing is faked — including the provider, which Charter builds from worker.yaml's provider: ollama rather than being handed by the test.

  • test_a_local_model_completes_a_task — the run succeeds and the injected response_format comes back filled.
  • test_a_run_out_of_calls_ends_on_the_finalizer — max_llm_calls: 1, so the first call is also the forced ending. Ollama can't force a tool at all, so the ending holds because the finalizer is the only tool left on offer.

Verified against Ollama 0.34.0, qwen2.5:7b, CPU only:

SDK Result
boundflow 0.7.4 2 passed
boundflow 0.7.0 both fail with TypeError: ... unexpected keyword argument 'max_tokens', the original bug
OLLAMA_TEST_MODEL unset 2 skipped

They skip without OLLAMA_TEST_MODEL, and carry their own timeout mark because the suite's 180s is shorter than a single call on CPU.

Not in CI, deliberately. A model small enough to pull on every run is too small to drive the harness — qwen2.5:3b answered the ticket by writing a file instead of calling submit_result. The SDK-level contract with Ollama is covered by BoundFlow's Python SDK (Ollama) job, which can use a 0.5B model because it asserts on the call rather than on the agent.

A failure reason that lost its number

0.7.4 offers only submit_result on the last call a budget allows. A model that answers in text instead ends the run there, rather than on a further call — and that further call was where BoundFlow raised reached max_llm_calls=N. So the reason read only "the agent stopped without calling submit_result", which doesn't tell an operator whether to raise the ceiling or fix the agent.

_publish now names the ceiling when the budget is spent. New test, test_a_budget_spent_on_an_answer_in_prose_still_names_the_ceiling, which fails without it.

This also turned tests/e2e/test_failures.py::test_a_spent_budget_says_which_ceiling_it_hit red — not for that reason. The scripted model tells a subagent's call from the parent's by the absence of task in the offered tools, so it read the narrowed finalize offer as a subagent's and answered prose a turn early. Fixed in the fake: a bind of the finalizer alone is the last call, not a subagent.

Worth noting main would fail the same way — CI installs boundflow>=0.7.0, which now resolves to 0.7.4.

Docs

A new docs/local-models.md, linked from the README's documentation list. Notes from the runs rather than prose: the config, and the three things measured against Ollama 0.34.0 on CPU.

  • max_call_seconds defaults to 60; a single quickstart call took about 170s.
  • An unpriced model reports no cost, so max_cost_usd never binds — max_llm_calls is the ceiling that does. Usage is still reported, so runs are metered at $0.
  • qwen2.5:7b completes the quickstart; qwen2.5:3b answered by writing a file instead of calling submit_result.

Plus the context window being the server's (OLLAMA_CONTEXT_LENGTH), and that forcing the ending is weaker on Ollama, which can't force a tool at all.

The README itself is unchanged except for one line in the documentation list.

Floors the pin at boundflow 0.7.4. Below it every Ollama run failed: the output
cap was sent as `max_tokens` and a spent `max_llm_calls` forced the ending tool
with `tool_choice`, neither of which Ollama's client takes.

Adds tests/e2e/test_ollama.py, the quickstart scaffold against a model on this
machine with nothing faked, and a README section on what differs from a hosted
provider — call timeouts, budgeting in calls rather than dollars, and model size.
boundflow 0.7.4 offers only submit_result on the last call a budget allows, so a
model that answers in text ends the run there rather than on a further call that
would have raised. The failure reason said only that nothing was submitted, which
doesn't tell an operator whether to raise the number or fix the agent.

The scripted test model read that narrowed offer as a subagent's, since it tells
them apart by the absence of `task`, and answered prose a turn early.
@arjunvlama
arjunvlama merged commit 1aa45cc into main Sep 16, 2026
3 checks passed
@arjunvlama
arjunvlama deleted the ollama-local-models branch September 16, 2026 00:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant