Repository navigation
Run agents on a local model - #36
Merged
Merged
Conversation
Floors the pin at boundflow 0.7.4. Below it every Ollama run failed: the output cap was sent as `max_tokens` and a spent `max_llm_calls` forced the ending tool with `tool_choice`, neither of which Ollama's client takes. Adds tests/e2e/test_ollama.py, the quickstart scaffold against a model on this machine with nothing faked, and a README section on what differs from a hosted provider — call timeouts, budgeting in calls rather than dollars, and model size.
arjunvlama
force-pushed
the
ollama-local-models
branch
from
September 16, 2026 00:06
34ae7a6 to
808e46e
Compare
boundflow 0.7.4 offers only submit_result on the last call a budget allows, so a model that answers in text ends the run there rather than on a further call that would have raised. The failure reason said only that nothing was submitted, which doesn't tell an operator whether to raise the number or fix the agent. The scripted test model read that narrowed offer as a subagent's, since it tells them apart by the absence of `task`, and answered prose a turn early.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Charter couldn't run on Ollama. Two call arguments that Ollama's client rejects stopped every run, both fixed in BoundFlow 0.7.4 (#114); this floors the pin there and covers the path from Charter's side.
TypeError: AsyncClient.chat() got an unexpected keyword argument 'max_tokens'max_llm_callsis spentTypeError: ... unexpected keyword argument 'tool_choice'tests/e2e/test_ollama.pyThe agent is
charter init's own scaffold, so this is the README's quickstart with the model line changed. Nothing is faked — including the provider, which Charter builds fromworker.yaml'sprovider: ollamarather than being handed by the test.test_a_local_model_completes_a_task— the run succeeds and the injectedresponse_formatcomes back filled.test_a_run_out_of_calls_ends_on_the_finalizer—max_llm_calls: 1, so the first call is also the forced ending. Ollama can't force a tool at all, so the ending holds because the finalizer is the only tool left on offer.Verified against Ollama 0.34.0,
qwen2.5:7b, CPU only:boundflow0.7.4boundflow0.7.0TypeError: ... unexpected keyword argument 'max_tokens', the original bugOLLAMA_TEST_MODELunsetThey skip without
OLLAMA_TEST_MODEL, and carry their owntimeoutmark because the suite's 180s is shorter than a single call on CPU.Not in CI, deliberately. A model small enough to pull on every run is too small to drive the harness —
qwen2.5:3banswered the ticket by writing a file instead of callingsubmit_result. The SDK-level contract with Ollama is covered by BoundFlow'sPython SDK (Ollama)job, which can use a 0.5B model because it asserts on the call rather than on the agent.A failure reason that lost its number
0.7.4 offers only
submit_resulton the last call a budget allows. A model that answers in text instead ends the run there, rather than on a further call — and that further call was where BoundFlow raisedreached max_llm_calls=N. So the reason read only "the agent stopped without calling submit_result", which doesn't tell an operator whether to raise the ceiling or fix the agent._publishnow names the ceiling when the budget is spent. New test,test_a_budget_spent_on_an_answer_in_prose_still_names_the_ceiling, which fails without it.This also turned
tests/e2e/test_failures.py::test_a_spent_budget_says_which_ceiling_it_hitred — not for that reason. The scripted model tells a subagent's call from the parent's by the absence oftaskin the offered tools, so it read the narrowed finalize offer as a subagent's and answered prose a turn early. Fixed in the fake: a bind of the finalizer alone is the last call, not a subagent.Worth noting
mainwould fail the same way — CI installsboundflow>=0.7.0, which now resolves to 0.7.4.Docs
A new docs/local-models.md, linked from the README's documentation list. Notes from the runs rather than prose: the config, and the three things measured against Ollama 0.34.0 on CPU.
max_call_secondsdefaults to 60; a single quickstart call took about 170s.max_cost_usdnever binds —max_llm_callsis the ceiling that does. Usage is still reported, so runs are metered at $0.qwen2.5:7bcompletes the quickstart;qwen2.5:3banswered by writing a file instead of callingsubmit_result.Plus the context window being the server's (
OLLAMA_CONTEXT_LENGTH), and that forcing the ending is weaker on Ollama, which can't force a tool at all.The README itself is unchanged except for one line in the documentation list.