Skip to content

feat: scenario-scoped MCP tool fakes in fixtures - #490

Merged
jpr5 merged 9 commits into
mainfrom
feat/mcp-fakes
Oct 5, 2026
Merged

jpr5 merged 9 commits into
mainfrom
feat/mcp-fakes

Conversation

@jpr5

@jpr5 jpr5 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

What & why

Fixture files can now declare MCP tool fakes for a scenario, in an mcpFakes block next to the LLM turns of the same test. Each block names its scenario (a test id or a context). Inside a block you can:

  • script ordered answers per tool: the first call gets a timeout, the retry gets a result;
  • match arguments exactly, or accept any args for a call;
  • run a closed world (undeclaredTools: "deny"), so a call to any undeclared tool is an error;
  • fail loud: a mismatch, an exhausted tool or an undeclared tool under deny returns a JSON-RPC error with a stable code in error.data.aimock.code. The failure is also written to the journal, a metric and the server log.

An agent test has two inputs that change from run to run: what the LLM says, and what the tools return. aimock already pins the LLM side. For MCP tools, aimock could only give one fixed answer per tool for the whole server process. A test could not say "in this scenario, get_weather(Seattle) returns rain the first time and a timeout the second time, and any other call is a bug." So teams either ran real MCP servers in tests (side effects, drift, flaky runs) or wrote code handlers that cannot be shared as data and do not reset per test.

This works at the wire, over Streamable HTTP, so it is the same for LangGraph, CrewAI, ADK, PydanticAI, Mastra or a plain SDK client, in any language. In-process tool mocks, such as Mastra's Experiment Tool Mocks, only cover agents of one framework. Because the fakes live in fixture files, a later recording mode can write real MCP tool results into the same file as the LLM turns, so one file holds the whole scenario.

How to use

tickets/retry.json (from src/__tests__/fixtures/mcp-fakes/tickets/retry.json):

{
  "fixtures": [
    {
      "match": { "userMessage": "open a refund ticket", "hasToolResult": false },
      "response": {
        "toolCalls": [{ "name": "create_ticket", "arguments": { "title": "Refund" } }]
      }
    }
  ],
  "mcpFakes": {
    "scope": { "testId": "tickets › retry on timeout" },
    "tools": [
      {
        "name": "create_ticket",
        "calls": [
          { "id": "first-try", "args": { "title": "Refund" }, "error": "upstream timeout" },
          { "id": "retry", "args": { "title": "Refund" }, "result": "TICKET-42" }
        ]
      }
    ]
  }
}

Start aimock with --fixtures on that directory. aimock mounts an MCP mock at /mcp if there is none. Point the agent's MCP client at /mcp?testId=tickets%20%E2%80%BA%20retry%20on%20timeout (or send X-Test-Id). The first create_ticket call gets isError: true "upstream timeout", the second gets TICKET-42, and a third fails with MCP_FAKE_EXHAUSTED.

What changed

  1. Fake model, fixture contracts and request helpers. The mcpFakes types, the validator, the fake store (scope, ordering, matching, consumption, eviction), the error codes and the request identity helpers (header, query, session).
  2. Serve fakes from the MCP mount. tools/call goes to fakes first, with the error bodies for mismatch, exhausted, not declared, evicted and invalid override. tools/list includes faked tools. Journal entries record the body, test id, context and fake outcome. There is a failure metric. GET on the mount answers 405.
  3. Load mcpFakes from fixture files. New loadFixtureFileWithServices / loadFixturesFromDirWithServices (in src/fixture-loader-services.ts). The old free loaders throw on a file that holds mcpFakes instead of dropping it silently.
  4. Server and control API. Start-time hand-off to existing MCP mounts, auto-mount at /mcp, mount-conflict checks. POST /__aimock/fixtures accepts mcpFakes. New GET /__aimock/mcp/fakes. Reset and DELETE /__aimock/fixtures also unload fakes.
  5. LLMock, CLI, validator and test plugins. LLMock loads fakes before and after start. The CLI supports fakes-only starts, --watch and remote URLs, and reports every bad block. aimock validate checks blocks and cross-file id collisions. The vitest and jest plugins load fakes and fail beforeAll on a bad block. The pytest plugin's load_fixtures posts mcpFakes and detects a server that is too old. The MCP SDK v1 and v2 clients are pinned as dev dependencies for the real-client tests.
  6. Public exports. The new loaders, McpFakesAddError, MCP_FAKE_ERROR_CODES and the fake types, from the root and the ./mcp subpath.
  7. Docs. MCPMock page section on scenario fakes, plus the fixtures, strict mode, sequential responses, control API and Mastra pages, the homepage MCP card, the README and the Unreleased changelog entries.

User-visible changes

  • GET on an MCP mount answers 405 with Allow: POST, DELETE instead of falling through to the LLM routes. This includes a standalone MCPMock.
  • MCP tools/call journal entries carry the request body (was null), and response.mcpFake = { id, outcome } when a fake answered or failed.
  • MCP journal entries carry testId and context. GET /__aimock/journal?testId= lists an MCP request under the test id bound at initialize, or under the decoded header value, instead of "__default__" or the encoded string.
  • MCP requests read X-Test-Id and X-AIMock-Context. Answers on a mount without fakes do not change.
  • An MCP request that repeats a test id, context or undeclared-tools override on one path gets HTTP 400 MCP_DUPLICATE_IDENTITY. Nothing changes for clients that never do that.
  • New opt-in X-AIMock-MCP-Undeclared header or ?undeclared= query parameter (allow or deny) overrides the undeclared-tool policy for a request, or for the session when sent at initialize. With deny, any tools/call to a tool that no fake declares fails with MCP_FAKE_NOT_DECLARED, even on a mount with no fakes loaded, so onToolCall handlers do not run.
  • Any other value for that header or query parameter gets HTTP 400 MCP_INVALID_UNDECLARED on every MCP request, also on a mount without fakes.
  • POST /__aimock/fixtures: fixtures is optional when mcpFakes is sent, and the response adds mcpFakesAdded. Existing callers still get {"added": n}.
  • POST /__aimock/reset, DELETE /__aimock/fixtures, LLMock.reset() and LLMock.clearFixtures() also clear fakes.
  • Under undeclaredTools: "deny", an undeclared tool fails even when an onToolCall handler or a config result exists for it.
  • The CLI and aimock validate accept fakes-only files.
  • The recorder's snapshot merge keeps top-level keys it does not own (such as mcpFakes).
  • A tools/call whose name is not a string (for example 123) gets Invalid tool name: expected a string, got number instead of Unknown tool: 123. A missing or empty name still gets Missing tool name.
  • A JSON-RPC batch POST to an MCP mount is journaled as one entry per message, each with its own body, instead of one entry per HTTP request. Journal and getRequests() counts go up for batch callers.
  • The pytest plugin's load_fixtures still raises requests.HTTPError on a 400 with a JSON body, but the message is now aimock rejected fixtures from <path>: <error>, plus one line per details item, instead of the default requests text.
  • aimock validate --json always includes run.mcpFakeBlocks, which is 0 when no file has mcpFakes.
  • A bad mcpFakes block, a mount conflict, or a --watch reload whose mcpFakes changed fails the load instead of being skipped. On --watch the previous fixtures stay loaded. No effect without mcpFakes.
  • The old loaders now reject files with mcpFakes: the free functions loadFixtureFile and loadFixturesFromDir throw on a file with a top-level mcpFakes key. Use the ...WithServices loaders or the LLMock methods.
  • New public field unreadable on FixturesWithServices, the return type of the services loaders: the sources of files that could not be read or were not valid JSON.
  • An older aimock ignores mcpFakes silently. The MCPMock docs warn about this.
  • The MCPMock code API (addTool, onToolCall, addResource, addPrompt) does not change.

Proof

Red-green, same command on the base and on this branch:

pnpm build && AIMOCK_MCP_FAKES_13_1=1 pnpm vitest run src/__tests__/mcp-fakes-redgreen.test.ts

RED (base origin/main plus the test only):

 ❯ src/__tests__/mcp-fakes-redgreen.test.ts (1 test | 1 failed)
   × 13.1 red-green > serves the scripted retry sequence, then fails loud when exhausted
     → Streamable HTTP error: Error POSTing to endpoint: {"error":{"message":"Not found","type":"not_found"}}
 Test Files  1 failed (1)
      Tests  1 failed (1)

The base mounts nothing at /mcp and does not load the mcpFakes block.

GREEN (this branch):

 ✓ src/__tests__/mcp-fakes-redgreen.test.ts (1 test) 86ms
 Test Files  1 passed (1)
      Tests  1 passed (1)
exit=0

The env-var gate is removed on this branch, so plain pnpm test runs this test as well (287 files, 9349 passed at that point). Typecheck, lint, format check and build exit 0. Mutation checks on the CLI paths (watch guard, URL source label, fakes-only start, print every error) each turn one test red.

Value check with the built CLI (node dist/cli.js --fixtures <dir>), no real LLM calls:

  • Real agent. A LangGraph createReactAgent using @langchain/mcp-adapters MultiServerMCPClient (http transport) called the faked get_weather, got {"tempF":60,"conditions":"rain"}, and printed FINAL: It is 60F and raining in Seattle. (exit 0). The journal entry has response.mcpFake = {"id":"weather/seattle.json:get_weather#0","outcome":"answered"}.
  • Broken args. Changing the declared args to Portland gave MCP_FAKE_MISMATCH in error.data.aimock.code, with firstDifference: "$.city: expected \"Portland\", received \"Seattle\"". The journal recorded outcome: "mismatch" and the server log printed an MCP-FAKE: mismatch line.
  • Error bodies. With the MCP SDK v1 client, the retry (exhausted, -31010), any-args and closed-world (not declared, -32602) scenarios returned error bodies that match the documented shapes field for field.
  • Homepage card. Screenshots of the "Everything you need" grid and the MCP Protocol card at 375, 430 and 1440 px were checked. The card text fits and the grid is intact.

Note from the value check: LangGraph's default ToolNode (handleToolErrors: true) turns the MCP error into a tool message, so a default agent can still print the expected answer. The journal outcome and the server log catch that case.

Known limitations

  • Adding fakes across several existing mounts is not all-or-nothing. If a later mount rejects a block (for example, an entry-id collision), earlier mounts keep the blocks they already got. This applies to LLMock after start, the start-time hand-off to existing MCP mocks, and control-API adds that span several mounts. This is documented on the MCPMock page.
  • A journal test-id cap between 0 and 1 keeps 0 test ids. This behavior predates this PR.

Follow-ups (backlog)

  • Comment wording in a few places.
  • Test hygiene: some weak assertions.
  • Hardening for loggers, metrics or internal lookups that throw.
  • The /health body shape after auto-mount.
  • docs/integrate-maf shows the wrong command (@copilotkit/aimock --fixtures as npx arguments).

Release

Release surfaces are deferred to the release commit. There is no version bump in this PR. The changelog entries are under Unreleased.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AKuRWApa4KRBia1X1F3anz

jpr5 added 7 commits October 2, 2026 16:50
Adds the mcpFakes fixture types, the fake store (claim, match, echo, error results), message and
echo text helpers, FixtureLoadError rules for MCP fakes, request identity helpers, and the journal
cap getter. Adds the MCP SDK dev dependencies and the real-client test harness.
MCPMock and the MCP handler answer tools/call from the fake store, with structured MCP_FAKE_*
errors, per-test-id scoping and journal records. Adds the auto-mount used by the server and LLMock.
Adds loaders that return fixtures together with service fixtures and MCP fake sources. The recorder
keeps top-level keys it does not own, such as mcpFakes, when it merges into an existing file.
The server owns the mount list, auto-mounts /mcp for loaded fakes, and adds the /__aimock MCP fakes
control routes and reset doors.
…ugins

LLMock and the CLI load mcpFakes blocks and auto-mount them. validate-cli checks mcpFakes. The
vitest, jest and pytest plugins load fake blocks and report load errors.
The package entry and the ./mcp subpath export the new loaders, FixtureLoadError, McpFakesAddError,
the error codes and the public types.
Adds the MCP fakes guide, control API reference, fixture and integration notes, and a test that runs
the documented samples against a real server.
@pkg-pr-new

pkg-pr-new Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@copilotkit/aimock@490

commit: e218c56

jpr5 and others added 2 commits October 4, 2026 12:33
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AKuRWApa4KRBia1X1F3anz
Document the MCPMock fake methods, the loaders and their return shape,
the exports, entry and block ids, the identity 400 body, the failure
metric, the test plugins and the pytest plugin, aimock --config loading,
and aimock validate with mcpFakes. Add mcpFakeBlocks to the validate
--json samples, and list the missing items in the changelog.
@jpr5
jpr5 marked this pull request as ready for review October 5, 2026 15:20
@jpr5
jpr5 merged commit 2dc1d81 into main Oct 5, 2026
29 checks passed
@jpr5
jpr5 deleted the feat/mcp-fakes branch October 5, 2026 15:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant