feat(coreai): serve Qwen3.5-9B via Core AI runtime, thinking on - #176
Merged
Merged
Conversation
…ing on, honest caveats - flips coreai_driver's thinking default to ON (the finding: thinking off is the variable that fails the train-arithmetic step on both int8 and fp16) - registers qwen3-5-9b-coreai in the model registry: bundle facts, gate history, and the two runner caveats shipped in 2.9.0's coreai gate work (warmup must be off for the pipelined bundle; no working stop-token path) - adds coreai to WIRED_RUNTIMES so route_for resolves it - regenerates the web registry and the landing snapshot This is serve-the-bundle-with-caps, not 'the long gate passed': the spec's long gate (multi-hour soak on a longer prompt with thinking on) has not run yet, and 'qwen3-5-9b-coreai-int8' carries measured-run rows only for what was actually run. Default model stays qwen3-8.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Capped scope for the 'serve-or-kill' decision on Qwen3.5-9B: the driver thinks ON by default, the model is registered, coreai is wired into WIRED_RUNTIMES, and the honest caveats travel with the registry row (no working stop-token path; warmup off; the G2-train failure belongs to the checkpoint not the bundle).
What this does NOT claim:
qwen3-8-27b-4bit. Core AI becomes reachable, not the default.Local gates: registry + model revisions + coreai driver tests green; full suite 1071 passed.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.