Reuse recorded tokens by default and preserve literal assistant content - #977
Draft
bradhilton wants to merge 64 commits into
Draft
bradhilton wants to merge 64 commits into
bradhilton wants to merge 64 commits into
Conversation
bradhilton
marked this pull request as ready for review
September 25, 2026 20:17
This was referenced Sep 26, 2026
bradhilton
marked this pull request as draft
September 26, 2026 15:46
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Complete recorded prompts and sampled outputs drive tokenization by default when their provenance and boundaries are valid. This preserves the policy's actual tokens and log probabilities, including length-stopped responses, without an opt-in or reconstructing them through the text tokenizer. Missing or edited evidence continues through rendering and validation; explicit template overrides remain respected. Standard thinking-disabled assistant content preserves literal
<think>/</think>text.Evidence is checked across renderer/provider callbacks and mutable inputs, including physical mapping, Pydantic, scalar-subclass and enum state. Proof refusal is distinguished from observer errors. Sticky errors preserve identity through nested tokenization and release internal traceback ownership at the outer operation boundary. Original Anthropic/Responses evidence is now checked between template selection and rendering; disposable converted messages retain their existing semantics. Historical-role fallback refuses unavailable original request projections consistently, and the docs explain safe copied-owner and conditioning refusals. This integrates with main's staged tokenizer and contains no process-result codec changes.
Validation includes 1,941 tests on the lifetime composition, 69 subsequent actual-model admission controls, 356 parser controls, and 224 final callback/error/lifetime controls. Three public selector/renderer corruption cases fail on the parent and pass with the correction; independent receiving also checks exception identity and immediate failed-input release. The historical-request correction has three parent
UnboundLocalErrorwitnesses and three fixed positives, with 35 existing role controls passing. Typing and lint checks pass. Fresh Astra and Fable whole-source inspections completed on the preceding composition; their concrete findings have been addressed in received children. The final parser restricts role-dependent whitespace rewriting to proved local renderer bindings under the standard chat-rendering environment; unknown/custom effects retain their original trim. The exact composed child passed 82 literal/role/callback controls and independent source receiving. The exact final-child Astra review identified a reproduced counter-free macro effect gap in the parser; its correction is pending while the other finite review and CI continue. Astra also disclosed unavailable local-read tooling, so missing caller-review coverage is not credited. Performance remains held. No experiment has adopted this branch.Performance remains unresolved. The last finite pre-integration authentic comparison measured 1.626 s versus its predecessor's 1.663 s, while the older frozen source took 0.886 s; validation schedules differ, so this is not a general speedup claim. The lifetime/type-admission composition preserved 29,605 tokens, all masks, aliases, input nonmutation and 1,313 callback events. Actual-template equivalence carries that exact-case proof through the prior parser-only children; the newest parser donor also preserves both actual Qwen templates. This does not cover the new exchange callback delta or establish performance parity. Earlier hosted CI failed the rerender scaling test at 3.086 times the 64-turn runtime for 128 turns (limit 3), with 3,209 other tests passing. A later passing CI result does not resolve the deterministic superlinear observation cost. No threshold was changed. The bounded cache limits snapshot bytes; retained exchange references can keep their graphs alive through the builder lifetime.