Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
68 commits
Select commit Hold shift + click to select a range
dda1108
Preserve literal assistant content in Qwen thinking templates
bradhilton Sep 25, 2026
ddab236
Constrain literal-thinking rewrite to its supported template syntax
bradhilton Sep 25, 2026
52cb457
Bind thinking-template edits to executable Jinja blocks
bradhilton Sep 25, 2026
225bb8a
Use recorded boundaries and preserve sampled conditioning by default
bradhilton Sep 25, 2026
09018c9
Test copied-context loss and routing after length filtering
bradhilton Sep 25, 2026
ab0b3ea
Preserve copied-context authority across native protocols
bradhilton Sep 25, 2026
b11ac6b
Reuse resolved sampled stop authority and mask alignments
bradhilton Sep 26, 2026
2e4ec3c
Preserve literal content in selected and equivalent chat templates
bradhilton Sep 26, 2026
3ed85fa
Retain whitespace coverage and safe named-template selection
bradhilton Sep 26, 2026
6c595a9
Preserve recorded request role boundaries and reuse scoped evidence
bradhilton Sep 26, 2026
9b1507b
Reuse native source evidence across the tokenization stop decision
bradhilton Sep 26, 2026
f62857b
Isolate the warning state in the callback fallback regression
bradhilton Sep 26, 2026
b310d69
Use final recorded output as the end of complete native Chat histories
bradhilton Sep 26, 2026
c5a30b7
Keep resolved STOP tests on histories that require role rendering
bradhilton Sep 26, 2026
afabb3c
Test literal rendering separately from complete native output
bradhilton Sep 26, 2026
d252b53
Preserve request roles and validate token evidence across callbacks
bradhilton Sep 26, 2026
360f126
Preserve complete tokenization context across renderer callbacks
bradhilton Sep 26, 2026
bd64400
Validate all consumed sources after historical role proofs
bradhilton Sep 26, 2026
5f39e1e
Validate consumed evidence through rendered token replacement
bradhilton Sep 26, 2026
4ece40f
Keep scoped reasoning normalization idempotent
bradhilton Sep 26, 2026
201fdf4
Retain consumed evidence and prove complete historical request context
bradhilton Sep 26, 2026
1466a4e
Check consumed evidence around prefix repair warning hooks
bradhilton Sep 26, 2026
f5bd90b
Retain distinct consumed sources and allow unused opaque native context
bradhilton Sep 26, 2026
1033b83
Preserve content bindings across conditional parser joins
bradhilton Sep 26, 2026
c27a555
Bind consumed protocol evidence before tokenizer callbacks
bradhilton Sep 26, 2026
e11885a
Guard rendered token evidence and callback inputs before reuse
bradhilton Sep 26, 2026
da6c866
Keep loader mutation tests aligned with consumed evidence lifetime
bradhilton Sep 26, 2026
6864c96
Prove request roles without reconstructing later native bodies
bradhilton Sep 26, 2026
6d14031
Fence dynamic STOP metadata and avoid unused template normalization
bradhilton Sep 26, 2026
26ffc66
Classify STOP callbacks through the internal rendering wrapper
bradhilton Sep 26, 2026
8a7b725
Preserve reusable projections and first sufficient native role proofs
bradhilton Sep 26, 2026
eafc6a9
Retain callback authority through completed history validation
bradhilton Sep 26, 2026
8dff469
Require decoder only for unresolved native boundaries
bradhilton Sep 27, 2026
55823f7
Include scalar and Enum slot storage in context snapshots
bradhilton Sep 27, 2026
02ca697
Clarify rendered role extent and restored-edit authority
bradhilton Sep 27, 2026
eb9e483
Preserve recorded fields and request rendering authority
bradhilton Sep 27, 2026
0fc97ed
Reuse unchanged evidence serialization within source validation
bradhilton Sep 27, 2026
5f818a3
Reuse revalidated fingerprints across tokenization phases
bradhilton Sep 28, 2026
aa39739
Avoid duplicate identity work in tokenization context snapshots
bradhilton Sep 28, 2026
ea9e747
Inline exact logprob snapshots while preserving observation order
bradhilton Sep 28, 2026
8468814
Preserve exact dict dispatch while narrowing its static type
bradhilton Sep 28, 2026
24ea4bd
Merge ART main and retain recorded-token guarantees in staged tokenizer
bradhilton Sep 28, 2026
bec23b1
Narrow tokenizer regression fixture types for static checking
bradhilton Sep 28, 2026
da8106f
Validate physical mapping state across tokenizer callbacks
bradhilton Sep 28, 2026
6878aa3
Treat recursive tokenization contexts as unsupported callback proofs
bradhilton Sep 28, 2026
2034198
Require unsupported-context error while retaining recursion cause
bradhilton Sep 28, 2026
dd5d745
Observe complete physical Pydantic context state across callbacks
bradhilton Sep 28, 2026
7951618
Release tokenization snapshot closures after each observation
bradhilton Sep 28, 2026
e32455a
Detect recursive context explicitly without hiding observer errors
bradhilton Sep 28, 2026
9b51a19
Observe native instance storage consistently across context types
bradhilton Sep 28, 2026
90ec720
Preserve context observer failures and native dictionary evidence
bradhilton Sep 28, 2026
112d30b
Use identity for exact type admission and context class tags
bradhilton Sep 28, 2026
f9203d2
Use identity for plain render-cache type eligibility
bradhilton Sep 28, 2026
52ae6fe
Assert retained identity in custom context type tags
bradhilton Sep 28, 2026
d09152f
Release sticky observer failures when tokenization ends
bradhilton Sep 28, 2026
74f0554
Preserve shared trim bindings across calls and parser fallback
bradhilton Sep 28, 2026
886f233
Narrow the rendered-template type in the macro regression
bradhilton Sep 28, 2026
759613c
Require actual model types for context source authority
bradhilton Sep 28, 2026
8bfb7d6
Conservatively preserve shared content across unknown template effects
bradhilton Sep 28, 2026
6d3682f
Limit transparent assignments to fresh loop-local bindings
bradhilton Sep 28, 2026
c5f0d6a
Track prior stores independently of template role pruning
bradhilton Sep 28, 2026
3e8962f
Validate original exchange evidence between template callbacks
bradhilton Sep 28, 2026
cca25f4
Refuse unavailable historical request role proofs consistently
bradhilton Sep 28, 2026
3d37269
Keep shared trim for unproved renderer bindings and context effects
bradhilton Sep 28, 2026
4aed0d4
Expect the earlier consumed-projection refusal in response controls
bradhilton Sep 28, 2026
f9d834e
Check whole-template effects for counter-free renderers
bradhilton Sep 28, 2026
90e36bf
Preserve trim around renderer parameter mutations
bradhilton Sep 28, 2026
7407b1b
Keep private counters bound to validated declarations
bradhilton Sep 28, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
125 changes: 123 additions & 2 deletions docs/features/additional-histories.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,24 @@ with `chat_template_kwargs={"preserve_thinking": False}`. Additional histories
remain useful for custom or externally managed templates that do not expose a
prior-thinking preservation option.

For supported Qwen templates, reasoning belongs in the structured
`reasoning_content` field. Assistant `content` remains literal, including
`<think>` and `</think>` anywhere in that content; ART does not infer reasoning
from those strings. This also applies when the next response has thinking
enabled or when prior reasoning is explicitly omitted. A legacy adapter that
knows its response uses a leading reasoning envelope should split that known
format into `reasoning_content` and `content` before rendering. A leading tag
pair alone cannot establish that format. Structured reasoning and the template's
generation-prompt defaults retain their existing behavior.

For named template dictionaries, ART uses the tokenizer’s default, tool, or
explicitly selected template before applying the same correction. The correction
recognizes the known inline-content parsing operations, including equivalent
quoting and spacing; it does not reinterpret arbitrary custom template logic.
This also preserves literal content when no recorded token IDs are available.
If a corrected template body is itself another dictionary entry’s name, ART
refuses that ambiguous selection rather than rendering a different template.

By splitting each turn into a separate history, you can preserve these tokens for training:

```python
Expand All @@ -44,7 +62,7 @@ trajectory = Trajectory(
messages_and_choices=[
# First turn with thinking
{"role": "user", "content": "What is 2+2?"},
{"role": "assistant", "content": "<think>I need to add 2 and 2</think>4"}
{"role": "assistant", "reasoning_content": "I need to add 2 and 2", "content": "4"}
],
additional_histories=[
LegacyHistory(
Expand All @@ -53,7 +71,7 @@ trajectory = Trajectory(
{"role": "user", "content": "What is 2+2?"},
{"role": "assistant", "content": "4"},
{"role": "user", "content": "What is 3+3?"},
{"role": "assistant", "content": "<think>I need to add 3 and 3</think>6"}
{"role": "assistant", "reasoning_content": "I need to add 3 and 3", "content": "6"}
]
)
]
Expand Down Expand Up @@ -134,6 +152,109 @@ trajectory = Trajectory(
)
```

## Recorded exchange histories

`art.tokenize` and `trajectory.tokenize` use recorded prompt and response token IDs
for unchanged, complete exchange histories. Recorded logprobs belong to those
exact conditioned tokens. Chat, Responses, Messages, and Completions keep their
existing protocol-specific projection rules; no separate tokenization API or
native-representation option is needed. `multi_history=True` preserves the
histories selected by the trajectory, including their order and model selection.

Complete unchanged Chat histories end at the final recorded response token,
including tool responses and responses stopped by a length limit. ART does not
reconstruct that sampled body from its text or structured tool projection, or add
an unobserved terminal footer. This deliberately excludes synthetic terminal
tokens from `OUTPUT`/SFT masks; recorded sampled tokens and logprobs are unchanged.
Explicit template overrides and incomplete or edited histories retain rendering.

Templates still prove nonterminal separators, role masks, and synthetic stop
tokens against the next recorded prompt.
If tokenizing another history in the same trajectory already resolves a tokenizer
for the same model, ART reuses that authority to label recorded sampled stop
tokens. This does not trigger a new tokenizer load or change rendering. Complete
native histories still work offline without a tokenizer; when neither recorded
stop metadata nor resolved tokenizer authority identifies a stop, ART leaves that
label unknown rather than guessing from the final token.

For supported Chat boundaries, ART decodes the recorded body and encodes only the
unrecorded separator instead of re-tokenizing the whole conversation. It checks
that the separator reproduces the next recorded prompt exactly. Edited contexts,
explicit template overrides, incomplete projections, and unsupported templates
continue through the generic rendering path and its source validation. That
fallback can still refuse copied response context when it cannot prove the
original conditioning of the sampled tokens that follow it.

Correcting literal-content rendering does not rewrite a recorded request. When
the original request and template reproduce its complete native prompt, ART can
recover historical assistant roles from that rendering. This uses the original
tool serialization order and preserves role labels through exact length-stop
assembly; it does not restore destructive parsing for new response content.
The role proof uses a complete recorded request covering the request-only role
spans being attributed, including its trailing query. The normalized path uses
the first sufficient request; the original-template fallback also locates the
last request-only assistant message. Later sampled native bodies remain
authoritative even when their rendered projections differ. An empty assistant
with no rendered token extent supplies no role or STOP authority for unmatched
native tokens; ART leaves those tokens' roles unknown rather than inventing them.
With a supplied renderer, request-owned assistant roles are proved throughout
the native stream, including between sampled responses. This requirement applies
even when template normalization makes no change. A request whose text
disagrees with its recorded assistant tokens can still use the offline native
path, but cannot claim a complete rendered role mask. Recovering those roles
requires the original request to reproduce the complete recorded prompt, with
only independently proved whitespace retokenization permitted. Matching an
assistant body alone does not prove an altered role header. If that proof fails,
ART refuses that rendered role certification; this check does not require
reconstructing text on the complete offline native path.

Custom STOP encoders, terminator decoders and metadata lookups must not change
already-consumed source tokens, logprobs, model or stop evidence. ART checks each source around its
STOP callback and checks all consumed evidence before returning, including
logprobs used only by the rendered fallback. Multi-history
tokenization also checks completed histories after later renderer callbacks.
That callback authority is retained even if a callback later removes the
tokenizer methods that made it callback-capable.
Evidence is bound when first read, including a later Responses generation read
to prove a copied suffix, rendered item IDs and logprobs, and text used for prefix
repair. STOP admission also retains the source keys read before its encoder.
It is not a snapshot of every future response before that response is used.
Ordered original requests and history context, including roles, tools, render
settings and protocol selectors, are captured before callbacks and checked
before reuse. Chat role rendering checks its working projection and actual
callback arguments before and immediately after callbacks, including optional
probes; a later encoder cannot hide an earlier edit by restoring it. Unexpected
callback exceptions keep their original identity.

When recorded prompt IDs are unavailable, a supplied renderer defines the prompt
tokens. ART cannot certify that arbitrary returned IDs implement a particular
text transformation. A renderer changing a disposable, single-use message copy
does not itself invalidate that authority. ART refuses inconsistent borrowed
provenance and changed semantic inputs that it reuses. An edit to an original
object that is restored before any later consumption, leaves final provenance
unchanged, and changes no tokenization decision is not itself a stale result.
A Responses message projection reused for prompt and completion rendering must
remain unchanged across those calls.
These checks refuse stale results; they do not make callbacks or source objects
immutable. Context snapshots reuse shared containers only within one observation;
callback-free native assembly retains its bounded response-evidence reuse.
Unused opaque request metadata remains valid on complete callback-free native
paths, including a supplied tokenizer exposing only plain EOS metadata;
callbacks require context that ART can compare.

A response copied into a later, shortened prompt is output provenance, but it is
not a fresh sample under that new prompt. ART keeps its `OUTPUT`, `ASSISTANT`,
`EXACT`, and proven `STOP` flags while removing `SAMPLED` and the old conditional
logprob. This requires the complete original sampled occurrence to remain in an
earlier selected history; otherwise unchanged native replay is refused. The
original occurrence retains its logprobs and ownership, including recorded NaNs
before finite-value filtering. Tokenizing only the shortened view cannot prove
that coverage; tokenize the containing trajectory with `multi_history=True`
and keep `reconcile_text_equivalent_tokenizations=False` (the default) so the
complete original occurrence remains in a selected history. Reconciliation can
collapse that owner into a shortened view, which ART safely refuses. This correction does
not change how generic output/SFT masks include copied assistant content.

## How It Works

### Tokenization Process
Expand Down
14 changes: 11 additions & 3 deletions src/art/trajectories/_render_cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
def _render_context_key(value: object) -> object:
"""Snapshot plain JSON without losing mapping order or scalar types."""
kind = type(value)
if kind in (str, int, bool, type(None)):
if kind is str or kind is int or kind is bool or kind is type(None):
return kind, value
if kind is float and math.isfinite(cast(float, value)):
return kind, repr(value)
Expand Down Expand Up @@ -51,7 +51,10 @@ def cacheable_chat_template(tokenizer, template, tools, kwargs, messages) -> boo
or base_module.render_jinja_template is not chat.render_jinja_template
or not cls.__module__.startswith("transformers.")
or getattr(module, cls.__name__, None) is not cls
or type(tokenizer.chat_template) not in (str, type(None))
or (
(template_type := type(tokenizer.chat_template)) is not str
and template_type is not type(None)
)
or inspect.getattr_static(cls, "special_tokens_map")
is not inspect.getattr_static(base, "special_tokens_map")
):
Expand All @@ -67,7 +70,12 @@ def cacheable_chat_template(tokenizer, template, tools, kwargs, messages) -> boo
added_token = sys.modules["tokenizers"].AddedToken
special = tokenizer._special_tokens_map
if type(special) is not dict or any(
type(key) is not str or type(value) not in (str, type(None), added_token)
type(key) is not str
or (
type(value) is not str
and value is not None
and type(value) is not added_token
)
for key, value in special.items()
):
return False
Expand Down
Loading
Loading