Repository navigation
feat: expose resolved renderer prefix stability - #184
Merged
Merged
Conversation
hallerite
marked this pull request as ready for review
October 9, 2026 08:36
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is an additive renderer metadata contract with configuration-aware boolean properties; existing rendering, bridging, and SFT behavior remain unchanged. Documentation and an opt-in audit corpus cover the new contract, with only the documented need for custom implementations to expose the property as a compatibility consideration. Notes:
No code changes detected at You can add or adjust custom eligibility rules. Learn more. |
hallerite
force-pushed
the
feat/prefix-stability
branch
from
October 9, 2026 09:24
62b235a to
6d8c744
Compare
hallerite
force-pushed
the
feat/prefix-stability
branch
from
October 9, 2026 09:35
6d8c744 to
0e4694a
Compare
hallerite
added a commit
to PrimeIntellect-ai/prime-rl
that referenced
this pull request
Oct 9, 2026
Warn when SFT constructs a renderer that does not guarantee prefix stability. The warning explains that each conversation is rendered once into one training sample, that N assistant turns are not expanded into N samples, and that earlier reasoning may be omitted by the template. Cached renderers warn only on construction; stable renderers stay quiet. Custom renderers without the capability are conservatively treated as unknown. Pin `deps/renderers` to merged main revision `41f9290`, which includes PrimeIntellect-ai/renderers#184. It provides the resolved capability and the manual, opt-in witness audit. Sample construction is unchanged. Validation: all **73 SFT dataset tests passed**, including stable/unstable renderer selection and cache-warning behavior. Prime-rl unit-test CI passed on the preceding revision; this update changes only the renderer gitlink to the merged revision. Dependency requirements are unchanged apart from renderer Python >=3.11, compatible with prime-rl Python 3.12. The manually invoked renderer audit passes **59 tests** and is excluded from normal CI discovery.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Expose
renderer.is_prefix_stableafter renderer configuration and chat-template kwargs are resolved. The flag describes whether full renders ending in an assistant remain token prefixes when messages are appended, with a fixed preamble and tools. Unknown custom templates are conservatively false.Configuration-sensitive renderers account for history reasoning retention, last-assistant wrappers, effort hints, and training mode. DeepSeek V4 and Inkling are conservatively false because supported message controls and reused tool-call IDs can rewrite earlier output.
A manual, opt-in audit uses
tests/fixtures/prefix_stability.json: 36 executable instability witnesses and 21 stable controls sharing six conversations. Each case stores the renderer/config, messages to append, and an explanation. Tests compare actual token prefixes and show the first differing tokens plus a decoded diff on failure. Coverage checks include every registered built-in renderer; opaque DefaultRenderer is tested as unknown. Finite stable controls are not a proof for arbitrary inputs.Run explicitly with
uv run pytest tests/audit_prefix_stability.py -q. The audit is excluded from normal pytest discovery and CI, and reuses the existing shared tokenizer cache.Validation: manual audit 59 passed in 9.80s; collection-only verification confirms the normal suite excludes the audit. Ruff lint/format and
git diff --checkpassed. The full suite was not rerun locally.Closes #41.
Rebased onto
origin/mainat376f54f(Python 3.11 typing update). The focused manual audit and Ruff checks pass after the rebase; all four PR commits are signed.Note
Add
is_prefix_stableproperty to the renderer protocol and all built-in renderersis_prefix_stableboolean property to the runtime-checkableRendererprotocol in base.py and implements it in every built-in renderer, classifying whether a full re-render preserves a completed conversation prefix.DeepSeekV3Rendereris stable;Qwen3Rendereris not). Others derive the value from configuration, such asGLM5Renderer(clear_thinking,enable_thinking),Qwen35Renderer(preserve_thinking),Hy3Renderer(raw_last_assistant,is_training), andNemotron3Renderer(truncate_history_thinking, effort hint).is_prefix_stablewill fail the runtime protocol check if it is used for isinstance-style validation.Macroscope summarized 0e4694a.
Note
Low Risk
Additive API and documentation; behavior of
render()and bridges is unchanged. Custom renderers without the property may needgetattrif consumers use strict protocol checks.Overview
Adds
renderer.is_prefix_stableto theRendererprotocol so callers (especially SFT) can tell whether a full re-render of a conversation that ends on an assistant keeps the same token prefix when more turns are appended (add_generation_prompt=False, fixed tools/config).DefaultRendereralways reportsFalse; each hand-coded renderer sets a constant or derives the flag from template knobs (e.g. GLM-5clear_thinking, Qwen3.6/3.8preserve_thinking, Hy3 training/preserved thinking, Nemotron effort/truncation).README documents the guarantee, how it differs from
thinking_retention/bridge_to_next_turn, and conservativegetattr(..., False)for custom renderers.Opt-in audit (
tests/audit_prefix_stability.py+tests/fixtures/prefix_stability.json): parametrized witnesses compare token prefixes before/after append and assert they match each renderer’s declared stability; includes registry coverage and a Qwen3.8chat_template_kwargscheck. Not run in default CI — manualuv run pytest tests/audit_prefix_stability.py -q.Reviewed by Cursor Bugbot for commit 0e4694a. Bugbot is set up for automated code reviews on this repo. Configure here.