fix(stage-router): warn when capable-first cannot offload via scorer - #711
ting-hong-shieh wants to merge 2 commits into
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughChangesStage threshold behavior
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Severity of issue fixed: Medium Merge Risk: ⚪ Minimal · up to The warning and documentation changes are ready to merge; no concrete regressions remain identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 1 files. (1 skipped: 1 unsupported.)
A rabbit reads each line, Comment |
|
++ @sabhatinas for review |
|
Small wording nit: direct |
|
Good catch, thanks. I narrowed the description: the warning shows on the runner and server paths, including the Python server. Direct |
Signed-off-by: Ting-Hong Shieh <32212900+ting-hong-shieh@users.noreply.github.com>
Signed-off-by: Ting-Hong Shieh <32212900+ting-hong-shieh@users.noreply.github.com>
dbde17d to
329dd60
Compare
What
Warn when a
capable_firststage router uses a confidence threshold that prevents the scorer from selecting the efficient tier. For example,0.5exceeds the efficient confidence ceiling of approximately0.462117.The warning runs once when
StageClassifieris constructed, not on each request. It appears in the runner and server, including the Python server path, because they set up tracing output. A directswitchyard.libsy.stage_router()call runs the same check, but that path sets up no tracing output, so the warning is not shown. It computes the ceiling through the existing scorer and uses the picker's closed probability band, including the exact boundary. Routing behavior is unchanged: an optional LLM classifier can still select efficient.The documentation adds a short section on this limit for
capable_first.Why
Closes #264.
The current TOML schema requires an explicit threshold, but
capable_firstwith0.5still silently disables scorer-driven offloading. This implements the issue's warning option without changing scoring weights or introducing new public APIs.Notes for reviewers
Start with
StageClassifier::newincrates/libsy/src/algorithms/util/stage.rs. One regression test checks warning output and picker behavior at0.45, the exact ceiling, and0.5.This branch is rebased onto
main.mainremoved hard de-escalation in #651, so the warning, test, and docs no longer mention it.Validation after the rebase:
cargo test -p switchyard-libsy: 322 passed.cargo clippy --workspace --all-targets -- -D warningsandcargo fmt --all --check: passed.Validation before the rebase:
cargo test --workspace: 757 passed, 1 ignored.uv run ruff check .: passed;uv run pytest tests/: 116 passed, 2 skipped.Validation used the project's Python 3.14 interpreter for PyO3 and allowed local ports for mock HTTP servers. No live model benchmark was run;
0.45is documented as below the ceiling, not as a calibrated recommendation.Summary by CodeRabbit
New Features
capable_firstmode when production scoring cannot reach the configured confidence threshold for efficient selection.Documentation
0.5threshold applies toefficient_first.