You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Seven reported reviews support the gate’s rubric behavior and supersession, but native registration was missing and no two-session experiment ran. Net shared-learning benefit, factual reliability, and protocol adherence remain unproven.
Priority
P2 within the user-feedback follow-up. This issue records planned work, not a claim that a fix or experiment has shipped.
Acceptance criteria
Write the evaluation plan and success criteria before the run: relevant recall reuse, missed findings, contradictions against source evidence, factual errors, skipped recall/save/promote steps, token and latency cost, and human maintenance.
Use two genuinely fresh native sessions with distinct ordinary identities and related tasks, verifying that the second retrieves approved findings from the first without seeing the first agent’s drafts.
Compare against the same tasks using committed project documentation. Include runs with explicit workflow instructions and a clearly identified condition without extra reminders; do not conflate the two.
Include well-structured but deliberately false synthetic findings and adjudicate them against evidence. A gate approval alone must not count as factual correctness.
Use synthetic or explicitly authorized data and an agreed review destination. Start only after actual registration and installed-artifact checks pass. Invite external participation rather than assuming the original evaluator has committed to another test.
Record actual session dates, observations, costs, and limitations. If a week is selected, collect a real week; never invent elapsed time or equate this evaluation with approval to begin a separate formal pilot.
Source review baseline: main 6c5278a40a86246014901a88417f3455a46cdfcc, checked 2026-09-09. Evaluator observations are reported evidence, not a reproduction on their private project. Feedback is summarized without their project details, local paths, or original attachment.
Problem
Seven reported reviews support the gate’s rubric behavior and supersession, but native registration was missing and no two-session experiment ran. Net shared-learning benefit, factual reliability, and protocol adherence remain unproven.
Priority
P2 within the user-feedback follow-up. This issue records planned work, not a claim that a fix or experiment has shipped.
Acceptance criteria
Evidence and existing work
Source review baseline: main
6c5278a40a86246014901a88417f3455a46cdfcc, checked 2026-09-09. Evaluator observations are reported evidence, not a reproduction on their private project. Feedback is summarized without their project details, local paths, or original attachment.