Outcome
Turn external evaluation feedback into a smaller installation burden, accurate claims, and a measurable cross-session learning test. Preserve the positive evidence for rubric-based promotion and supersession without treating it as proof of factual truth or overall value.
First blockers
- JJ defaults to 0.28.0; current official stable is 0.45.1 as of 2026-09-09. Compatibility must be tested, and existing compatible installations must not be silently downgraded.
- Governed recall and MCP registration fixes are merged, but the latest public package remains 0.6.0. Prepare and verify the released artifact before telling evaluators to upgrade for these fixes.
- The evaluator exercised MCP directly, but native client registration and the two-session experiment did not complete. No elapsed trial time or net benefit is established.
Tasks
Existing delivery, not duplicate implementation
Order and dependencies
Start with JJ compatibility and release preparation. Documentation, safe onboarding, context measurement, local-only review validation, and local sync UX can progress independently. The two-session evaluation depends on a verified installed artifact, actual native registration, distinct identities, and an agreed data/review destination. Review candidate secret handling before inviting tests with sensitive data.
Decisions before an external evaluation
Agree who will participate, which synthetic or explicitly authorized data may be used, the local or cloud review destination, and the duration and success criteria. No new participation, provider consent, or formal pilot approval is implied by this issue. Historical data deletion or history rewriting is outside this follow-up.
Evidence limits
The reported test had seven LLM reviews: three approvals and four rejections. It supports the rubric and lineage behavior observed in that run. It does not prove prevention of initial secret storage or remote transmission, factual correctness of approvals, consistent agent adherence, or better outcomes than committed documentation.
Source baseline: main 6c5278a40a86246014901a88417f3455a46cdfcc, live GitHub/PyPI release metadata checked 2026-09-09, and summarized external feedback. The public issues exclude the evaluator’s private project details and original attachment. Creating this checklist does not claim its implementation is complete.
Outcome
Turn external evaluation feedback into a smaller installation burden, accurate claims, and a measurable cross-session learning test. Preserve the positive evidence for rubric-based promotion and supersession without treating it as proof of factual truth or overall value.
First blockers
Tasks
Existing delivery, not duplicate implementation
Order and dependencies
Start with JJ compatibility and release preparation. Documentation, safe onboarding, context measurement, local-only review validation, and local sync UX can progress independently. The two-session evaluation depends on a verified installed artifact, actual native registration, distinct identities, and an agreed data/review destination. Review candidate secret handling before inviting tests with sensitive data.
Decisions before an external evaluation
Agree who will participate, which synthetic or explicitly authorized data may be used, the local or cloud review destination, and the duration and success criteria. No new participation, provider consent, or formal pilot approval is implied by this issue. Historical data deletion or history rewriting is outside this follow-up.
Evidence limits
The reported test had seven LLM reviews: three approvals and four rejections. It supports the rubric and lineage behavior observed in that run. It does not prove prevention of initial secret storage or remote transmission, factual correctness of approvals, consistent agent adherence, or better outcomes than committed documentation.
Source baseline: main
6c5278a40a86246014901a88417f3455a46cdfcc, live GitHub/PyPI release metadata checked 2026-09-09, and summarized external feedback. The public issues exclude the evaluator’s private project details and original attachment. Creating this checklist does not claim its implementation is complete.