Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Preserve the saved d5a129190000abbe3c6299050a4dfbc86bcc06b2 implementation recovered as 3bbd7fd on the newer cdb9dc5 base. Persist trusted read ages across queue removal, schedule outstanding work before completed rechecks, and bound consistency passes for discussion and timeline evidence. Recognize Markdown-local references without fetching arbitrary URLs. Inherited baseline: 62 passing tests. Recovery RED: 28 failures out of 107; human-decision hash mutation also fails both required tests. Final GREEN: 123 tests, covering default-budget drain, restart, moving pages, metadata races, and linked-only corrections. Classification corpus unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Preserve adjacent Markdown links while masking URL fragments, recognize punctuation and emphasis around issue references, and share label validation between snapshots and listings. Reject malformed listing state or labels before filtering and retain only successful page boundaries. RED: 26 new acceptance failures on inherited 6af22cd. GREEN: both prescribed Node suites pass 149 tests; saved baseline passes 62 tests. Human-decision hash mutation fails both human cases while bot cases pass. Node syntax, diff checks and Fantomas pass; frozen corpus and recovered/newer-base history are unchanged. The requested dotnet Release command cannot start because the pinned 11.0.100-rc.1.26420.103 SDK is absent; this sprint requires no F# product build. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Validate a single model proposal batch against trusted run-bound evidence, persist publication intents, and reconcile guarded label and clarification effects. Add versioned signed state commits, a no-write staged sink, and production adapter recovery tests. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Separate prior clarification attempts from current analysis publication, recheck the target after linked evidence reads, and validate durable clarification and intent fields. Cover crash recovery, malformed memory and known-unsent incomplete evidence reads. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Retain unresolved prior label attempts independently so their receipts cannot complete or block newer clarification and classification work. Cover crash transitions, same-input retries, late receipts and the final target-read handoff; consolidate duplicate authentication and policy tests. Validation: both Node suites pass 373 tests; syntax, formatter and diff checks pass. Release FSharp.slnx build passes. Full solution tests finish with 35 failures (19501 passed, 537 skipped); no product files are changed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Persist a recovery-time boundary instead of suppressing all future questions. Keep ambiguous histories pending, migrate legacy terminal noops, and reconcile recovered creation timestamps without replaying unresolved comment attempts. Add existing-ledger controls and missing branch/file, boundary, migration and receipt recovery coverage. Both prescribed Node suites pass all 396 tests; syntax, formatting and diff checks pass. The clean Release solution build passes with zero warnings or errors. Full solution tests have reported failures; results are retained in the session logs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Accept the pinned framework's empty ingestion-error envelope and preserve JSON proposal citations through its typed validation config. Keep failed ingestion fail-closed. Include qualified GitHub dependencies in freshness checks and reserve analysis capacity by pending age independently of snapshot reads. Cover the actual pinned MCP/ingestion path, continuous backlog churn, linked corrections, staged restarts and exact failed-request attempt counts. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Check model content bounds before fair candidate selection and refill batch-capacity rejections from already-read snapshots. Keep incomplete work pending without increasing read budgets. Cover mixed stable and changing backlogs, staged restarts, exact UTF-8 entry bounds and batch refill. All 475 deterministic tests and 22 actual-model fixtures pass; pinned workflow compilation remains unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Cap proposal batches at 64000 UTF-8 bytes before publication, matching the pinned HTTP string-offload threshold. Exercise official tool generation, HTTP requests, ingestion and staged publication at ASCII, Unicode and escaped-content boundaries. Validated 495 deterministic tests with no skips, 22 fresh gpt-5.6-sol fixtures and staged checks, and reproducible GH AW v0.76.1 compilation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add timestamp-only restart checks, interleaved clarification claim races, and cross-repository identity and citation coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rotate the reserved classification slot by persisted selection age so unresolved historical work cannot starve later reports. Use canonical API identities for redirected evidence while preserving root publication scope and freshness guards. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Preserve incomplete pending history while allowing in-repository reports and discovery progress to publish across restarts. Keep linked redirects and the independent publisher scope guard intact. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Carry clarification history across issue transfers, require human provenance for durable corrections, and fail active workflow runs without successful publication. Cover isolated reference origins and restart behavior. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep the forwarding characterization independent of FSharp.Core inlining policy by using an explicitly non-inline recursive target. Preserve both existing allocation assertions and test bodies; List.forall2 is now inline in the locally built Core. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Merge the recovered feature and Fixup #3 history into the checkout used by verification, preserving current main changes. The previous isolated-only delivery left required implementation and tests absent from this branch. Includes transfer-safe clarification history, human-only durable corrections, the trusted missing-output completion guard, linked-reference regression coverage, and the verified closure fixture repair. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
T-Gro
left a comment
There was a problem hiding this comment.
🤖 🕵️ AI review — verify independently.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Retain the non-inlining guard together with main's allocation-free prelude assertion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
| if (proposal.correction) { | ||
| const source = sources(snapshot).get(proposal.correction.sourceId); | ||
| record.humanCorrection = { ...proposal.correction, createdAt: source.createdAt ?? null }; |
There was a problem hiding this comment.
[P2] 🤖 An older citation can erase a newer persisted human veto. I reproduced this with a Sep 1 rejection, Sep 2 human Regression application, Sep 3 rejection saved to memory, and Sep 4 bot label removal. After deleting the Sep 3 comment, a proposal citing Sep 1 replaces the saved correction; a later positive proposal adds Regression because Sep 2 now appears to override the rejection. Preserve the newest durable correction when accepting older evidence, including when its original comment was deleted.
| const deadline = /^\d+$/.test(retryAfter) ? Date.parse(now) + Number(retryAfter) * 1000 : Date.parse(retryAfter); | ||
| const reset = Number(headers["x-ratelimit-reset"]) * 1000; | ||
| const fallback = Date.parse(now) + 60000 * 2 ** (intent.rejections - 1); |
There was a problem hiding this comment.
[P2] 🤖 Relative retry delays are measured from batch-start now, so they can expire before the rejection arrives. With a batch starting at 18:00 and a 429 Retry-After: 60 received at 18:05, this saves retryAt = 18:01; the next run sends another request at 18:05:01 instead of waiting until 18:06. Calculate numeric Retry-After and fallback backoff from a trusted clock at rejection time, while retaining absolute server deadlines.
| pendingLabelPublication: priorLabel ? prior.pendingPublication : prior.pendingLabelPublication ?? null, | ||
| pendingPublication: unresolved && !priorComment && !priorLabel ? prior.pendingPublication : { operationId, phase: "prepared" }, |
There was a problem hiding this comment.
[P2] 🤖 Returning to the same positive operation loses its retained retry state. After a 429 with Retry-After: 3600, an intervening uncertain proposal moves the label intent into pendingLabelPublication. If that uncertainty stays pending because clarification history is unknown, the next positive proposal for the same fingerprint creates a fresh intent here. I reproduced four label requests before the deadline, all with the same operation ID and rejections reset to 1. Restore the matching retained label intent so its deadline and retry budget survive reclassification.
| add(item.titleSourceId, { url, body: item.title, createdAt: item.updatedAt, human }); | ||
| add(item.bodySourceId, { url, body: item.body, createdAt: item.updatedAt, human }); |
There was a problem hiding this comment.
[P2] 🤖 updatedAt is the issue's update time, not the correction's edit time. Save a rejecting body on Sep 1, apply Regression manually on Sep 2, then remove it through a bot on Sep 3 without editing the body. Re-citing that unchanged body dates the correction Sep 3, and a subsequent positive proposal is incorrectly suppressed as human-veto. Keep the saved chronology for unchanged correction text; unrelated issue activity must not turn an old rejection into a new human decision.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: efda83b5-9bca-463f-9f91-af5b6883ee3e
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: efda83b5-9bca-463f-9f91-af5b6883ee3e
Adds one compact GitHub Agentic Workflow for regression triage.
It runs only for an open, non-PR issue when it is opened/reopened with
Needs-Triage, or whenNeeds-Triageis applied. It reads the report, human discussion, and directly linked GitHub evidence, then addsRegressiononly when the evidence supports previously working behavior becoming broken.The agent has no shell, cannot comment, remove labels, close issues, or modify code, and can add at most one
Regressionlabel to an issue that still hasNeeds-Triage.