refactor(steerwatch): verify follow-up run provenance and split the delta by authority - #7447
waynesun09 wants to merge 3 commits into
Conversation
PR Summary by QodoVerify steer provenance and split deltas by authority
AI Description
Diagram
High-Level Assessment
Files changed (20)
|
|
🤖 Review · Commit: |
Code Review by Qodo
1.
|
c47571c to
cdcc59e
Compare
|
🤖 Finished Review · ✅ Success · Started 2:42 PM UTC · Completed 3:10 PM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $12.20 |
|
Risk Assessment: moderate (2/5) DetailsRe-review anchoring holds: Tier 1 signals are unchanged in kind and magnitude from the prior round (25 files, ~5974 lines, large blast radius, zero protected/security/CI/dependency touches, 0.28 test ratio, known non-bot non-first-time author) giving Tier 1 ~1.9; Tier 2 git-history churn on the pre-existing shared files (forge.go, github.go, fake.go, run.go, gitlab/issue.go, gitlab/mr.go, ADR0106, architecture.md, security-threat-model.md) lands at ~2.9-3.0 once diluted by very fresh code age, near-zero true reverts, and zero hack/workaround sentiment despite superficially high raw commit/fix-grep counts driven by hub files touched by nearly every PR in this active monorepo; with no same-repo linked issue (62/38 metadata/git-history weighting), composite = 0.621.9 + 0.383.0 = 2.25 -> rounds to 2, preserving the score that has been stable across 19+ prior rounds. Previous runRisk Assessment: moderate (2/5) DetailsNo same-repo linked issue (62/38 metadata/git-history weighting); Tier 1 stays low (~1.9) since the large raw diff is dominated by the still-unwired, net-new internal/steerwatch package plus other genuinely new files, with zero protected/security/CI/dependency touches and an adequate 0.28 test ratio, by a known non-first-time, non-bot author; Tier 2 (~3.0) reflects real but not extreme churn/author-contention on the handful of pre-existing shared files touched (forge.go, github.go, fake.go, run.go), tempered by very fresh code age, low true revert counts, and no hack/workaround sentiment. Composite rounds to 2 (moderate), consistent with re-review anchoring against the prior stable score of 2 across 17+ rounds. Previous run (2)Risk Assessment: moderate (2/5) DetailsNo linked issue (62/38 metadata/git-history weighting); Tier 1 stays low (~1.9) because the large raw diff is dominated by the still-unwired, net-new internal/steerwatch package plus several other genuinely new files, with zero protected/security-manifest/CI/dependency touches and an adequate 0.28 test ratio; Tier 2 (~2.75) reflects elevated but not extreme churn/author-contention on the small set of pre-existing shared files this PR touches, tempered by low revert counts and minimal hack/workaround sentiment. Composite rounds to 2 (moderate), reproducing and confirming the stable score assigned across 17+ prior review rounds on this PR via re-review anchoring. Previous run (3)Risk Assessment: moderate (2/5) DetailsNo linked issue in this repository (62/38 metadata/git-history weighting applied); metadata signals stay low since the large raw diff is dominated by the still-unwired, net-new internal/steerwatch package with no protected/security-manifest/CI/dependency touches and an adequate 0.28 test ratio, while git-history churn on the small set of pre-existing shared files this PR touches (run.go, forge.go, github.go, architecture.md, gitlab issue.go/mr.go) is elevated but offset by zero actual reverts and zero hack/workaround sentiment; composite lands at 2 (moderate), consistent via re-review anchoring with the same score assigned across all prior review rounds on this PR. Previous run (4)Risk Assessment: moderate (2/5) DetailsComposite ~2.3/5 (moderate, weights 62% Tier1 / 38% Tier2, no linked issue): Tier1 metadata is materially unchanged from the prior review (25 files, ~5.9k lines, no protected/security-sensitive/CI/dependency changes, 0.28 test ratio, established non-bot author) giving a low Tier1 average despite the large raw diff size being dominated by the still-unwired, net-new internal/steerwatch package; Tier2 is pulled up by high recent churn/fix-revert rates and author contention on the small set of pre-existing shared files this PR touches (run.go, forge.go, github.go, architecture.md) but pulled back down by very recent single-purpose commits with no revert or hack/todo sentiment. This lands the composite in the same moderate band as the prior assessment (score 2), preserved per re-review anchoring. Previous run (5)Risk Assessment: moderate (2/5) DetailsComposite ~2.2/5 (moderate): the diff is large (5.9k lines, 25 files) but almost entirely a net-new, not-yet-wired internal/steerwatch package plus additive forge.Client methods -- no protected paths, no security-sensitive path matches, no CI/dependency changes, and an adequate test ratio (0.28), keeping Tier 1 low; Tier 2 is pulled up modestly by recent churn and fix/revert rates on the small set of pre-existing shared files this PR touches (forge.go, github.go, cli/run.go) but pulled back down by very recent, single-purpose commit history with no revert or hack/todo sentiment, landing the composite in the moderate band despite the PR's raw size and its self-applied 'risk/elevated' label. Previous run (6)Risk Assessment: elevated (3/5) DetailsWeighted Tier1/Tier2 average (no reliably scoreable cross-repo Tier3) lands near 2-3, but the change touches security-sensitive provenance/authority-verification and untrusted-content-defanging code central to this repo's external-injection-first threat model and has undergone 10 prior review rounds, which rounds the score up to 3 (elevated, requires careful review). Previous run (7)Risk Assessment: moderate (2/5) DetailsTier 1 signals are unchanged from all 12 prior assessments on this PR (25 files, 5814 lines, large blast radius, no protected/security/CI/dependency touches, 0.28 test ratio, established non-bot/non-first-time author); no same-repo linked issue (fullsend-ai/agents#1163 is in a companion repo), so 62%/38% Tier1/Tier2 redistribution applies; Tier 2 remains elevated due to churn on shared hotfiles but the bulk of new lines sit in the net-new, uncoupled steerwatch package, so the weighted composite still rounds to 2, preserving the stable moderate score via re-review anchoring. Previous run (8)Risk Assessment: moderate (2/5) DetailsTier 1 signals are unchanged from 11 prior assessments on this PR (25 files, ~5.6-5.8K lines, large blast radius, no protected/security/CI/dependency touches, 0.28 test ratio, established non-bot/non-first-time author), and although Tier 2 churn/author/fix-revert activity on shared hot files (run.go, architecture.md, forge.go) is elevated, the Tier1 62%/Tier2 38% weighted composite (no linked issue) still rounds to 2, preserving the stable moderate score via re-review anchoring. Previous run (9)Risk Assessment: moderate (2/5) DetailsTier 1 and Tier 2 signals are unchanged from the prior 11 assessments on this PR (same 25 files / ~5619 lines / large blast radius, no protected/security/CI/dependency changes, established non-bot author, 0.28 test ratio, ordinary churn with no reverts or hack sentiment), so anchoring preserves the stable composite score of 2 (moderate) under the Tier1 62%/Tier2 38% weighting with no linked issue. Previous run (10)Risk Assessment: moderate (2/5) DetailsTier 1 and Tier 2 signals are unchanged from the prior round (same 25 files / ~5576 lines / large blast radius, no protected/security/CI/dependency changes, established non-bot author, 0.28 test ratio, and shared forge/CLI/docs churn with ordinary fix-labeled commits but no actual reverts or hack/workaround sentiment), so anchoring preserves the stable composite score of 2 (moderate) under the Tier1 62%/Tier2 38% weighting with no linked issue. Previous run (11)Risk Assessment: moderate (2/5) DetailsRe-review anchored to the stable prior score of 2: Tier 1 is essentially unchanged from the prior round (25 files, ~5567 lines, large blast radius, no protected/security/CI/dependency changes, established non-bot author, 0.28 test ratio), and Tier 2 confirms the new steerwatch/runtime files carry no independent history while churn remains concentrated in shared forge/CLI infrastructure with ordinary fix commits and no reverts or hack/workaround sentiment; with Tier1 62%/Tier2 38% weighting (no linked issue) the composite stays at 2 (moderate). Previous run (12)Risk Assessment: moderate (2/5) DetailsRe-review, anchored to the stable prior score of 2: Tier 1 is essentially unchanged (25 files, 5535 lines, large blast radius, no protected/security/CI/dependency changes, established author, 0.28 test ratio), and Tier 2 confirms churn confined to shared forge/CLI infrastructure rather than the new steerwatch code, with no reverts or hack/workaround sentiment; with Tier1 62%/Tier2 38% weighting (no linked issue) the composite remains 2 (moderate). Previous run (13)Risk Assessment: moderate (2/5) DetailsRe-review round 7, anchored to the stable prior score of 2: Tier 1 signals (25 files, 5506 lines, large blast radius, no protected/security/CI/dependency changes, established-member author, 0.28 test ratio) are unchanged from the immediately preceding review, and Tier 2 confirms heavy churn/multi-author/fix-revert-worded activity confined to shared, actively-maintained forge/CLI infrastructure rather than the new steerwatch code itself, with zero actual reverts and zero hack/workaround sentiment; with Tier1 62%/Tier2 38% weighting (no linked issue) the composite remains 2 (moderate), unchanged from all 6 prior assessments on this PR. Previous run (14)Risk Assessment: moderate (2/5) DetailsRe-review round 6, anchored to the prior score: Tier 1 (25 files, 5487 lines, large blast radius, no protected/security/CI/dependency changes, established-member author, 0.28 test ratio) is identical to the immediately preceding review, and Tier 2 shows heavy churn/multi-author/fix-revert activity confined to shared, actively-maintained forge/CLI infra rather than the new steerwatch code itself, with no reverts or hack/workaround sentiment; with Tier1 62%/Tier2 38% weighting (no linked issue) the composite remains 2 (moderate), unchanged from all five prior assessments on this PR. Previous run (15)Risk Assessment: moderate (2/5) DetailsNo same-repo linked issue, so weights redistribute to Tier1 62%/Tier2 38%; Tier1 (25 files, 5487 lines, large blast radius, no protected paths, no CI/dependency changes, established-member author, 0.28 test ratio) is essentially unchanged from the prior five assessments, and Tier2 (computed over the subset of changed files with resolvable git history) shows heavy churn/multi-author/fix-revert activity in shared forge/CLI infra offset by zero revert/sentiment flags and very recent last-touch dates; weighted composite rounds to 2, matching and reconfirming the anchored moderate score. Previous run (16)Risk Assessment: moderate (2/5) DetailsNo same-repo linked issue, so weights redistribute to Tier1 62%/Tier2 38%; Tier1 (25 files, 5452 lines, large blast radius, no protected paths, no CI/dependency changes, established-member author, 0.28 test ratio) is essentially unchanged from the prior run, and Tier2 (computed over the subset of changed files with resolvable git history, since the actual base branch steer-followup-runs is not fully fetchable here) shows heavy churn/multi-author/fix-revert activity in shared infra offset by zero revert/sentiment flags and very recent last-touch dates; weighted composite rounds to 2, matching and confirming the prior anchored moderate score. The large volume of genuinely new, unreviewable-by-history steerwatch provenance/authority-splitting code is not captured by Tier2 churn signals and still warrants deliberate reviewer attention despite the moderate composite. Previous run (17)Risk Assessment: moderate (2/5) DetailsNo linked issue, so weights redistribute to Tier1 62% / Tier2 38%; Tier1 signals (22 files, 5092 lines, large blast radius, no protected paths, no CI/dependency changes, established member author, 0.27 test ratio) are essentially unchanged from the prior run, and Tier2 confirms heavy churn/regression/multi-author activity concentrated in shared infra offset by net-new steerwatch code with no prior history and zero revert/sentiment red flags, yielding the same composite; per re-review anchoring the prior moderate score of 2 is preserved. The security-adjacent provenance/authority-splitting logic in steerwatch still warrants extra reviewer attention despite the moderate composite. Previous run (18)Risk Assessment: moderate (2/5) DetailsNo linked issue, so weights redistribute to Tier1 62% / Tier2 38%; composite ≈ 2.14 → 2 (moderate). Large by raw size (22 files, ~4952 lines) and touches heavily-churned shared infra (internal/cli/run.go, internal/forge/forge.go, github.go), but no protected paths, no CI/workflow or dependency changes, an established MEMBER author, and a reasonable 0.27 test ratio; most added content is net-new steerwatch/provenance code with no prior git history to inflate churn. Qualitative caveat: the new steerwatch package implements provenance verification and authority-based delta splitting, which is security-adjacent logic warranting extra reviewer attention despite the moderate composite score. Re-review anchoring not applied (no tooling available to fetch the prior sticky risk-assessment comment). Previous run (19)Risk Assessment: moderate (2/5) DetailsRe-review anchoring preserved: Tier 1 signals are essentially unchanged from the prior assessment (22 files/+4830-88 lines vs prior 21/~4841, still no protected or security-sensitive paths, no CI/dependency changes, non-bot non-first-time author, adequate test ratio); Tier 2 git history on the touched forge/CLI files shows high churn and multi-author contention but no reverts or hack/workaround sentiment; with no linked issue (62/38 Tier1/Tier2 weighting) the composite remains a moderate score of 2, matching the prior run. Previous run (20)Risk Assessment: moderate (2/5) DetailsLarge diff (21 files, ~4841 lines, mostly a new steerwatch package with heavy test coverage) but low Tier 1 signals (no protected/security-sensitive paths, no CI/dependency changes, non-first-time non-bot author, adequate test ratio); Tier 2 git history on the touched forge files shows recent high-frequency, multi-author churn with some past fix commits but no reverts or hack/workaround sentiment; no same-repo linked issue, so Tier1/Tier2 weighting (62/38) yields a moderate composite score of 2. |
|
Looks good to me Previous runReviewFindingsLow
Next steps:
Previous run (2)ReviewFindingsLow
The prior review round's medium-severity finding on Next steps:
Previous run (3)ReviewFindingsMedium
Low
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or the three commit messages (all carry Next steps:
Previous run (4)ReviewFindingsMedium
Low
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or commit messages (all three commits carry Next steps:
Previous run (5)Looks good to me Previous run (6)ReviewFindingsMedium
Info
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or commit messages. All three commits carry Next steps:
Previous run (7)ReviewFindingsLow
Next steps:
Previous run (8)ReviewFindingsMedium
Low
Next steps:
Previous run (9)ReviewFindingsMedium
Low
Next steps:
Previous run (10)ReviewFindingsMedium
Low
Next steps:
Previous run (11)ReviewFindingsMedium
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or commit messages. The prior round's Next steps:
Previous run (12)ReviewFindingsMedium
Low
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or commit messages. The two correctness findings from the immediately prior review round (the Next steps:
Previous run (13)ReviewFindingsMedium
Low
No files under the repository's protected-path prefixes are touched by this PR, and no prompt-injection or non-rendering-Unicode patterns were found in the PR title, body, or commit messages. Next steps:
|
cdcc59e to
e6e8b1a
Compare
|
🤖 Review · ❌ Terminated · Started 3:19 PM UTC · Ended 3:48 PM UTC Commit: |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
🤖 Finished Review · ✅ Success · Started 3:19 PM UTC · Completed 3:48 PM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $10.90 |
e6e8b1a to
d0b68b8
Compare
|
🤖 Review · ❌ Terminated · Started 5:54 PM UTC · Ended 6:21 PM UTC Commit: |
|
🤖 Finished Review · ✅ Success · Started 4:18 AM UTC · Completed 4:41 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $10.94 |
|
On the LOW for 0a8afec (AGENTS.md row for steering.md): the row exists on the base branch — |
0a8afec to
50bda1b
Compare
|
🤖 Finished Review · ✅ Success · Started 5:18 AM UTC · Completed 5:47 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $12.68 |
|
🤖 Review · Commit: |
|
🤖 Finished Review · ✅ Success · Started 5:58 AM UTC · Completed 6:21 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $9.74 |
|
🤖 Finished Review · ✅ Success · Started 6:37 AM UTC · Completed 7:00 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $10.92 |
|
🤖 Review · Commit: |
|
🤖 Finished Review · ✅ Success · Started 7:14 AM UTC · Completed 7:34 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $8.79 |
|
🤖 Review · Commit: |
|
🤖 Finished Review · ✅ Success · Started 7:41 AM UTC · Completed 8:04 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $8.94 |
|
🤖 Finished Review · ✅ Success · Started 8:56 AM UTC · Completed 9:22 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $13.89 |
…elta The runner in a CI job has no inbound path, but every update to the work item already fires the shim, and that run's Route job already applied ADR 0054's authorization. This package turns those run records into steers: `candidateChecks` and `jobChecks` verify a candidate's provenance from server-side fields the sender cannot write (same repository, shim path and event allowlist, referenced_workflows by path and ref, Route concluded success, my stage job not skipped, bound to my work item, judged once), and `buildDelta`/`buildText` render what changed on the item since the baseline into a runner-authored envelope, split into amendments — for each accepted issue_comment run, the one comment its Route job evaluated (the newest by the run's actor created at or before the run, its text unchanged since; a re-run confers nothing) — and context the agent must not obey. A moved head is handed over as context read from the compare API, complete or declared unreadable, never as an instruction to fetch, which the agent definitions forbid. Nothing calls it yet: the watcher loop, listing and runner wiring follow in the queue-monitoring change. Carries the package types and setup (Config, Watcher, New, Start, resolveItem, resolveStageJob, markSeen/markSteered) so the package builds and its tests run; the rule resolveStageJob applies is decided one change up, with the runner that passes its hints. The forge model gains the run provenance fields (ReferencedWorkflow, WorkflowRun.Path/DisplayTitle/Actor/TriggeringActor/PullRequestNumbers/ ReferencedWorkflows), IssueComment.UpdatedAt and Issue.IsPullRequest, and the GitHub client decodes them; ListIssueCommentsSince lands end to end and CompareChanges land end to end (interface, GitHub, fake, GitLab not-supported stub) because the Client assertions force a method and all its implementations into one change. ListWorkflowRunJobs paginates so a matrix job past the first page is not read as absent. security.SanitizeAgentText is the one Unicode sanitizer both the validation-feedback prompt and the steer envelope go through, so the two cannot drift. The steer envelope regains the authority sentence the interface change left out because nothing established provenance then: with an actor, that the route job verified that actor's authorization by the same permission check that authorized this run; without one, that the update came through an authorized follow-up run; and, with no follow-up run at all, that no permission check stands behind it. How to weigh Amendments against Work-item context stays in the body, whose opening paragraph varies with what the batch holds. The test that pinned the sentence's absence is replaced by one that requires it, and the envelope contract's Structural tokens and Defanging sections lose their DRAFT markers now that this package implements them. Signed-off-by: Wayne Sun <gsun@redhat.com> Assisted-by: Claude
The delta dropped every author whose login ended in `[bot]`. That threw away the reviews of Apps the repository installed — coderabbit, qodo — which are exactly the context the review agent needs, and it decided trust from the shape of a login, which any user account can imitate: a person can be named `fullsend-ai-review` or end in `-bot`. Replace it with `ownOutput`, which reads no login shape. An item is the runner's own when its author is an exact, case-insensitive match against `Config.SelfLogins` — the logins the runner resolves at start: its own App login from the forge token, and the review App logins exactly as reusable-dispatch.yml builds REVIEW_BOT and SHARED_REVIEW_BOT (`ReviewBotLogins`, pinned to the workflow file by a test) — or when the forge itself says the author is an App (`user.type == "Bot"`, now decoded into `forge.IssueComment.AuthorIsApp` and `PullRequestReview.AuthorIsApp`) and the body carries a `<!-- fullsend:` marker. Exact login is the primary rule — the post-fix and post-code comments carry no marker, so only the login keeps the coder App's own output out — and SelfLogins is required: Start refuses an empty list, because a runner that cannot name its own login would steer itself. The marker rule supplements it for markered bodies from any App, whatever the login. A human is never own output, so an authorized `/fs-fix` that quotes a status comment is still delivered as an amendment. The failure direction is fixed by construction: a miss leaves an App's text in context, where it is data the agent must not obey, and nothing here can move an author toward amendments — those remain solely the actors the Route job authorized. Tests cover the four author classes through buildDelta, the user-account-shaped-like-a-bot cases classified by type, and the refusal to start without a resolved login; each of the three tempting shortcuts (suffix, prefix, marker without the App gate) fails them, and the GitHub decoder tests pin AuthorIsApp to `user.type` at the wire. Signed-off-by: Wayne Sun <gsun@redhat.com> Assisted-by: Claude
ADR 0106 ends its Decision by deferring "how content provenance and actor authority constrain agent behavior" to a separate ADR. This is that ADR, for the updates a run in flight absorbs: a steer is authorized once, in the follow-up run's Route job, and the runner verifies provenance only. An amendment is text from an actor that job authorized, bound to the text it saw; everything else is context. Context excludes only fullsend's own output — by exact resolved login or by the forge's App verdict plus a fullsend marker, never by the shape of a login — so a repository-installed App's review reaches the agent as context. The two author rules are laid out with the alternatives they were chosen over, and the ADR 0098 alignment is stated: the Route job is that ADR's evaluation for the event. steering.md gains the provenance table, the amendment boundary with its per-arm audit, the new "Whose text is context" rules, and the work item's baseline; architecture.md records the decision and closes the question ADR 0106 left open; the threat model's "treat all PR content as untrusted?" question is annotated with the part this decides. Signed-off-by: Wayne Sun <gsun@redhat.com> Assisted-by: Claude
|
🤖 Finished Review · ✅ Success · Started 9:45 AM UTC · Completed 10:11 AM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $12.58 |
What this adds
A run in flight can now tell a legitimate update to its work item from anything else, and render what changed into text the agent can act on safely. Nothing calls it yet.
internal/steerwatchverifies a candidate follow-up run's provenance from run records the sender cannot write — same repository, the shim's path and an allowlisted event,referenced_workflowsequal by path and ref,Routeconcluded success, this run's stage job not skipped, bound to this work item, judged once — and builds the delta: the item's current state against a baseline, split intoRoutejob authorized for an acceptedissue_commentrun, bound to the text that authorization saw (an edit needs a new one); the agent acts on these;Context excludes only fullsend's own output, and never by the shape of a login (a person can be named
fullsend-ai-reviewor end in-bot): an exact match against the logins the runner resolves at start (Config.SelfLogins: its own App login andReviewBotLogins(owner), which is pinned to theREVIEW_BOT/SHARED_REVIEW_BOTstrings inreusable-dispatch.ymlby a test), or an App-authored body —user.type == "Bot", now decoded intoforge.IssueComment.AuthorIsApp/PullRequestReview.AuthorIsApp— carrying a<!-- fullsend:marker. A repository-installed App's review therefore reaches the review agent as context; a human is never excluded, so an authorized/fs-fixthat quotes a status comment is still an amendment.Startrefuses an emptySelfLogins.ADR 0118 records the decision — it is the ADR that ADR 0106's Decision deferred "content provenance and actor authority" to — with the two author rules and the alternatives they were chosen over. steering.md gains the provenance table, the amendment boundary with its per-arm audit, "Whose text is context", and the work item's baseline.
Decided in the PRs above
Config.SelfLoginswiring (the runner resolves its login viaGetAuthenticatedUserand declines steering when it cannot), and the ruleresolveStageJobapplies (GITHUB_JOBfirst, harness slug second, fails closed) — is the queue-monitoring PR (ADR 0119).resolveStageJoblands here only because the package's setup needs it to build.Scope notes
resolveItemcomment; no code path for it changes.Also carried, because the
forge.Clientassertions force itListIssueCommentsSinceend to end (interface, GitHub, fake withatOrAfter, GitLab not-supported stub), the run provenance fields onforge.WorkflowRun,Issue.IsPullRequest,IssueComment.UpdatedAt,ListWorkflowRunJobspagination, andsecurity.SanitizeAgentText— the one Unicode sanitizer the validation-feedback prompt and the steer envelope share.Verification
go build ./...,go vet ./...,go test -raceoninternal/steerwatch,internal/forge,internal/forge/github,internal/security. Three deliberate mutations of the author rule — the old[bot]suffix, afullsend-ai-prefix, the marker rule without the App gate — each fail the new tests.make lintwith docs staged.