Repository navigation
Merged
Conversation
…he owning entity's rows
- queue RowKey is {runEpoch}|{run}|{task}; the epoch is stored once per run so re-enqueues keep one row
- schema v3 migration re-keys an existing backlog, writing partitions in parallel
- the queue epoch is the run's StartedUtc, so concurrent enqueues of a run can no longer write two rows for a task - schema v3 migration keys runs by their Run row's start time and skips rows already re-keyed - claim reads one page sized to the batch - run-scoped reads (dispatch check, claim release) read the run's contiguous key range instead of one GET per task - RemoveRunAsync is scoped to one outing of a recurring run name and deletes in batches - status snapshot streams with a projection and keeps only the head rows; re-scans wait 4x the last scan - CancelRunAsync cancels the live run and drops its queue rows; re-drive requeues orphans in one batch
- each run is one partition of a Work table; a task's state is its row, and claim, finish and the aggregation barrier are single partition transactions - the pump claims from a Ready list ordered by band then run age; lapsed leases are reclaimed and a task interrupted MaxRetries times is failed - PostExecution is an ordinary leased work item; child runs are placeholders the parent waits on - a run name still in progress is not started again - removes the re-drive, recovery pass, status writer, queue index and remaining counter, and their settings; the previous tables are dropped at startup - per-partition write limiter at 1,900 ops/s
…pts out - QueueOrchestration/QueueOrchestrationFromFile take allowCollision (default true); false skips the run while another run of that name is unfinished - the stamped operation context carries RunKey, so a child registers under its exact parent; a run name still resolves to the newest outing - cancel, reprioritize and task cancel by name apply to every unfinished run of that name - pump job ids are unique per claim; status counts sum runs sharing a name
…CraftOrchestrator
…s active - OrchestratorBridge.IsRunActive(name) lets a caller check before queueing - a start skipped at drain time logs a warning - Start-CraftOrchestrator skips with a warning and returns "-Skipped" when AllowCollision is false and the name is active
… and stop blocked runs starving the queue - MaxConcurrency (0 = no limit) caps a run's tasks running at once; the pump applies it against the claims it holds, at no storage cost; dropped for sequential runs - StopOnFailure makes a sequential run cancel its remaining steps at the first failure (including a step that keeps crashing); carrying on stays the default - a run whose pending tasks ran out is skipped until its counts move, a claim is released, or its earliest lease is due; other empty runs back off 30 s doubling to 15 min, so runs waiting at the head no longer starve the runs behind them - Ready rows carry the run's mode and are read 1,000 to a page; the running-row scan per claim is bounded - bridge and Start-CraftOrchestrator take maxConcurrency and stopOnFailure - tests: shared harness, mode tests, storage-cost pins, Azurite end-to-end runs, wrapper inheritance pins
…her dequeued it - a cancel that no longer finds the job in the queue is now reported by the dispatcher's skip (or the job's start) instead of being dropped, so the claim is finished as Cancelled rather than lapsing and running half an hour later - tests: lease renewal, graceful-stop release, buffered cancel, reprioritise, stale Ready entries, legacy table drop, result cleanup, and the log lines the health tooling parses
…orker threw - invocations pass their own output collection; after a failed async invocation the next EndInvoke returned null, so the following call on that worker lost its output (a sequential run dropped the result of the step after each failure)
…dexes and run diagnostics - instance lock row: the pump claims nothing until it holds it, keeps it while its own claimed tasks still run on shutdown, then releases it; claims and sequential drivers of any other process are taken back on sight while held - owner ids are unique per process start - in-process run-changed events wake the pump on create, finish, release, child and band changes, and also refill as soon as the job buffer drains - active-run row written first and removed last; Ready and Finished are rebuilt from it at startup and whenever an index write has failed for good - index writes retried; failed finishes retried in the background for 30 min - a task finished twice in one batch is applied once (one entity per transaction) - range reads send a page size; status, claim and cancel reads are bounded; queued rows carry their run key and position so cancelling one is a point read - status views no longer count claims made after the cached snapshot twice - InspectRun and RepairIndexes on the bridge
- PerfApi probes: recording task/PostExecution, run starter, durable run reader, bridge and legacy-table endpoints - checks for PostExecution contract, failures, child gating and inheritance, collisions, cancel, sequential runs, priority order, status consistency, restart recovery and perf gates - MaxConcurrency/StopOnFailure checks gated on image capability (SKIP when absent)
…, and warn on ignored run options - the worker summary and metrics snapshot added the live local queue to the cached snapshot's unclaimed count, double-counting claims made since; all merged counts now use one helper - Start-CraftOrchestrator warns when MaxConcurrency is set on a sequential run or StopOnFailure on a non-sequential one - e2e status check bounds the durable queue rather than the global one
…e2e perf gates - the pump no longer claims tasks of a run being cancelled (it raced the cancel for the same rows and header); it cancels a page per visit only when no cancel is in progress, so an interrupted cancel still finishes - cancel reuses the rows it just read and keeps going until the pending range is empty; CancelRun reports what the run records - e2e perf gates recalibrated to the faster pump, plus a cancel-5000 gate
…g the lock release - WithdrawJob takes a queued job out under the record lock, so it cannot start after being handed back - Report the instance lock release (or why it lapses) at shutdown
…ite or a storage account
…ne transaction - Header, step and next pending row read together; the aggregation is claimed directly at the barrier - Falls back to the batched finish and a separate claim if the combined write fails - e2e gate on per-step cost
… collisions
Runs named Name-{guid} (a queue id appended per outing) collide with each other and with Name when collisions are off.
…ts, job lists and the queue page - The driving job carries each step's name and keeps a finished record per step; the last step's outcome is the job's - Run summaries and task lists cover the latest outing of a name only - e2e check orch-seq each-step-visible
- app-only and signed-in principals both carry the validated token's claims (azp, scp, ...) - pin the app-only vs user split and the rewrite with middleware tests
…last GC of any kind - Worker.MemoryTrimIntervalMinutes (default 5, 0 disables) runs TrimMemory periodically so idle hosts return freed heap - Memory breakdown uses GCKind.Any instead of FullBlocking, which reported zeros until a blocking gen2 ran
JohnDuprey
approved these changes
Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This pull request introduces several enhancements and refinements to the Craft orchestrator and supporting infrastructure, focusing on improved orchestration control, better parent-child run tracking, and more accurate metrics. The most significant changes include new orchestration options (like collision control and concurrency limits), improved parent/child run key handling, new orchestration inspection utilities, orchestration and recovery logic is simplified while offering more advanced queue logic, and some bug fixes and API improvements.
Orchestration control and options:
Added new orchestration options to
Start-CraftOrchestratorand the orchestrator bridge:AllowCollision(to prevent overlapping runs with the same name),MaxConcurrency(limit parallelism within a run), andStopOnFailure(sequential runs stop on the first failure). These are now passed through the PowerShell entry point and into the backend orchestration logic. [1] [2] [3] [4] [5] [6] [7]The orchestrator bridge now exposes
IsRunActiveto check if a run with a given name is still active, and will skip new runs with the same name ifAllowCollisionis false.Parent/child run tracking and lineage:
RunKey(unique outing identifier) instead of justRunName, ensuring precise tracking of overlapping runs and correct parent/child relationships. This includes updates to job records and orchestration queueing. [1] [2] [3]Diagnostics and orchestration utilities:
InspectRun(returns JSON diagnostics for a run) andRepairIndexes(rebuilds internal indexes and returns a summary), available from PowerShell for live troubleshooting.Metrics and queue status improvements:
API and bug fixes:
QueueStatusBridgeto use the correct method (GetRunJobs) for fetching jobs by run name.These changes collectively make orchestration more robust, observable, and configurable, especially in scenarios involving overlapping or recurring runs.