Skip to content

chore: rebuild the orchestration queue and table mechanism to be faster, more robust and offer more granular queue options - #47

Merged
Zacgoose merged 24 commits into
mainfrom
dev
Oct 7, 2026
Merged

Zacgoose merged 24 commits into
mainfrom
dev

Conversation

@Zacgoose

@Zacgoose Zacgoose commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

This pull request introduces several enhancements and refinements to the Craft orchestrator and supporting infrastructure, focusing on improved orchestration control, better parent-child run tracking, and more accurate metrics. The most significant changes include new orchestration options (like collision control and concurrency limits), improved parent/child run key handling, new orchestration inspection utilities, orchestration and recovery logic is simplified while offering more advanced queue logic, and some bug fixes and API improvements.

Orchestration control and options:

  • Added new orchestration options to Start-CraftOrchestrator and the orchestrator bridge: AllowCollision (to prevent overlapping runs with the same name), MaxConcurrency (limit parallelism within a run), and StopOnFailure (sequential runs stop on the first failure). These are now passed through the PowerShell entry point and into the backend orchestration logic. [1] [2] [3] [4] [5] [6] [7]

  • The orchestrator bridge now exposes IsRunActive to check if a run with a given name is still active, and will skip new runs with the same name if AllowCollision is false.

Parent/child run tracking and lineage:

  • Improved parent/child run tracking by passing RunKey (unique outing identifier) instead of just RunName, ensuring precise tracking of overlapping runs and correct parent/child relationships. This includes updates to job records and orchestration queueing. [1] [2] [3]

Diagnostics and orchestration utilities:

  • Added orchestration inspection and repair utilities: InspectRun (returns JSON diagnostics for a run) and RepairIndexes (rebuilds internal indexes and returns a summary), available from PowerShell for live troubleshooting.

Metrics and queue status improvements:

  • Improved queue metrics: now uses a more accurate count of jobs waiting in durable storage, and fixes the memory breakdown GC mode for more representative results. [1] [2]

API and bug fixes:

  • Fixed bug in QueueStatusBridge to use the correct method (GetRunJobs) for fetching jobs by run name.
  • Added a new method to encode normalized EasyAuth principals, preserving claims for downstream consumers.

These changes collectively make orchestration more robust, observable, and configurable, especially in scenarios involving overlapping or recurring runs.

- queue RowKey is {runEpoch}|{run}|{task}; the epoch is stored once per run so re-enqueues keep one row

- schema v3 migration re-keys an existing backlog, writing partitions in parallel
- the queue epoch is the run's StartedUtc, so concurrent enqueues of a run can no longer write two rows for a task
- schema v3 migration keys runs by their Run row's start time and skips rows already re-keyed
- claim reads one page sized to the batch
- run-scoped reads (dispatch check, claim release) read the run's contiguous key range instead of one GET per task
- RemoveRunAsync is scoped to one outing of a recurring run name and deletes in batches
- status snapshot streams with a projection and keeps only the head rows; re-scans wait 4x the last scan
- CancelRunAsync cancels the live run and drops its queue rows; re-drive requeues orphans in one batch
- each run is one partition of a Work table; a task's state is its row, and
  claim, finish and the aggregation barrier are single partition transactions
- the pump claims from a Ready list ordered by band then run age; lapsed leases
  are reclaimed and a task interrupted MaxRetries times is failed
- PostExecution is an ordinary leased work item; child runs are placeholders the
  parent waits on
- a run name still in progress is not started again
- removes the re-drive, recovery pass, status writer, queue index and remaining
  counter, and their settings; the previous tables are dropped at startup
- per-partition write limiter at 1,900 ops/s
…pts out

- QueueOrchestration/QueueOrchestrationFromFile take allowCollision (default
  true); false skips the run while another run of that name is unfinished
- the stamped operation context carries RunKey, so a child registers under its
  exact parent; a run name still resolves to the newest outing
- cancel, reprioritize and task cancel by name apply to every unfinished run
  of that name
- pump job ids are unique per claim; status counts sum runs sharing a name
…s active

- OrchestratorBridge.IsRunActive(name) lets a caller check before queueing
- a start skipped at drain time logs a warning
- Start-CraftOrchestrator skips with a warning and returns "-Skipped" when
  AllowCollision is false and the name is active
… and stop blocked runs starving the queue

- MaxConcurrency (0 = no limit) caps a run's tasks running at once; the pump
  applies it against the claims it holds, at no storage cost; dropped for
  sequential runs
- StopOnFailure makes a sequential run cancel its remaining steps at the first
  failure (including a step that keeps crashing); carrying on stays the default
- a run whose pending tasks ran out is skipped until its counts move, a claim is
  released, or its earliest lease is due; other empty runs back off 30 s doubling
  to 15 min, so runs waiting at the head no longer starve the runs behind them
- Ready rows carry the run's mode and are read 1,000 to a page; the running-row
  scan per claim is bounded
- bridge and Start-CraftOrchestrator take maxConcurrency and stopOnFailure
- tests: shared harness, mode tests, storage-cost pins, Azurite end-to-end runs,
  wrapper inheritance pins
…her dequeued it

- a cancel that no longer finds the job in the queue is now reported by the
  dispatcher's skip (or the job's start) instead of being dropped, so the claim
  is finished as Cancelled rather than lapsing and running half an hour later
- tests: lease renewal, graceful-stop release, buffered cancel, reprioritise,
  stale Ready entries, legacy table drop, result cleanup, and the log lines the
  health tooling parses
…orker threw

- invocations pass their own output collection; after a failed async invocation
  the next EndInvoke returned null, so the following call on that worker lost
  its output (a sequential run dropped the result of the step after each failure)
…dexes and run diagnostics

- instance lock row: the pump claims nothing until it holds it, keeps it while
  its own claimed tasks still run on shutdown, then releases it; claims and
  sequential drivers of any other process are taken back on sight while held
- owner ids are unique per process start
- in-process run-changed events wake the pump on create, finish, release, child
  and band changes, and also refill as soon as the job buffer drains
- active-run row written first and removed last; Ready and Finished are
  rebuilt from it at startup and whenever an index write has failed for good
- index writes retried; failed finishes retried in the background for 30 min
- a task finished twice in one batch is applied once (one entity per transaction)
- range reads send a page size; status, claim and cancel reads are bounded;
  queued rows carry their run key and position so cancelling one is a point read
- status views no longer count claims made after the cached snapshot twice
- InspectRun and RepairIndexes on the bridge
- PerfApi probes: recording task/PostExecution, run starter, durable run reader, bridge and legacy-table endpoints
- checks for PostExecution contract, failures, child gating and inheritance, collisions, cancel, sequential runs, priority order, status consistency, restart recovery and perf gates
- MaxConcurrency/StopOnFailure checks gated on image capability (SKIP when absent)
…, and warn on ignored run options

- the worker summary and metrics snapshot added the live local queue to the
  cached snapshot's unclaimed count, double-counting claims made since; all
  merged counts now use one helper
- Start-CraftOrchestrator warns when MaxConcurrency is set on a sequential run
  or StopOnFailure on a non-sequential one
- e2e status check bounds the durable queue rather than the global one
…e2e perf gates

- the pump no longer claims tasks of a run being cancelled (it raced the cancel
  for the same rows and header); it cancels a page per visit only when no
  cancel is in progress, so an interrupted cancel still finishes
- cancel reuses the rows it just read and keeps going until the pending range
  is empty; CancelRun reports what the run records
- e2e perf gates recalibrated to the faster pump, plus a cancel-5000 gate
…g the lock release

- WithdrawJob takes a queued job out under the record lock, so it cannot start after being handed back
- Report the instance lock release (or why it lapses) at shutdown
…ne transaction

- Header, step and next pending row read together; the aggregation is claimed directly at the barrier
- Falls back to the batched finish and a separate claim if the combined write fails
- e2e gate on per-step cost
… collisions

Runs named Name-{guid} (a queue id appended per outing) collide with each other and with Name when collisions are off.
…ts, job lists and the queue page

- The driving job carries each step's name and keeps a finished record per step; the last step's outcome is the job's
- Run summaries and task lists cover the latest outing of a name only
- e2e check orch-seq each-step-visible
- app-only and signed-in principals both carry the validated token's claims (azp, scp, ...)
- pin the app-only vs user split and the rewrite with middleware tests
…last GC of any kind

- Worker.MemoryTrimIntervalMinutes (default 5, 0 disables) runs TrimMemory periodically so idle hosts return freed heap
- Memory breakdown uses GCKind.Any instead of FullBlocking, which reported zeros until a blocking gen2 ran
@Zacgoose
Zacgoose merged commit 9efa184 into main Oct 7, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants