Summary
The controller creates and manages the child Deployment for each worker version, but there's no way to configure that Deployment's rolling-update strategy. Since the controller never sets spec.strategy, owned Deployments fall back to the Kubernetes default (RollingUpdate, maxUnavailable/maxSurge 25%/25%). On large fleets this makes in-place rolling restarts of a Current version disruptive: up to ~25% of pods can go unavailable at once, which spikes schedule_to_start.
Users would like a field on the WorkerDeployment spec to set maxUnavailable/maxSurge, which the controller then applies to the Deployments it owns and reconciles.
Current behavior
TemporalWorkerDeploymentSpec exposes replicas, template, minReadySeconds, progressDeadlineSeconds, rollout, sunset, and workerOptions — there is no Deployment-level strategy field. rollout.strategy (Manual/AllAtOnce/Progressive) controls Temporal traffic routing across versions, not how pods roll within a single version's Deployment. The builder in internal/k8s/deployments.go doesn't set spec.strategy on the owned Deployment, so it inherits the cluster default of 25%/25%.
Why this matters
Users may periodically do in-place rolling restarts of Current workers with the same build ID (e.g. non-workflow config/bug-fix changes — the UnsafeCustomBuildID path, plus ordinary pod-template drift on the current version). In those cases the restart happens on the same Deployment, so its own strategy.rollingUpdate is the only thing governing availability. With 25% of a large fleet cycling at once, they lose enough pollers to build a backlog. They'd like to declare something conservative like maxUnavailable: 5% and have it stick.
Filing on behalf of a Temporal Cloud customer running large worker fleets in production.
Summary
The controller creates and manages the child
Deploymentfor each worker version, but there's no way to configure that Deployment's rolling-update strategy. Since the controller never setsspec.strategy, owned Deployments fall back to the Kubernetes default (RollingUpdate,maxUnavailable/maxSurge25%/25%). On large fleets this makes in-place rolling restarts of a Current version disruptive: up to ~25% of pods can go unavailable at once, which spikesschedule_to_start.Users would like a field on the
WorkerDeploymentspec to setmaxUnavailable/maxSurge, which the controller then applies to the Deployments it owns and reconciles.Current behavior
TemporalWorkerDeploymentSpecexposesreplicas,template,minReadySeconds,progressDeadlineSeconds,rollout,sunset, andworkerOptions— there is no Deployment-levelstrategyfield.rollout.strategy(Manual/AllAtOnce/Progressive) controls Temporal traffic routing across versions, not how pods roll within a single version's Deployment. The builder ininternal/k8s/deployments.godoesn't setspec.strategyon the owned Deployment, so it inherits the cluster default of 25%/25%.Why this matters
Users may periodically do in-place rolling restarts of Current workers with the same build ID (e.g. non-workflow config/bug-fix changes — the
UnsafeCustomBuildIDpath, plus ordinary pod-template drift on the current version). In those cases the restart happens on the same Deployment, so its ownstrategy.rollingUpdateis the only thing governing availability. With 25% of a large fleet cycling at once, they lose enough pollers to build a backlog. They'd like to declare something conservative likemaxUnavailable: 5%and have it stick.