fix: tolerate node not-ready/unreachable forever on managed pods#21
Merged
Conversation
VM pods are pinned to their virtual node, so taint eviction cannot reschedule them; it can only delete them, and vk-cocoon then removes the VM. Without explicit entries the DefaultTolerationSeconds plugin injects the 300s pair, so any node flap past 5 minutes (e.g. a control-plane DB lockup failing lease renewals) batch-destroys every VM on the node.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #20.
Problem
newManagedPodtolerates only the vk provider taint, so the apiserver'sDefaultTolerationSecondsadmission plugin injects the default 300s tolerations fornode.kubernetes.io/not-ready/unreachable. When a virtual node stays NotReady past 5 minutes, the taint-eviction-controller batch-deletes every VM pod on it, and vk-cocoon'sDeletePoddestroys each VM; the recreated pods boot fresh, so running VM state is reset.That is what happened in the 2026-07-18 mindos incident: a Kine DB lockup failed the vk node's lease renewals → NotReady → 4 VM pods deleted in the same second → VMs reset.
Eviction is pure loss for these pods: they are pinned to one virtual node via
NodeSelector, so eviction cannot reschedule them; the node's readiness reflects vk-cocoon's control-plane connectivity, not VM health; and the CocoonSet operator is the component that heals genuinely lost pods.Fix
Add permanent
Exists/NoExecutetolerations for both taints innewManagedPod(niltolerationSeconds), covering agent and toolbox pods — the same pair the DaemonSet controller injects for node-pinned workloads. Explicit entries suppress the injected 300s pair.Tests
TestNewManagedPodToleratesNodeFlapForeverpins key,Exists,NoExecute, and assertsTolerationSecondsis nil, so a regression restoring the 300s default fails.Gates
go build/go vet/go test ./... -racegreen;make lint0 issues on GOOS=linux and GOOS=darwin.