BUG FIX: Shut down relaunched cluster workers when the plan is stopped - #839
Open
Arthur031221 wants to merge 1 commit into
Open
Arthur031221 wants to merge 1 commit into
Arthur031221 wants to merge 1 commit into
Conversation
When a cluster worker is terminated (for example by cancel()) and then relaunched, the new node was stored in the backend but not in the cluster registry. stopCluster() then closed the old node and left the relaunched worker process and its connection open until garbage collection.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A user of
plan(multisession), or ofplan(cluster, interrupts = TRUE), who cancels a future and then starts another one gets a relaunched worker thatplan(sequential)never shuts down; its process and connection stay open until garbage collection, which also produces the "closing unused connection" warning. The same happens withplan(cluster)when a worker died by other means and was relaunched.When a terminated worker is relaunched (in
requestNode()andhandleInterruptedFuture()), the new node is stored in the backend but not inclusterRegistry.clusterRegistry$stopCluster()then closes the old, already closed, node. This change addsclusterRegistry$replaceNode()and calls it at both relaunch sites, when the backend's workers are the cluster the registry started. A cluster passed in by the user, for exampleplan(cluster, workers = cl), is not owned by the registry, so nothing changes for it.Reproduction, before the change (Linux, one worker, the backend object held so garbage collection cannot hide it):
After the change the worker is gone and there are 0 connections.
The new test
test-cancel-relaunch-cluster.Rcounts open connections around this sequence. On the parent commit it fails (before = 0, after = 1, thenError: n <= n0 is not TRUE); with the change it passes (before = 0, after = 0).R CMD check --as-cranon the change ran 112 test scripts; 111 pass, andtest-multisession-libpaths.Rfails the same way on the parent commit in my environment. The existingcancel,cluster,worker-termination,interrupts-from-worker-itself,value-error-cancels-set,cluster-connection-clashes,planandresettests are among those that pass.This also touches the symptom in #830. The loop from #820 (100 iterations of 10 erroring futures on 2 multisession workers,
options(warn = 1), counting the warning lines) printed "closing unused connection" between 7 and 9 times per run on the parent commit across five runs, and 0 times in every run with the change. The leak is intermittent, so I only claim the cancel and relaunch case shown above. I did not check whether other causes of that warning remain.