[core] Optimize tag reads during snapshot expiration - #9967
Merged
JingsongLi merged 1 commit intoSep 19, 2026
Merged
Conversation
leaves12138
approved these changes
Sep 18, 2026
leaves12138
left a comment
Contributor
There was a problem hiding this comment.
LGTM. The candidate-scoped tag reads reduce retained state while preserving tag protection and changelog-decoupled cleanup behavior. Repeated scanning across batches is an acceptable trade-off for reducing memory pressure.
Validated with Java 8: 161 existing test cases across ManifestFileTest, ExpireSnapshotsTest, FileDeletionTest, ChangelogExpireTest, and TagTest, plus a local cross-batch tag-retention regression (162 passed in total). No blocking issues found.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Snapshot expiration previously built a full data-file index for every referenced tag. On large tables, concurrent tag reads retained memory proportional to the number of tags multiplied by all table files and could cause OOM.
This change:
The filename filter is streaming rather than storage-index based; manifests still need to be scanned within matching partitions and buckets.
Tests