Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Loading a newer checkpoint can leave many older epochs resident in Triton until the idle timeout, exhausting GPU memory.
Keep the newest
TRITON_MAX_CHECKPOINTS_PER_EIDepochs per EID, defaulting to 1, in the sharedsetup_triton_modelloader. Artifacts of the same epoch share a slot. Serialize eviction and loading with an EID lock, including cached loads; reject requests outside the retained window. Names without an EID/epoch bypass the retention rule.Wait for unloading to finish before deleting model files or allocating the replacement. If it does not finish within 60 seconds, raise instead of continuing. This avoids the old/new GPU allocation overlap observed with native Triton version replacement. Existing tasks using evicted checkpoints can fail on subsequent inference.
Validation: 24 temporary focused checks passed, including concurrency, cached backlog trimming, unload failures/timeouts, and numeric epoch ordering. An isolated Triton 2.69.0 GPU probe with Redis and ONNX/Python backends confirmed mixed-backend eviction, latest-N behavior, stale inference failures, old Python GPU allocations finalized before replacement allocation, and cleanup. All commit hooks, including Ruff and full Miniray ty, passed.