[telemetry] fix: clean up platform probes - #73
Conversation
|
Warning Review limit reached
More reviews will be available in 21 minutes and 43 seconds. Learn how PR review limits work. Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file). ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits. 🚦 How do rate limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (11)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request implements telemetry probe hygiene fixes and adds best-effort Apple Silicon/MPS telemetry support. Specifically, it ensures that platform detection probes for CUDA and ROCm clean up and shut down their respective vendor libraries immediately after initialization to avoid leaving open handles. It also aligns the ROCm detection probe to use the correct rsmi_init and rsmi_shut_down APIs. Additionally, it introduces best-effort MPS memory telemetry for Apple Silicon using PyTorch's MPS backend, returning nullable fields for unsupported metrics like utilization. Finally, the PR updates relevant documentation, removes GPU-presence constraints on mocked NVML tests to allow them to run in no-GPU CI environments, and adds comprehensive unit tests for the new probe lifecycles and MPS telemetry. There are no review comments, so I have no feedback to provide.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
|
Local re-review after the torch fallback guard fix found no Critical, Important, or Minor issues. Verified locally with focused telemetry tests ( |
Summary
rsmi_init/rsmi_shut_downAPI used by the rest of KeepGPU.macmdevice with best-effort memory counters and nullable unsupported fields.Test Plan
PYTHONPATH=src pytest tests/utilities/test_platform_manager.py tests/utilities/test_gpu_info.py -qPYTHONPATH=src pytest tests -qPYTHONPATH=src mkdocs buildpre-commit run --all-files