Skip to content

[telemetry] fix: clean up platform probes - #73

Merged
Wangmerlyn merged 2 commits into
mainfrom
codex/fix-telemetry-probes
Jun 27, 2026
Merged

[telemetry] fix: clean up platform probes#73
Wangmerlyn merged 2 commits into
mainfrom
codex/fix-telemetry-probes

Conversation

@Wangmerlyn

@Wangmerlyn Wangmerlyn commented Jun 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • Shut down NVML and ROCm SMI vendor probes after detection so platform discovery does not leave telemetry handles open.
  • Align ROCm detection with the rsmi_init / rsmi_shut_down API used by the rest of KeepGPU.
  • Add guarded MPS telemetry so Mac M series reports a macm device with best-effort memory counters and nullable unsupported fields.
  • Update AGENTS.md, README, architecture, MCP, CLI reference, and the implementation plan docs.

Test Plan

  • PYTHONPATH=src pytest tests/utilities/test_platform_manager.py tests/utilities/test_gpu_info.py -q
  • PYTHONPATH=src pytest tests -q
  • PYTHONPATH=src mkdocs build
  • pre-commit run --all-files

@coderabbitai

coderabbitai Bot commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@Wangmerlyn, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 21 minutes and 43 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits.

🚦 How do rate limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 3abc098d-32e4-4bf2-89ce-760bae2f3c60

📥 Commits

Reviewing files that changed from the base of the PR and between 17053dc and 9671242.

📒 Files selected for processing (11)
  • AGENTS.md
  • README.md
  • docs/concepts/architecture.md
  • docs/guides/cli.md
  • docs/guides/mcp.md
  • docs/plans/telemetry-probe-hygiene.md
  • docs/reference/cli.md
  • src/keep_gpu/utilities/gpu_info.py
  • src/keep_gpu/utilities/platform_manager.py
  • tests/utilities/test_gpu_info.py
  • tests/utilities/test_platform_manager.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/fix-telemetry-probes

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements telemetry probe hygiene fixes and adds best-effort Apple Silicon/MPS telemetry support. Specifically, it ensures that platform detection probes for CUDA and ROCm clean up and shut down their respective vendor libraries immediately after initialization to avoid leaving open handles. It also aligns the ROCm detection probe to use the correct rsmi_init and rsmi_shut_down APIs. Additionally, it introduces best-effort MPS memory telemetry for Apple Silicon using PyTorch's MPS backend, returning nullable fields for unsupported metrics like utilization. Finally, the PR updates relevant documentation, removes GPU-presence constraints on mocked NVML tests to allow them to run in no-GPU CI environments, and adds comprehensive unit tests for the new probe lifecycles and MPS telemetry. There are no review comments, so I have no feedback to provide.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@Wangmerlyn

Copy link
Copy Markdown
Owner Author

Local re-review after the torch fallback guard fix found no Critical, Important, or Minor issues. Verified locally with focused telemetry tests (12 passed, 1 skipped), full test suite (62 passed, 12 skipped), mkdocs build, and git diff --check; GitHub checks are green.

@Wangmerlyn
Wangmerlyn merged commit 5a88735 into main Jun 27, 2026
5 checks passed
@Wangmerlyn
Wangmerlyn deleted the codex/fix-telemetry-probes branch June 27, 2026 12:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant