Summary
Requalify scale-sensitive settings for AICR's current NFD, NVSentinel, Topograph, Slurm, and monitoring components before adopting them as scale defaults.
Problem
Large-cluster tuning results cannot be copied directly when component versions or deployment topology differ. In particular, AICR deploys standalone NFD while some prior configurations tuned GPU Operator's embedded NFD. Current public pins must be measured directly.
Work
At representative qualification bands, measure and evaluate:
- CPU and memory requests and limits
- replica count, concurrency, worker, and queue settings
- Kubernetes client QPS and burst
- rollout and reconciliation duration
- scheduler and topology convergence
- monitoring discovery, scrape coverage, and cardinality
- failure behavior under partial rollout and control-plane throttling
For each component, record the exact AICR version, rendered values, cluster size, workload, and observed resource envelope.
Success criteria
Non-goal
This task does not tune managed control planes, etcd, kubelets, host sysctls, DNS, registries, or CNI infrastructure.
Summary
Requalify scale-sensitive settings for AICR's current NFD, NVSentinel, Topograph, Slurm, and monitoring components before adopting them as scale defaults.
Problem
Large-cluster tuning results cannot be copied directly when component versions or deployment topology differ. In particular, AICR deploys standalone NFD while some prior configurations tuned GPU Operator's embedded NFD. Current public pins must be measured directly.
Work
At representative qualification bands, measure and evaluate:
For each component, record the exact AICR version, rendered values, cluster size, workload, and observed resource envelope.
Success criteria
Non-goal
This task does not tune managed control planes, etcd, kubelets, host sysctls, DNS, registries, or CNI infrastructure.