You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Seven hands-on labs for running AI workloads on Kubernetes: GPU scheduling, distributed training, model serving, and agents. Runs on a local kind cluster, no GPU required. Companion to a KubeCon EU 2026 talk.
Multi-backend LLM serving and training platform — vLLM/Triton/Ray Serve/KServe/BentoML behind one contract, Kueue/Karpenter GPU orchestration, Ray Train/FSDP/DeepSpeed with LoRA/PEFT and DVC, MLflow/W&B tracking, and a tool-grounded LangGraph advisor — CI-validated without real GPU cost.
Make Kueue the single admission controller for mixed Slurm, Ray, and Kubernetes workloads on one shared pool. Experimental, live-validated prototype. Apache-2.0.
Production-style model reliability platform with drift checks, durable incidents, transactional notifications, RCA and lineage, a tested SignalOps console, and a study guide.
Production-style Kubernetes MLOps control plane with MLflow 3, Airflow 3.3, KServe/Kueue architecture labs, a tested KubeOps console, failure drills, and a study guide.
Production-style KServe platform with Open Inference V2, canary and rollback controls, a tested ServeOps console, observability, and guided study artifacts.
Production-style Metaflow and Airflow training control plane with partitioned backfills, checkpoint recovery, Kueue design, a tested TrainOps console, and a study guide.