You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
AMD Strix Halo / Ryzen AI Halo local LLM setup and benchmark guide for Ryzen AI MAX+ 395 and Radeon 8060S: Ollama, llama.cpp Vulkan/RADV, ROCm, 101 t/s Qwen3-Coder, CHADROCK MTP, 120B GGUF, and raw evidence.
Claude Code skill for AMD Strix Halo (Ryzen AI MAX+ 395) ML setup. Handles PyTorch installation (official wheels don't work with gfx1151), GTT memory config, and environment setup. Enables 30B parameter models.
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
vLLM on ROCm 7.2.2 for AMD Strix Halo (gfx1151): source-build recipe, upstream patches, environment lockfile, systemd unit — and both the failed ROCm 7.1.3 and successful 7.2.2 build logs so you skip the wrong half. PolyForm noncommercial.
Keeping a 128 GB unified-memory APU fleet from eating itself: the KFD restore-worker thrash case study (51-minute model loads that should take 71 seconds), cgroup v2 budget architecture, and triage procedures for AMD Strix Halo. PolyForm noncommercial.
Production ROCm llama.cpp build recipe for AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — plus seven-model serving runbooks, KV-cache sizing against a unified-memory budget, sanitized systemd units, and a pitfalls doc of the failures that cost real days. PolyForm noncommercial.
Docker stack: Ollama v0.21.0 built from source against ROCm 7.2.2 with native gfx1151 (Strix Halo) — serves Gemma 4 up to 256K context on AMD Ryzen AI MAX+ 395 / Radeon 8060S. Includes a 9-layer make validate ladder for the host firmware, ROCm runtime, container, and long-context inference.
Three-surface (NPU + CPU + iGPU) concurrency architecture for the AMD Strix Halo APU: Lemonade/FLM on the XDNA 2 NPU (36-39 tok/s), a class-based request router, and the honest 39.1% accuracy-gate failure that keeps small-NPU routing shadow-only. PolyForm noncommercial.