A two-day, non-expert guide to running local LLMs on an AMD Strix Halo (Ryzen AI Max+ 395) box — full ~120GB memory pool, ROCm backend, NPU in parallel via FastFlowLM, one OpenAI-compatible endpoint.
-
Updated
Jul 25, 2026
A two-day, non-expert guide to running local LLMs on an AMD Strix Halo (Ryzen AI Max+ 395) box — full ~120GB memory pool, ROCm backend, NPU in parallel via FastFlowLM, one OpenAI-compatible endpoint.
Flm Companion is a lightweight Windows desktop app that simplifies managing your local FastFlowLM CLI server. It offers an intuitive interface to start, stop, and monitor the LLM, manage models, and check updates without command line use.
Windows desktop tool for local-LLM hotkeys, chat, grammar fixes, and note capture.
A converter for transferring gguf Q8_0, Q4_0, Q4_1, MXFP4, or raw tensors to FLM Q4NX
Solving high school math questions with local LLMs
On-device semantic search & RAG over any folder — running on your NPU (AMD Ryzen AI / Intel). Private, local-first, low-power. The memory layer of the Antumbra platform.
Unlocking the power of the Halo NPU: NPU + iGPU live verification, compression, and routing research on AMD Strix Halo
Add a description, image, and links to the fastflowlm topic page so that developers can more easily learn about it.
To associate your repository with the fastflowlm topic, visit your repo's landing page and select "manage topics."