ToolSlack schedules agent memory and exact-prefix KV preparation during real tool execution. The anonymous artifact includes the dense Qwen3-8B / LangGraph / LangMem implementation, CPU correctness checks, and a native GPU benchmark workflow.
Code and reproduction instructions
git clone https://github.com/ToolSlack/code.git
cd code
./run.sh cpuThe CPU command bootstraps the pinned environment and runs correctness checks on Linux or macOS. Python 3.9 or newer is required to start it; no API key, model weights, CUDA installation, or GPU allocation is needed.
For the native GPU workflow, first read the
hardware and reproduction prerequisites.
Use Linux x86_64 with one independently reserved, empty GPU, and replace 0 with
its physical nvidia-smi index:
./run.sh smoke --gpu-indices 0 --exclusive-gpus
./run.sh full --gpu-indices 0 --exclusive-gpusCPU validation does not establish GPU performance. Fresh native-engine installation, GPU smoke, native DMA correctness, and paper performance reproduction remain unverified; see the validation status and scope.
