-
Notifications
You must be signed in to change notification settings - Fork 1.8k
All issues
Issue creation is restricted in this repository
- #19 · ZacharyZcR opened
on Jul 10, 2026 19 - #537 · ZacharyZcR opened
on Jul 22, 2026 4
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#586 In JustVugg/colibri;
- Status: Open.#585 In JustVugg/colibri;
[Performance]: GLM-5.2 (744B MoE) on a single AMD R9700 32 GB: 0.15 tok/s, 80% expert hit — capacity-bound benchmark
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwareperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#583 In JustVugg/colibri;[Analysis] The DRAM bandwidth wall: decode performance ceiling for streaming MoE, measured paths forward
performanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#537 In JustVugg/colibri;[Performance]: Vulkan (#418) vs ROCm/HIP on RDNA4 (RX 9070 XT) — CORRECTED: Vulkan is 19–24% faster (original comparison was confounded)
hardware-owner-neededServe verifica su silicio specificoServe verifica su silicio specificovulkanBackend Vulkan/AMDBackend Vulkan/AMDStatus: Open.#523 In JustVugg/colibri;RDNA2 (gfx1030) field report: AMD backend works; greedy decode not token-stable across mixed CPU/GPU expert tiers
hardware-owner-neededServe verifica su silicio specificoServe verifica su silicio specificovulkanBackend Vulkan/AMDBackend Vulkan/AMDStatus: Open.#510 In JustVugg/colibri;[Feature]: external draft-model hook (DeepSpec/DSpark-style sequential drafts) to raise acceptance on disk-bound hosts
featureNuova funzionalitàNuova funzionalitàStatus: Open.#494 In JustVugg/colibri;Does anyone want to collaborate to train a model that would be efficient specifically with such an engine?
discussionProposta / discussione aperta, non un taskProposta / discussione aperta, non un taskStatus: Open.#493 In JustVugg/colibri;Feature request: speculative decoding (MTP) in the HTTP
coli servepath for single-slot (--kv-slots 1)featureNuova funzionalitàNuova funzionalitàStatus: Open.#492 In JustVugg/colibri;Benchmark datapoint: GLM-5.2 744B on Ryzen AI Max+ 395 (Strix Halo, 128 GB, Ubuntu) — ~1.0 tok/s CPU-only
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwareStatus: Open.#486 In JustVugg/colibri;[Bug]: glm_tiny TF oracle is 30/32, not the documented 32/32 — forward-pass near-tie at 2 positions
bugDifetto verificato nel codiceDifetto verificato nel codiceStatus: Open.#482 In JustVugg/colibri;Datapoints: GLM-5.2 int4 fully RAM-resident on 2-socket Ice Lake (48C, 661 GiB) — 3.85 tok/s; DRAFT=0 beats MTP at full residency
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwareStatus: Open.#472 In JustVugg/colibri;