An independent, evidence-graded read on how much you own the open models and inference providers you rely on.
-
Updated
Aug 11, 2026 - Astro
An independent, evidence-graded read on how much you own the open models and inference providers you rely on.
A systems-security framework for evaluating trust in open-weight LLM deployments — air-gapped environments, hidden backdoors, and software supply chain integrity.
Solving the amnesiac problem for LLM agents. Research series on agents that compound knowledge across sessions — first measurement: +4.6 pp accuracy lift on Terminal-Bench 2.1 with an open-weight executor and a single failure-derived skill file.
Preregistered cross-lingual evidence-sufficiency probing in Qwen3 — 6 languages, 29,206 paired QA constructions, frozen English probes, and a sealed one-shot evaluation.
Add a description, image, and links to the open-weight-llm topic page so that developers can more easily learn about it.
To associate your repository with the open-weight-llm topic, visit your repo's landing page and select "manage topics."