Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
-
Updated
Sep 2, 2026 - Python
Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.
Efficient LLM inference on Slurm clusters.
Phi-Bench (Φ-Bench): 85 open-source LLM-infrastructure engineering tasks for frontier LLMs & coding agents (KFC/LH/E2E). Self-contained public Dockerfile + offline scoring + reference solution per task. Website: llminfrabench.com
PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Multi-model AI agent runtime. Define agents in YAML, route each role to a model, orchestrate with 7 patterns (ReAct, Plan & Execute, Fan-Out, Pipeline, Supervisor, Swarm, Glyph), and deploy as a REST/WebSocket API with RAG, memory, MCP tools, guardrails and OpenTelemetry observability.
AI Workload Control Layer for routing deterministic, reusable, retrieval-needed, tool-needed, and provider-needed work before model invocation.
A browser-based UI for launching, monitoring, clustering and managing multiple llama.cpp server instances from inside a Docker container. Includes an Ollama-compatible API proxy
Reliability control for the Anthropic Python SDK
Persistent Cognitive Memory Infrastructure — durable, multi-tenant memory layer for AI agents. HTTP + gRPC, hybrid retrieval, background workers, observability. Go.
A production-grade, schema-aware PostgreSQL MCP server for enterprise AI. Features Zero-Trust SQL validation, multi-tier permissions, and real-time schema introspection for secure, autonomous database operations.
nvProbe — Open-source NVIDIA GPU benchmark suite for CUDA workload automation, Slurm HPC cluster profiling, and MLPerf reporting
Krako 2.0 – Energy-efficient, triadic multi-tier inference infrastructure enabling adaptive routing across heterogeneous edge–cloud nodes.
A humble exploratory PoC for a hardware-native, optical timing-frozen control plane engine. Fuses inline CUDA PTX, RAII memory tunnels, and 4D JAX shard_map structures to investigate 0ns-overhead fault-tolerant routing for hyperscale distributed AI.
One command. Full LLM stack. Zero config.
An intelligent gateway for Claude APIs that dynamically routes requests to the most cost-efficient model, caches responses, and escalates based on confidence signals — reducing LLM spend without sacrificing quality.
Forward-Only Homeostasis Accelerator Kernel for 0ns Distributed Overlapping & Static O(1) VRAM Memory Wall Liquidation via JAX/XLA & PyTorch.
High-performance Triton kernels for NVIDIA H100. Implements fused FP8 LayerNorm, tiled FlashAttention, and SRAM-optimized memory primitives for Hopper architecture.
Fail-fast admission control for LLM workloads, with concurrency limits, token budgeting, deduplication, streaming support, and overload protection.
To associate your repository with the llm-infrastructure topic, visit your repo's landing page and select "manage topics."