| 2022-09 |
carperai/openelm |
进化式 Prompt 优化 |
进化/搜索循环 → 评估器/打分器 |
github_api |
| 2023-01 |
stanfordnlp/dspy |
声明式 Prompt 优化 |
反馈-精炼 → 进化/搜索循环 → 评估器/打分器 |
github_api |
| 2023-03 |
Significant-Gravitas/AutoGPT |
自主 Agent 平台 |
智能体编排 |
github_api |
| 2023-03 |
camel-ai/camel |
角色扮演 Agent 框架 |
智能体编排 → 反馈-精炼 |
github_api |
| 2023-03 |
noahshinn/reflexion |
反思记忆 |
进化/搜索循环 → 反思记忆 → 反馈-精炼 → 评估器/打分器 → 训练/数据循环 |
github_api |
| 2023-03 |
madaan/self-refine |
反馈精炼 |
反馈-精炼 |
github_api |
| 2023-04 |
microsoft/CoML |
ML 知识库驱动 |
反馈-精炼 → 评估器/打分器 |
github_api |
| 2023-06 |
FoundationAgents/MetaGPT |
多 Agent 协作框架 |
智能体编排 → 反馈-精炼 |
github_api |
| 2023-07 |
aiwaves-cn/agents |
数据驱动 Agent 进化 |
进化/搜索循环 → 评估器/打分器 → 智能体编排 |
github_api |
| 2023-08 |
langchain-ai/langgraph |
图式 Agent 编排 |
智能体编排 |
github_api |
| 2023-08 |
microsoft/autogen |
多 Agent 对话框架 |
智能体编排 |
github_api |
| 2023-10 |
google-deepmind/opro |
LLM 作为优化器 |
进化/搜索循环 → 评估器/打分器 |
github_api |
| 2023-10 |
crewAIInc/crewAI |
多 Agent 协作框架 |
智能体编排 |
github_api |
| 2023-11 |
google-deepmind/funsearch |
进化式数学发现 |
进化/搜索循环 → 评估器/打分器 |
github_api |
| 2024-03 |
All-Hands-AI/OpenHands |
AI 软件开发平台 |
智能体编排 |
github_api |
| 2024-04 |
princeton-nlp/SWE-agent |
软件工程 Agent |
反馈-精炼 → 评估器/打分器 |
github_api |
| 2024-07 |
shengranhu/adas |
Agent 架构自动搜索 |
进化/搜索循环 → 智能体编排 → 评估器/打分器 |
github_api |
| 2025-05 |
DeepAuto-AI/automl-agent |
多 Agent AutoML |
智能体编排 → 进化/搜索循环 → 评估器/打分器 |
github_api |
| 2025-05 |
algorithmicsuperintelligence/openevolve |
进化式代码优化 |
进化/搜索循环 → 评估器/打分器 |
github_api |
| 2025-07 |
JARVIS-Xs/SE-Agent |
代码智能体自进化 |
进化/搜索循环 → 评估器/打分器 → 智能体编排 |
github_api |
| 2025-10 |
inter-co/science-codeevolve |
科学代码进化 |
进化/搜索循环 → 评估器/打分器 |
github_api |
| 2025-11 |
modelscope/AgentEvolver |
Agent 进化框架 |
进化/搜索循环 → 评估器/打分器 → 智能体编排 → 训练/数据循环 |
github_api |
| 2025-12 |
JarvisPei/SCOPE |
上下文/Prompt 进化 |
进化/搜索循环 |
github_api |
| 2026-03 |
OPPO-Mente-Lab/LLM-Self-Judge |
自评判训练 |
进化/搜索循环 → 评估器/打分器 → 智能体编排 → 训练/数据循环 |
github_api |
| 2026-04 |
ZJU-LLM-Safety/DARWIN |
安全策略进化 |
进化/搜索循环 → 反思记忆 |
github_api |
| unknown |
0xNyk/lacp |
Agent Context Protocol and Interoperability Tooling |
define lightweight context exchange boundaries for agent workflows -> normalize session and memory payload handoff across tools -> reduce protocol mismatch cost in multi-agent chains -> preserve context continuity while keeping transport overhead low |
unknown |
| unknown |
aayoawoyemi/ori-mnemos |
Agent Memory Substrate and Runtime Tracing Harness |
capture agent interactions and prompts into typed memory records -> run compression and retrieval policies over prior episodes -> inject relevant traces back into current execution context -> expose memory operations as runtime primitives for reproducible long-horizon behavior |
unknown |
| unknown |
addyosmani/agent-skills |
Production Engineering Skill Pack for Coding Agents |
encode senior engineering workflows as reusable agent skills -> map tasks to explicit process gates -> force verification evidence before ship -> compound quality through consistent command-level behavior |
unknown |
| unknown |
ag2ai/ag2 |
多 Agent 协作框架 |
多 Agent 对话 → 编排 → 协作 |
unknown |
| unknown |
agent0ai/agent-zero |
Autonomous Agent Runtime |
project workspace -> Linux/tools/browser/memory/skills -> inspectable agent work -> reusable state |
unknown |
| unknown |
Agenta-AI/agenta |
LLM 评测平台 |
Prompt 管理 → 测试集 → 评估器 → 可观测性 |
unknown |
| unknown |
agentevals-dev/agentevals |
Benchmark and Evaluation Framework for Agent Systems |
standardize evaluation tasks for autonomous agents -> execute benchmark suites across versions and prompts -> compare outcomes using shared scoring protocols -> keep performance claims auditable across continuous updates |
unknown |
| unknown |
agentmemoryworld/awesome-agent-memory |
Agent Memory Resource Survey Index |
collect memory papers and systems -> organize them by mechanism and scope -> point readers to benchmark and implementation anchors -> keep the memory landscape navigable as a survey resource |
unknown |
| unknown |
agentralabs/agentic-memory |
Agentic Memory Runtime Framework for Persistent Context |
model memory as a first-class runtime subsystem -> persist and recall long-horizon context artifacts -> expose deterministic memory hooks to agent loops -> reduce context loss across iterative autonomous tasks |
unknown |
| unknown |
agentscope-ai/ReMe |
Long-Term Agent Memory and Context Compression Framework |
compress long context into structured summaries -> persist long-term memory in file/vector backends -> recall relevant memory with hybrid retrieval -> reuse memory traces across sessions and benchmark loops |
unknown |
| unknown |
agiprolabs/claude-trading-skills |
Domain Agent Skill Workflow Pack |
domain skill pack -> market data / backtesting / risk / tax workflows -> reusable agent task procedures |
unknown |
| unknown |
agiresearch/A-mem |
Agentic Memory Architecture for LLM Agent Long-Term Context Retention |
build autonomous memory lifecycle for LLM agents -> store and retrieve long-horizon context with salience control -> update memory store through usage feedback -> improve continuity and task grounding across iterative agent runs |
unknown |
| unknown |
ai-agents-2030/darwin-mobile-agent |
Mobile Agent Self-Evolution Framework |
run mobile task execution loops -> record failures and intervention traces -> evolve prompts/skills/action plans -> replay against app tasks to measure iterative gains |
unknown |
| unknown |
ai4co/awesome-fm4co |
基础模型+组合优化综述 |
文献综述 |
unknown |
| unknown |
ai4co/reevo |
反射式进化搜索 |
进化/搜索循环 → 反思记忆 → 评估器/打分器 |
unknown |
| unknown |
ai4co/rl4co |
RL 组合优化基准 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
aiming-lab/AutoHarness |
Automated Agent Harness Engineering Framework |
treat harness as the controllable layer around model reasoning -> enforce multi-step governance and risk checks on tool execution -> track costs, logs, and sessions -> feed failures back into harness policies to improve future agent runs |
unknown |
| unknown |
aiming-lab/ClawArena |
Interactive Computer-Use Benchmark Harness Arena |
run interactive browser/desktop tasks in benchmark arenas -> score agent behavior across controlled environments -> compare policy and harness variants with reproducible evaluation traces -> feed benchmark deltas back into harness and skill updates |
unknown |
| unknown |
alfa-group/tutorial_gp_llm |
GP+LLM 教学 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
alibaizhanov/mengram |
Semantic/Episodic/Procedural Memory Runtime for Agents |
model semantic, episodic, and procedural memories as explicit agent assets -> learn procedures from failures and feedback traces -> integrate memory services into LangChain/CrewAI/OpenClaw flows -> improve adaptation quality through structured memory retention |
unknown |
| unknown |
Alienfader/continuity-benchmarks |
Execution-Intent Memory Benchmark Harness |
agent action intent -> retrieval keyed by execution intent vs prompt intent -> benchmark runners score recall/alignment -> report deltas and confidence for memory strategy selection |
unknown |
| unknown |
ALucek/agentic-memory |
Memory Methods Library for Cognitive Agent Architectures |
translate cognitive-memory concepts into implementation templates -> organize memory techniques by operational use-case -> provide runnable method patterns for agent builders -> improve practical memory design choices through comparative examples |
unknown |
| unknown |
apify/agent-skills |
Reusable Skills Library for Coding Agents and Automation Workflows |
package reusable skill prompts and procedures for coding agents -> map repeated engineering tasks into skill modules -> compose skill units into longer autonomous workflows -> improve reliability and transferability of agent execution behavior |
unknown |
| unknown |
AQ-MedAI/MedMemoryBench |
Personalized Healthcare Agent Memory Benchmark |
construct longitudinal healthcare episodes -> require agents to recall and apply patient-specific context -> score temporal memory quality and downstream task success -> expose where memory retrieval helps or harms clinical reasoning |
unknown |
| unknown |
Arc-Computer/CL-Bench |
Stateful Continual-Learning Benchmark for LLM Agents |
place agents inside stateful multi-turn workflows -> mutate persistent entities under production-style constraints -> evaluate adaptation and reliability under cross-turn dependencies -> use continual-learning pressure instead of one-shot benchmark snapshots |
unknown |
| unknown |
ArcadeAI/openclaw-arcade-plugin |
OpenClaw Skill Plugin for Arcade Tool Connectivity |
bridge OpenClaw runtime calls to Arcade.dev tool endpoints -> register reusable tool skills through plugin contracts -> execute constrained external actions with typed interfaces -> feed tool outcomes back into agent planning loops for iterative skill reuse |
unknown |
| unknown |
arthurmgraf/graphmind |
Knowledge-Graph Agentic RAG Runtime |
receive query -> choose LangGraph or CrewAI engine -> retrieve over hybrid graph layer -> self-evaluate the answer -> retry when score stays below threshold |
unknown |
| unknown |
AutoJunjie/awesome-agent-harness |
Harness Curation and Reading Map |
collect harness repositories and papers -> classify them by lifecycle, runtime, memory, protocols, and workflows -> provide a quick browse path for reproducibility and safety trends |
unknown |
| unknown |
automl/auto-sklearn |
AutoML 框架 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
AutoX-AI-Labs/AutoR |
Human-Centered Research Harness |
human research intent -> staged agent execution -> approval checkpoints -> artifact-backed run directory -> resume/redo/rollback |
unknown |
| unknown |
axiomhq/agent-memory |
Persistent Agent Memory Runtime |
capture user/agent interaction state -> extract and store memory artifacts in redis-backed structures -> retrieve context through memory APIs -> feed subsequent agent decisions and orchestration flows |
unknown |
| unknown |
back1ply/agent-skill-loader |
Runtime Agent Skill Loader |
ingest skill bundles with uniform loader interfaces -> resolve runtime dependencies and skill metadata -> mount skills into agent execution contexts -> support iteration through modular updates |
unknown |
| unknown |
beeevita/EvoPrompt |
进化式 Prompt 优化 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
benchflow-ai/skillsbench |
Agent Skills Benchmark Harness |
define skill-centric tasks -> run agent plus skill compositions -> score verifier outputs and task success -> compare per-task and per-skill behavior across models and runtimes |
unknown |
| unknown |
BerriAI/litellm |
LLM 基础设施 |
统一接口 → 100+ LLM → 代理网关 |
unknown |
| unknown |
BerriAI/self-improving-agent |
Self-Improving Coding Agent Loop |
agent executes coding tasks -> evaluates outcomes via benchmarks and user feedback -> writes workflow/self changes -> reruns tasks to measure iterative improvement |
unknown |
| unknown |
block/agent-skills |
Enterprise Agent Skills and Playbook Library |
standardize reusable skills as enterprise playbooks -> encode engineering policy and risk checks into skill procedures -> distribute skills across codex/claude workflows -> reduce variance and onboarding cost in large agent teams |
unknown |
| unknown |
BlockRunAI/awesome-OpenClaw-Money-Maker |
Agent Monetization Workflow and OpenClaw Use-Case Index |
catalog repeatable OpenClaw monetization workflows and supporting tools -> connect skills, automations, and infrastructure choices to concrete earning scenarios -> expose cost and risk controls for agent deployment decisions -> provide a practical lens for product-level agent usability |
unknown |
| unknown |
browser-use/browser-harness |
Self-Healing Browser Agent Harness |
attach one websocket to Chrome -> let the agent call or write browser helpers -> execute repeatable browser tasks -> keep the harness editable so the next run can reuse stronger helpers |
unknown |
| unknown |
Chainlit/chainlit |
LLM 聊天框架 |
LLM 聊天 UI → 快速构建 → 部署 |
unknown |
| unknown |
CharlesQ9/Self-Evolving-Agents |
自进化 Agent 综述 |
文献综述 |
unknown |
| unknown |
cheshire-cat-ai/core |
AI 聊天框架 |
插件式 AI → 模块化 → 可扩展 |
unknown |
| unknown |
Chorus-AIDLC/Chorus |
AI-Human Collaboration Harness |
requirements/task state -> sub-agent orchestration -> permissions/context injection -> observability/failure recovery -> OpenSpec archival |
unknown |
| unknown |
christinminor459/OnionClaw |
Security/Privacy Agent Plugin with Tooling and Channel Hardening |
wrap agent actions with OPSEC constraints and privacy-first tool defaults -> provide hardened plugins for network, command, and context operations -> keep operational traces bounded while still enabling autonomous workflows -> improve safe deployment readiness for agent teams |
unknown |
| unknown |
clawland-ai/geneclaw |
Safe Self-Evolving Agent Framework |
observe failures -> diagnose root causes -> propose constrained diffs -> validate through five safety gates -> branch, test, and apply only after approval or configured autopilot |
unknown |
| unknown |
clawsouls/soulclaw |
OpenClaw Fork with Multi-Tier Memory and Persona Runtime |
extend OpenClaw with persistent identity and multi-tier memory boundaries -> separate immutable soul identity from working and session memories -> synchronize persona and memory policies across channels and teams -> reduce drift while enabling long-horizon behavior consistency |
unknown |
| unknown |
cloudllm-ai/mentisdb |
Durable Agent Memory Graph Database and Skill Registry Runtime |
store append-only semantic memory and thought chains in a durable graph substrate -> version skills as integrity-checked artifacts similar to a registry -> retrieve and merge high-signal historical context into active agent decisions -> preserve learning continuity across sessions, models, and team handoffs |
unknown |
| unknown |
CMA-ES/pycma |
经典进化策略 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
CodeFuse-ML/awesome-code-llm |
代码 LLM 综述 |
文献综述 |
unknown |
| unknown |
coleam00/Archon |
Deterministic AI Coding Harness Builder |
workflow-defined coding pipeline -> deterministic phases and validation gates -> isolated worktree execution -> artifacted review/PR generation with mixed deterministic and AI nodes |
unknown |
| unknown |
composio-community/awesome-openclaw-plugins |
OpenClaw Plugin Ecosystem Index and Skill Resource Map |
curate installable OpenClaw plugins into structured categories -> map memory, security, observability, and multi-agent capabilities in one index -> provide concrete install commands as reusable operational skills -> accelerate capability discovery for agent workflow evolution |
unknown |
| unknown |
ComposioHQ/agent-orchestrator |
Production Coding-Agent Swarm Orchestrator |
route coding tasks into specialized agents -> isolate changes in Git worktrees -> reuse skills and memory across execution steps -> coordinate MCP/tool calls and review gates -> merge accepted work back into the main engineering flow |
unknown |
| unknown |
ComposioHQ/awesome-agent-clis |
Agent CLI Orchestration Resource Index |
aggregate production-ready agent CLIs -> expose setup/docs/ecosystem compatibility -> guide teams to reusable command-line workflows for coding and operations agents |
unknown |
| unknown |
Corbell-AI/evalmonkey |
Agent Evaluation Harness and Regression Pipeline |
define task-level agent evaluation suites -> run LLM-based and deterministic regression checks -> aggregate quality metrics into repeatable reports -> feed benchmark regressions back into skill/harness improvement loops |
unknown |
| unknown |
CortexReach/memory-lancedb-pro |
OpenClaw Long-Term Memory Plugin |
auto-capture memory -> vector+BM25 retrieval -> rerank/context injection -> scoped memory boundaries -> CLI backup and migration |
unknown |
| unknown |
cuga-project/cuga-agent |
Enterprise Generalist Agent Harness |
enterprise agent config -> tools/MCP/OpenAPI -> policies/HITL -> optional memory/knowledge/skills -> trajectory visualization |
unknown |
| unknown |
cxxz/awesome-agent-memory |
Agent Memory Resource Index |
memory systems -> tools/patterns/research -> agent memory taxonomy |
unknown |
| unknown |
Da1yuqin/SEAD |
Self-Evolving Agent Design Benchmark |
benchmark architecture-level agent design quality -> compare model-generated system designs under controlled tasks -> score design quality and completion behavior -> reveal where self-evolving design loops fail |
unknown |
| unknown |
dataelement/bisheng |
LLM 应用平台 |
LLM 应用平台 → 可视化编排 → 知识库 |
unknown |
| unknown |
datalayer/agent-skills |
Composable Agent Skills Pack and Runtime Recipes |
collect reusable skills as installable packs -> map skills to real workflows and runtime contexts -> version skill definitions to preserve reproducibility -> compose skills into controllable agent workflows with lower setup cost |
unknown |
| unknown |
Dataojitori/nocturne_memory |
Context-Aware Long-Term Memory Engine for AI Agents |
capture and rank interaction context for long-term retention -> retrieve semantically relevant memory at inference time -> reinforce memory quality through ongoing usage feedback -> improve continuity and personalization of autonomous agent behavior |
unknown |
| unknown |
dceoy/speckit-agent-skills |
Spec-Driven Agent Workflow Skills |
spec-driven workflow -> constitution/specify/plan/tasks/implement skills -> multi-runtime agent process discipline |
unknown |
| unknown |
DEAP/deap |
经典进化算法框架 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
DEEP-PolyU/Awesome-GraphMemory |
Graph-Based Agent Memory Index |
graph memory papers -> techniques/applications -> memory substrate map |
unknown |
| unknown |
desplega-ai/agent-swarm |
Compounding Lead-Worker Agent Runtime |
ingest tasks from external channels -> lead agent plans and delegates -> workers run inside isolated Docker environments -> shared memory and identity accumulate across sessions -> pages, PRs, replies, and scheduled workflows turn learnings into reusable operations |
unknown |
| unknown |
dotnet/skills |
Cross-IDE .NET Agent Skills Runtime Pack |
publish reusable coding-agent skill packs for multiple runtimes -> provide strict .NET engineering workflows and reusable prompts -> score and compare skill quality with benchmark harness integration -> transfer high-quality skill behavior across sessions and tools |
unknown |
| unknown |
e2b-dev/e2b |
代码执行沙箱 |
AI 代码 → 安全沙箱 → 隔离执行 |
unknown |
| unknown |
EESIZ/clawdreamer |
OpenClaw Automation App and Productivity Workflow Plugin |
package practical automation actions as OpenClaw-callable capabilities -> bind daily productivity workflows to reusable agent commands -> persist operational context between invocations -> turn repeated manual steps into compounding assistant skills |
unknown |
| unknown |
elizaOS/agentmemory |
Agent Memory Plugin for ElizaOS Runtime and Persistent Context Handling |
attach memory plugin into ElizaOS runtime pipeline -> persist memory records and expose retrieval hooks to agents -> apply configurable memory operations per interaction -> provide reusable memory module boundary for agent ecosystems |
unknown |
| unknown |
elizaOS/eliza |
Autonomous Agent Framework |
autonomous-agent framework -> plugins/CLI/web lifecycle -> deployed agent applications |
unknown |
| unknown |
evalops/agent-harness |
Cross-Provider Agent Harness Adapter |
register tools once -> normalize json schema and response shape -> lazy-load provider adapters -> run the same task across multiple agent backends for comparison |
unknown |
| unknown |
EverMind-AI/EverOS |
自进化 Agent 记忆系统 |
反思记忆 → 智能体编排 |
unknown |
| unknown |
EvoAgentX/EvoAgentX |
自进化 Agent 生态系统 |
进化/搜索循环 → 智能体编排 → 评估器/打分器 → 反馈-精炼 |
unknown |
| unknown |
EvoMap/awesome-agent-evolution |
External Awesome List and Taxonomy Comparator |
field taxonomy -> curated project/paper/benchmark sections -> related awesome-list pointers -> reader-facing ecosystem navigation |
unknown |
| unknown |
EvoMap/evolver |
Self-Evolving Memory and Reasoning Map Framework |
represent knowledge as evolving graph maps -> apply reinforcement and relation updates from interaction feedback -> optimize memory retrieval paths for downstream reasoning -> persist the updated map structure as a reusable long-horizon state substrate |
unknown |
| unknown |
facebookresearch/nevergrad |
无梯度优化框架 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
FeiLiu36/LLM4Opt |
LLM 驱动算法设计综述 |
文献综述 |
unknown |
| unknown |
first-fluke/oh-my-agent |
Open Multi-Agent Runtime and Benchmark Harness |
build agent workflows through composable nodes -> run memory, planner, and tool execution as auditable modules -> benchmark agents with repeatable evaluation entries -> iterate runtime policies using harness feedback and workflow traces |
unknown |
| unknown |
flagos-ai/skills |
Open Agent Skill Registry |
skill package spec -> registry publishing -> install hooks -> versioning -> cross-agent reuse |
unknown |
| unknown |
FlowiseAI/Flowise |
可视化 LLM 平台 |
拖拽 UI → LLM 链 → 可视化编排 |
unknown |
| unknown |
FreedomIntelligence/Tiermem |
Provenance-Aware Memory Benchmark Framework |
construct knowledge-memory tasks with provenance labels -> run language-agent memory retrieval and generation pipelines -> score both answer quality and citation provenance -> compare memory frameworks under standardized settings |
unknown |
| unknown |
future-agi/future-agi |
自改进 Agent |
自改进循环 → 评估 → 迭代优化 |
unknown |
| unknown |
garrytan/gbrain |
Agent Company Brain and Memory OS |
ingest multi-source signals -> synthesize and link entities -> persist memory graph -> query/retrieve for next actions -> recurring maintenance jobs |
unknown |
| unknown |
GCWing/BitFun |
Desktop Agent Runtime and Multi-Mode Execution Environment |
provide desktop-native agent runtime with code, cowork, and computer-use modalities -> preserve memory and personality state across sessions -> support long-running service mode for continuous operation -> compound capabilities through repeated task execution and context retention |
unknown |
| unknown |
Genentech/OpenTreeSearch |
LLM 引导代码进化 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
gofenix/nex-agent |
Elixir/OTP Self-Evolving Agent Runtime |
supervised runtime -> memory/tools/skills -> subagents/jobs -> source-level upgrades |
unknown |
| unknown |
google-gemini/gemini-cli |
Agent CLI Auto-Memory and Skills |
session transcripts -> auto-memory mining -> reviewable patches / SKILL.md drafts -> approved durable memory or skill assets |
unknown |
| unknown |
google/ax |
Production Agent Runtime and Context Engineering Framework |
build modular agent execution runtime primitives -> encode context engineering into reusable runtime components -> integrate evaluation and tracing for production reliability -> iterate runtime behavior using measured workflow outcomes |
unknown |
| unknown |
harness/harness-evals |
Agent Reliability Evaluation Framework |
evaluate cases with normalized 0-1 scores -> configurable pass thresholds -> optional llm judged metrics and telemetry sinks -> regression export to CI observability pipelines |
unknown |
| unknown |
henrikrexed/openclaw-observability-plugin |
Agent Runtime Observability and Trace Monitoring Plugin |
instrument OpenClaw runtime events into observable traces -> route telemetry into inspection dashboards and logs -> correlate agent actions with execution outcomes -> reduce blind spots in harness debugging and regression analysis |
unknown |
| unknown |
holaboss-ai/holaOS |
Long-Horizon Agent Environment |
agent environment as execution substrate -> continuity-oriented context and memory management -> MCP-compatible tooling for long-horizon work -> self-evolving workflow emphasis through environment-level adaptation |
unknown |
| unknown |
howdymary/hermes-agent-metaharness |
Hermes Benchmark Outer-Loop Harness |
select candidate -> evaluate on TBLite/TB2 -> parse archives -> compare baseline vs candidate -> update frontier |
unknown |
| unknown |
huggingface/smolagents |
Agent 框架 |
轻量 Agent → 工具调用 → HuggingFace 集成 |
unknown |
| unknown |
humanitylabs-org/obsidianclaw |
OpenClaw Knowledge and Notes Integration Plugin |
connect OpenClaw sessions with Obsidian-backed knowledge notes -> store interaction artifacts as structured vault entries -> retrieve linked notes into future planning/context windows -> preserve durable memory across episodic agent runs |
unknown |
| unknown |
hwfengcs/dm-code-agent |
Auditable Local-First Code Agent Baseline |
plan and replan -> call tools with JSONL trace capture -> enable optional reflexion or critic modules -> run maintenance benchmark harness -> replay and diff the trajectory |
unknown |
| unknown |
hyperspell/hyperspell-openclaw |
OpenClaw Memory and Context Enhancement Runtime |
capture workspace memory artifacts from OpenClaw sessions -> sync and normalize context into an external memory substrate -> retrieve high-signal notes into follow-up prompts and tool calls -> reinforce long-horizon consistency across evolving agent tasks |
unknown |
| unknown |
icaros-usc/pyribs |
质量多样性优化 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
iflytek/skillhub |
Agent Skill Registry and Open Runtime Platform |
structured skill package definition -> runtime orchestration and multi-agent routing -> deployment and plugin integration -> reusable skill asset lifecycle management |
unknown |
| unknown |
iliaal/ai-skills |
Agent Process Skill Library |
portable agent skills -> planning, debugging, review and verification discipline -> reusable behavior layer |
unknown |
| unknown |
im4codes/imcodes |
Shared Agent Context, Memory, and Supervised Execution Layer |
establish shared context and memory channels across agent providers -> supervise execution and record cross-agent audit trails -> standardize communication primitives for multi-agent collaboration -> reduce fragmentation and improve reproducibility in mixed-agent systems |
unknown |
| unknown |
InfiAgent/InfiAgent |
Framework for Self-Improving Agent Loops |
break complex goals into planner-executor-reflection stages -> execute tasks with tool-use traces -> distill successful trajectories into reusable policies -> iterate to improve completion quality over time |
unknown |
| unknown |
inngest/utah |
Event-Driven Agent Harness Runtime |
incoming event -> think/act/observe loop -> durable retries and singleton control -> memory/session trace updates -> channel response |
unknown |
| unknown |
InternLM/WildClawBench |
Agent 评测基准 |
真实场景任务 -> 多轮动态交互 -> anti-overfitting 设计 -> 端到端评分 -> agent 能力画像 |
unknown |
| unknown |
InternScience/Awesome-Scientific-Skills |
Scientific Agent Skill and Tooling Index |
collect scientific task skills and tools -> map reusable research procedures -> connect benchmarks and methodology references -> support skill transfer into research agents |
unknown |
| unknown |
InternScience/InternAgent |
Autonomous Scientific Discovery Agent Framework |
idea generation -> method construction -> experiment planning and execution -> benchmark evaluation -> memory-informed next iteration |
unknown |
| unknown |
itgoyo/awesome-agent-skills |
Cross-Platform Agent Skills Resource Index |
curate agent-skill resources across runtimes -> map official and community skill ecosystems -> link install/readme paths -> help teams bootstrap reusable skill workflows quickly |
unknown |
| unknown |
junminhong/awesome-agent-skills |
Cross-Platform Agent Skill Index |
curate platform-specific skills -> define skill frontmatter and folder template -> enumerate design patterns and evaluation checklists -> map official docs for reusable implementation |
unknown |
| unknown |
Kenotic-Labs/ATANT |
Agent Continuity Evaluation |
agent narrative checkpoints -> continuity tests -> self/identity drift evidence |
unknown |
| unknown |
knowall-ai/mcp-neo4j-agent-memory |
Graph-Memory MCP Server for Long-Horizon Agents |
capture chat events and tool outputs into a graph memory store -> expose semantic and structural retrieval through MCP endpoints -> keep temporal and entity relations queryable -> feed retrieved memories back into agent planning and execution loops |
unknown |
| unknown |
kodustech/awesome-agent-skills |
Curated Agent Skill Catalog and Prompt Workflow Patterns |
curate reusable coding-agent skills into an index with reproducible examples -> normalize prompt and workflow patterns across domains -> make skills discoverable and composable for harness integration -> accelerate agent improvement by reusing validated skill modules |
unknown |
| unknown |
kweaver-ai/kweaver-core |
Enterprise Decision Agent Harness |
business knowledge network -> governed context loader -> tool curation/path guidance -> decision agent execution -> TraceAI feedback evidence |
unknown |
| unknown |
kyegomez/swarms |
Production Multi-Agent Orchestration Runtime |
define agents and swarm topology -> run sequential/concurrent/hierarchical orchestration -> attach tools, memory, and protocol adapters -> keep workflow behavior inspectable through runtime boundaries and reusable swarm patterns |
unknown |
| unknown |
langchain-ai/agentevals |
Agent Evaluation Harness with LangGraph Integrations |
define executable eval datasets and scoring rules -> run deterministic and model-graded checks against agent trajectories -> surface regressions through repeatable evaluation loops -> harden agent releases with test-like quality gates |
unknown |
| unknown |
langchain-ai/deepagents |
Batteries-included Agent Harness Runtime |
opinionated harness runtime -> sub-agent delegation and filesystem actions -> persistent memory plus context management -> evaluation and deployment paths via LangGraph/LangSmith |
unknown |
| unknown |
langchain-ai/memory-agent |
Memory-Aware Agent Workflow and Evaluation App |
long-running conversation and user context -> memory extraction and consolidation via LangMem -> LangGraph workflow execution -> memory-grounded follow-up behavior and replayable traces |
unknown |
| unknown |
langflow-ai/langflow |
可视化 Agent 平台 |
拖拽可视化 → LangChain 组件 → Agent 编排 |
unknown |
| unknown |
langgenius/dify |
LLM 应用平台 |
可视化工作流 → LLM 编排 → 应用部署 |
unknown |
| unknown |
LearnPrompt/cc-harness-skills |
Codex/Claude Harness Skill Playbooks |
encode recurring coding-agent execution patterns as reusable markdown skills -> bind each playbook to harness-level commands and quality checks -> reuse these skills across sessions to reduce setup entropy -> iterate skill prompts based on failure and review feedback |
unknown |
| unknown |
letta-ai/learning-sdk |
Continual Learning And Long-Term Memory SDK |
wrap an existing LLM client -> capture conversation traces -> persist and inject relevant memory -> make the original agent stateful without retraining the base model |
unknown |
| unknown |
lhl/agentic-memory |
Pluggable Agentic Memory Module for Any Agent System |
package memory primitives as pluggable modules -> slot memory handling into existing agent systems -> retain interaction context across sessions -> provide a lightweight memory building block for broader agent stacks |
unknown |
| unknown |
LHL3341/awesome-claws |
OpenClaw Ecosystem Collection and Skill/Tool Index |
organize OpenClaw products, skills, communities, and ecosystem resources into scenario-based sections -> keep practical references and links continuously discoverable -> reduce search overhead for builders selecting tools -> accelerate workflow assembly for multi-agent and channel deployments |
unknown |
| unknown |
LLMSecurity/awesome-agent-skills-security |
Agent Skill Security Resource Index |
collect benchmark and attack references -> map skill-level vulnerabilities and mitigations -> publish curated defense pathways for agent-skill engineering teams |
unknown |
| unknown |
longmans/self-evolve |
Self-Evolving OpenClaw Workflow Playground and Benchmark Harness |
feedback detection -> reward scoring and learning gates -> Q-value updates plus episodic memory append -> local and remote retrieval on later turns |
unknown |
| unknown |
luo-junyu/Awesome-Agent-Papers |
Agent 研究综述 |
论文索引 → LLM Agent 研究追踪 |
unknown |
| unknown |
manthanguptaa/water |
Self-Improving Coding Agent with Benchmark-Oriented Execution |
run coding-agent trajectories under measurable evaluation loops -> compare strategy revisions against task outcomes and regression signals -> retain higher-performing behaviors while dropping failing patches -> improve coding reliability through iterative self-correction |
unknown |
| unknown |
Martian-Engineering/lossless-claw |
Persistent Context and Memory Orchestration for OpenClaw |
capture task and conversation traces into structured context artifacts -> rank and compress context for retrieval fidelity -> inject curated context into follow-up agent/tool calls -> keep a replayable context lineage that reduces drift across long-running workflows |
unknown |
| unknown |
matevip/mateclaw |
OpenClaw Runtime Extension with Memory Control and Automation Rules |
inject memory and rule-engine controls into OpenClaw execution -> bind prompts, tools, and context flows to policy checks -> automate repeatable task pipelines with persistent state -> reduce drift while compounding operator experience across runs |
unknown |
| unknown |
MaximeRobeyns/self_improving_coding_agent |
Self-Improving Coding Agent |
coding agent -> own-codebase modification -> tests/review signal -> improved next agent iteration |
unknown |
| unknown |
MCKRUZ/openclaw-langfuse |
OpenClaw Tracing Plugin for Langfuse Observability |
hook OpenClaw tool and model events into Langfuse traces -> attach session metadata for replayable diagnostics -> surface latency and failure patterns across agent runs -> support regression spotting before releasing updated agent workflows |
unknown |
| unknown |
mem9-ai/mem9 |
Persistent Memory Layer for Multi-Agent Runtimes |
memory write/search/get/update/delete API -> runtime plugins and skills -> cross-session recall -> shared multi-agent memory reuse |
unknown |
| unknown |
memodb-io/Acontext |
Agent Skill Memory Layer and Runtime Context Engine |
skill and behavior trace ingestion -> memory indexing and retrieval -> context-aware execution with long-term persistence -> memory-informed agent behavior adaptation |
unknown |
| unknown |
MemTensor/MemOS-Cloud-OpenClaw-Plugin |
Hosted Agent Memory Runtime Plugin |
intercept agent execution context before task start -> recall long-term memories from hosted MemOS service -> run tasks with enriched context -> persist post-run conversations for cumulative memory growth |
unknown |
| unknown |
MemTensor/skills-vote |
Self-Evolving Skill Selection and Benchmark Pipeline |
generate multiple candidate skill mutations -> score candidates with voting-style evaluators on benchmark tasks -> retain winning variants in the skill pool -> iterate selection to improve downstream agent performance across tasks |
unknown |
| unknown |
memtomem/memtomem |
Hierarchical Agent Memory Framework |
capture episodic and semantic memory traces -> structure memory in hierarchical graphs -> retrieve context by relevance and recency -> feed persistent memory context back into ongoing autonomous task execution |
unknown |
| unknown |
mgechev/skillgrade |
Agent Skill Evaluation Harness |
SKILL.md package -> eval.yaml tasks and graders -> sandboxed agent trials -> pass-rate gate |
unknown |
| unknown |
mgechev/skills-best-practices |
Agent Skill Authoring Methodology |
skill need -> trigger-optimized frontmatter -> lean SKILL.md -> references/scripts/assets -> discovery/logic/edge-case validation -> regression-aware skill iteration |
unknown |
| unknown |
microsoft/agent-lightning |
Reinforcement-Learning Agent Training Framework |
decouple agent execution from RL training through unified trajectories -> build a training-agent disaggregation architecture -> optimize downstream agent policies with LightningRL credit assignment -> feed validated gains back into agent runtime loops |
unknown |
| unknown |
microsoft/SkillOpt |
Self-Evolving Agent Skill Optimizer |
collect trajectories -> propose skill edits -> validate on held-out tasks -> keep stronger best_skill artifacts -> repeat like epochs and mini-batches without touching base model weights |
unknown |
| unknown |
microsoft/waza |
Waza Agent Skill Evaluation CLI |
SKILL.md asset -> eval scaffold -> benchmark run -> grader/coverage report -> skill quality gate |
unknown |
| unknown |
mindfold-ai/Trellis |
Cognitive Workspace Agent Runtime |
agent workspace with visual browser timelines -> workspace state and memory graph persistence -> explicit logic layer for plan/edit/review loops -> local execution with web and tool integrations |
unknown |
| unknown |
mlcommons/modelbench |
Model Safety Benchmark and Reporting Framework |
model responses and annotator judgments -> hazard aggregation into benchmark scores -> safety report generation -> benchmark governance feedback into model evaluation pipeline |
unknown |
| unknown |
mnemon-dev/mnemon |
Persistent Memory Substrate for Cross-Session Agent Recall |
store agent knowledge in graph-shaped persistent memory -> enable cross-session recall with LLM-supervised consolidation -> feed historical memory into current task reasoning -> improve continuity for multi-agent CLI operations over time |
unknown |
| unknown |
Modelcode-ai/mcode-benchmark |
Repository-Scale Agent Translation Benchmark |
source repository workspace -> agent performs cross-language/framework translation -> hidden tests evaluate functional equivalence -> benchmark outputs per-language/task reliability |
unknown |
| unknown |
momo-personal-assistant/openclaw-plugin |
Personal Assistant Plugin for OpenClaw Workflows |
encode assistant behaviors as plugin capabilities inside OpenClaw -> retain user/task state and reminders across sessions -> orchestrate personal productivity actions through constrained tool calls -> evolve assistant behavior from interaction feedback and retained context |
unknown |
| unknown |
murataslan1/ai-agent-benchmark |
Multi-Domain Agent Benchmark Pack |
define multi-domain task suites -> evaluate coding/math/memory/translation and safety behavior -> score cross-model outcomes -> expose benchmark schema for reproducible comparisons |
unknown |
| unknown |
mvanhorn/last30days-skill |
Reproducible Agent Skill Benchmark and Evaluation Harness |
define benchmark tasks and scoring protocol for skill-driven agent runs -> execute task trajectories under controlled harness settings -> compare variants over reproducible metrics and historical windows -> retain high-performing skill behaviors while flagging regressions |
unknown |
| unknown |
n8n-io/n8n |
工作流自动化 |
可视化工作流 → 节点编排 → AI Agent 节点 |
unknown |
| unknown |
najeed/ai-agent-eval-harness |
Enterprise Multi-Agent Evaluation and Verification Harness |
simulate business workflows through benchmark scenarios and shims -> replay deep traces for verification and debugging -> compare agent reliability across environments and workflows -> close the agentic reliability gap with explicit eval infrastructure |
unknown |
| unknown |
nemori-ai/nemori |
Episodic Agent Memory Substrate and Knowledge Store |
episodic interaction capture -> memory graph indexing and retrieval -> semantic recall for future agent plans -> persistent memory feedback into subsequent actions |
unknown |
| unknown |
NirDiamant/Agent_Memory_Techniques |
Agent Memory Technique Cookbook |
memory need -> 30 runnable techniques -> taxonomy/decision tree -> evaluation and production notebooks |
unknown |
| unknown |
nomic-ai/aec-bench |
Agentic Context Engineering Benchmark Suite |
construct realistic context-heavy agent tasks -> compare retrieval, memory, and orchestration strategies -> benchmark long-context reasoning under controlled settings -> convert benchmark outcomes into actionable harness/memory optimizations |
unknown |
| unknown |
NousResearch/hermes-agent-self-evolution |
On-Policy RL Self-Evolution Pipeline for Agent Models |
collect trajectories from environment interaction -> run self-play and reward-driven filtering -> distill improved policy behavior into Hermes checkpoints -> iterate closed-loop updates to increase task-level coding and reasoning performance |
unknown |
| unknown |
nowledge-co/community |
OpenClaw Community Skills and Runtime Integration Hub |
curate community-authored OpenClaw skills and runtime practices -> map reusable skill packages to real execution scenarios -> standardize onboarding paths for contributors and operators -> accelerate practical adoption of multi-skill agent workflows |
unknown |
| unknown |
NVIDIA/skills |
Enterprise Agent Skill Registry and Runtime Templates |
package task procedures as reusable skills -> define declarative metadata for retrieval and execution -> map skill modules to agent workflow runtimes -> reduce repeated prompt engineering by codifying high-signal operational playbooks |
unknown |
| unknown |
oceanbase/powermem |
Agent Memory Plugin and Retrieval Augmentation Layer |
augment agent pipelines with explicit memory plugin boundaries -> optimize recall quality and retrieval cost across workflows -> provide reusable memory layer for multi-step decisions -> increase agent consistency through persistent context integration |
unknown |
| unknown |
ollama/ollama |
LLM 基础设施 |
本地推理 → 模型管理 → API 服务 |
unknown |
| unknown |
Olshansk/agent-skills |
Reusable Agent Skill Library |
package reusable operational skills -> validate with local tests and linting -> install into agent runtimes as modular capability units -> iterate via versioned skill updates |
unknown |
| unknown |
open-gitagent/gitagent |
Git-Native Agent Framework |
git repository -> agent identity/rules/memory/tools/skills/hooks -> auditable agent runtime |
unknown |
| unknown |
open-webui/open-webui |
自托管 AI 平台 |
自托管 → 多 LLM → RAG → 插件 |
unknown |
| unknown |
openai/swarm |
Experimental Multi-Agent Orchestration Framework |
compose lightweight routines and handoffs -> route user tasks across specialized agents -> keep tool usage explicit and inspectable -> treat orchestration as a simple educational baseline rather than a production control plane |
unknown |
| unknown |
OpenBMB/AgentVerse |
多 Agent 仿真平台 |
智能体编排 → 反思记忆 |
unknown |
| unknown |
OpenBMB/ChatDev |
多 Agent 协作框架 |
虚拟公司 → 角色对话链 → 软件开发 |
unknown |
| unknown |
OpenBMB/ClawXMemory |
OpenClaw Long-Term Memory Module |
background indexing of chat sessions -> markdown file memories + sqlite control-plane -> model-guided recall selection -> dashboard traces for recall/index/dream lifecycle |
unknown |
| unknown |
openclaw/acpx |
State-Preserving Agent Runtime and Session Handoff |
preserve session state across agent switches -> coordinate skills through ACP-compatible runtime contracts -> maintain continuity for long-running workflows -> turn ad hoc tool use into transferable orchestration assets |
unknown |
| unknown |
openclaw/clawbench |
Trace-Scored Full-Stack Agent Benchmark |
run container-isolated tasks -> capture full execution traces -> score deterministic completion, trajectory quality, and behavior -> quantify noise and failure regimes -> compare harness/model/config combinations |
unknown |
| unknown |
openclaw/clawhub |
OpenClaw Package Catalog and Skill Distribution Hub |
package discovery and rating hub -> curated OpenClaw package metadata -> install and publish flow -> reusable skill and harness package circulation |
unknown |
| unknown |
openclaw/clownfish |
Maintainer Codex Harness for Issue Clusters |
crawl and cluster large issue queues -> route clusters into codex-maintainer workflows -> execute fixes and verification steps -> feed resolved cases back into maintainable engineering loops |
unknown |
| unknown |
openclaw/crabbox |
Browser Agent Benchmark and Evaluation Harness |
define browser-agent task suites -> execute agent runs with standardized tooling and policies -> score outcomes with reproducible evaluators -> provide benchmark signals for agent iteration and harness tuning |
unknown |
| unknown |
openclaw/crabpot |
OpenClaw Plugin Compatibility Testbed |
assemble community plugin scenarios -> run compatibility checks across plugin boundaries -> surface breakage and integration regressions -> guide stable skill/plugin release pipelines |
unknown |
| unknown |
openclaw/crawlkit |
Shared Crawl Infrastructure Toolkit |
provide shared crawl primitives and storage abstractions -> standardize archive generation across crawler services -> reduce duplicate data-ingest logic -> improve reuse for harness and memory data pipelines |
unknown |
| unknown |
openclaw/discrawl |
Discord Archive and Memory Ingest Harness |
crawl Discord channels through CLI workflows -> persist conversations into SQLite archives -> expose query-ready organizational memory artifacts -> support agent learning and maintainer context retrieval loops |
unknown |
| unknown |
openclaw/gitcrawl |
Local-First GitHub Crawl and Archive Harness |
crawl GitHub issues and pull requests locally -> normalize and archive repository discourse -> expose structured artifacts for triage and analysis -> feed downstream maintainer and agent memory workflows |
unknown |
| unknown |
openclaw/openclaw-windows-node |
Windows Companion Runtime for Agent Execution |
bridge OpenClaw workflows into Windows environments -> run companion node services with OS-specific integration -> keep agent operations portable across platform constraints -> feed platform results back into broader harness reliability |
unknown |
| unknown |
OpenDevin/OpenDevin |
AI 软件开发平台 |
智能体编排 → 反馈-精炼 |
unknown |
| unknown |
openmemoryspec/oms |
Portable Agent Memory Interoperability Standard |
define immutable memory grain containers -> standardize query and context-assembly languages -> enforce portable and auditable memory exchange across runtimes -> reduce lock-in and memory migration friction for long-lived agents |
unknown |
| unknown |
opensquilla/opensquilla |
Token-Efficient Agent Runtime with OpenClaw/MCP/Memory Integration |
optimize agent intelligence density under fixed token budgets -> combine runtime controls with memory and MCP connectivity -> keep execution quality stable while reducing context waste -> improve long-horizon self-improving loops with explicit efficiency constraints |
unknown |
| unknown |
openswarm-ai/openswarm |
Multi-Agent Swarm Orchestration Framework with Lightweight Runtime Control |
assemble multiple role-specific agents into swarm workflows -> dispatch tasks across shared control boundaries -> coordinate context and handoff logic through lightweight runtime primitives -> support reproducible multi-agent execution loops in production-like settings |
unknown |
| unknown |
Orchestra-Research/AI-research-SKILLs |
Agent Research Skill Library |
research skill library -> autoresearch orchestration -> evaluation, agents, prompting and paper workflow skills |
unknown |
| unknown |
OWASP/www-project-agent-memory-guard |
Agent Memory Poisoning Defense and Guard Layer |
screen every memory read/write through detectors -> enforce declarative security policy -> emit forensics-ready events and snapshots -> block persistent memory poisoning before it propagates across sessions |
unknown |
| unknown |
paradigmxyz/centaur |
Secure Team Agent Runtime |
Slack/API request -> durable control plane -> sandboxed harness -> tools/workflows -> replayable team result |
unknown |
| unknown |
paradigmxyz/evmbench |
Smart Contract Agent Benchmark Harness |
contract upload -> sandboxed Codex detect worker -> JSON vulnerability report -> UI/report validation |
unknown |
| unknown |
pegasi-ai/reins |
Self-Improving Agent Policy Framework and Training Harness |
constrain and optimize agent behavior with explicit reinforcement policies -> score behavior against undesired actions and alignment constraints -> update control policies as reusable guardrails -> compound safer self-improving behavior in repeated execution loops |
unknown |
| unknown |
phidatahq/phidata |
Agent 框架 |
Agent → 记忆 + 知识 + 工具 → 执行 |
unknown |
| unknown |
Picrew/awesome-agent-harness |
Awesome Agent Harness Landscape |
aggregate benchmark suites and harness runtimes -> map evaluation dimensions and reliability criteria -> link open-source implementation references -> maintain rapid ecosystem comparison entrypoint |
unknown |
| unknown |
pinchbench/skill |
Real-World Agent Task Benchmark |
task suite -> OpenClaw agent execution -> automatic and/or LLM judging -> transcript retention -> optional leaderboard upload |
unknown |
| unknown |
plaited/agent-eval-harness |
CLI Agent Evaluation Harness with Schema-Driven Trial Pipelines |
define adapter schemas for any CLI agent -> capture raw trajectories over task suites -> grade and compare multi-run outputs -> turn agent release quality into repeatable pass@k-style evidence |
unknown |
| unknown |
pureples/pureples |
GP+LLM 代码进化 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
pwrdrvr/openclaw-codex-app-server |
Harness-Oriented App Server for OpenClaw and Codex Workflows |
host OpenClaw runtime endpoints behind an app server layer -> connect Codex and external model providers through unified interfaces -> control task execution and environment state with server-level policies -> enable reproducible harness loops for agent coding workflows |
unknown |
| unknown |
QF-Bench/QuantitativeFinance-Bench |
State-Aware Financial Agent Benchmark Suite |
package stateful quantitative tasks in reproducible sandboxes -> run oracle and real-agent evaluation paths -> enforce test-based numeric verification -> use benchmark deltas to tune agent harness and reasoning reliability |
unknown |
| unknown |
qpiai/Proced_mem_bench |
Procedural Memory Retrieval Benchmark |
task trajectory corpus -> procedural retrieval methods -> LLM-as-judge plus IR metrics -> benchmark reports for procedural memory quality |
unknown |
| unknown |
QuantaAlpha/GitTaskBench |
Repo-Level Code Agent Benchmark Harness |
repo-level task suites -> environment setup and incremental bug-fixing traces -> cost-aware alpha metrics for code-agent performance -> multi-agent runner comparison across real repositories |
unknown |
| unknown |
QuantClaw/QuantClaw |
Quantitative Agent Harness Runtime |
market data ingestion -> autonomous analysis and execution planning -> tool and strategy orchestration -> runtime feedback loop for quantitative task workflows |
unknown |
| unknown |
RangeKing/self-evolving-agent |
OpenClaw Self-Evolving Skill |
agent run -> .evolution workspace -> evaluation/curriculum -> promoted capability |
unknown |
| unknown |
razroo/state-trace |
state-trace Agent Memory Engine |
agent log step -> typed memory node/edge -> capacity-aware decay -> graph traversal retrieval |
unknown |
| unknown |
redis/agent-memory-server |
Agent Memory Runtime and Context Service |
capture agent events and context -> store and retrieve memory through Redis-backed services -> expose memory operations via MCP and client APIs -> feed retrieved context into later agent loops |
unknown |
| unknown |
reworkd/AgentGPT |
自主 Agent 平台 |
自主循环 → 任务分解 → 执行 → 学习 |
unknown |
| unknown |
rohitg00/awesome-openclaw |
OpenClaw Plugin and Agent Skills Resource Index |
aggregate plugin, memory, observability, deployment, and benchmark resources around OpenClaw workflows -> structure links by operational problem class -> give installable and reusable pathways for skills and channels -> increase reproducibility of agent capability composition |
unknown |
| unknown |
rucbm/laser |
Last-Token Self-Rewarding Reinforcement Learning Recipe |
optimize RLVR objective -> learn last-token self-reward signal -> reuse auxiliary reward during training and testing -> improve reasoning and reward calibration together |
unknown |
| unknown |
RyanAlberts/best-of-Agent-Harnesses |
Ranked Agent Harness Landscape Index |
collect harness projects -> score and rank ecosystem coverage -> expose category tags and update cadence -> provide comparative entrypoint for reliability-oriented harness selection |
unknown |
| unknown |
sachinsharma9780/memweave |
Persistent Agent Memory Substrate |
agent writes memory markdown -> sqlite vector+fts index build -> hybrid retrieval and reranking -> persistent memory feedback for next agent turns |
unknown |
| unknown |
SamurAIGPT/awesome-openclaw |
OpenClaw Ecosystem Curation and Skill Resource Index |
curate OpenClaw tools, skills, tutorials, and ecosystem modules into one navigable index -> group resources by workflow role and operating context -> shorten adoption path for newcomers and operators -> turn fragmented community assets into reusable skill discovery infrastructure |
unknown |
| unknown |
sd0xdev/sd0x-dev-flow |
Claude Code Harness Safety Runtime |
hook lifecycle -> state-machine gates -> dual-review approvals -> fail-closed enforcement -> reusable skill pack |
unknown |
| unknown |
seb1n/awesome-ai-agent-skills |
Cross-Agent Skill Index and Install Guide |
collect reusable agent skills in one registry -> map installation paths for multiple agent runtimes -> standardize skill formatting and metadata -> reduce bootstrapping friction for reproducible skill reuse |
unknown |
| unknown |
sevenschulte/agentic-harness |
Python Agent Workflow Testing Harness |
define reusable workflow nodes for agent pipelines -> attach deterministic tests to workflow behavior -> run harness checks before production rollout -> use failures as fast feedback for agent workflow evolution |
unknown |
| unknown |
shareAI-lab/kbench |
Agent Harness Benchmark CLI |
benchmark bridge -> kbench CLI -> built-in or custom agent harness -> standardized run artifacts |
unknown |
| unknown |
shareAI-lab/learn-claude-code |
Claude Code Skill Learning Curriculum |
structured learning path -> daily skill tasks with runnable examples -> command and workflow rehearsal -> advanced orchestration patterns for practical coding-agent productivity |
unknown |
| unknown |
siyuyuan/evoagent |
进化式多 Agent 系统 |
进化/搜索循环 → 智能体编排 |
unknown |
| unknown |
skillmatic-ai/awesome-agent-skills |
Cross-Framework Agent Skills Registry |
curate reusable skills in one index -> map installation paths for multiple agent runtimes -> standardize skill packaging and integration hints -> reduce onboarding and reproducibility cost for skill reuse |
unknown |
| unknown |
smol-ai/developer |
AI 开发助手 |
最小 Agent → 代码生成 → 迭代 |
unknown |
| unknown |
soimy/openclaw-channel-dingtalk |
Agent Channel Plugin for Enterprise Messaging Runtime |
wrap enterprise messaging APIs as OpenClaw channel adapters -> route conversations and action callbacks through typed plugin interfaces -> persist channel events as reusable context for follow-up workflows -> turn communication endpoints into reusable agent operation skills |
unknown |
| unknown |
sola-st/repairagent |
Autonomous Java Bug Repair Agent |
read failing test -> localize bug -> analyze code -> generate patch -> run tests -> iterate until a correct fix survives validation |
unknown |
| unknown |
sourcegraph/CodeScaleBench |
Enterprise-Scale Coding Agent Benchmark Harness |
enterprise codebase tasks -> Harbor/Claude harness with baseline vs MCP retrieval configs -> dual-verifier scoring and cost tracking -> auditable snapshots for benchmark governance |
unknown |
| unknown |
SponsioLabs/Sponsio |
Workflow Automation and Multi-Agent Control Infrastructure |
coordinate multi-agent workflow execution with explicit orchestration contracts -> maintain shared task state and control checkpoints across agents -> expose repeatable automation patterns for teams and pipelines -> improve end-to-end delivery reliability under autonomous execution |
unknown |
| unknown |
stanford-iris-lab/meta-harness |
Meta-Harness Framework and Reference Experiments |
define domain spec -> search harness candidates -> run reference experiments -> compare outcomes -> retain stronger harness |
unknown |
| unknown |
stitionai/devika |
AI 软件工程师 |
智能体编排 → 反馈-精炼 |
unknown |
| unknown |
sunnja69/akephalos |
Local-First Agent Passport Memory Bundle |
local passport init -> markdown/jsonl memory updates -> multi-agent bundle sync -> optional mcp serving |
unknown |
| unknown |
supabase/agent-skills |
Agent Skill Packs and Prompt Compression Patterns |
encode domain-specific playbooks as compact skill assets -> inject them into agent context only when needed -> preserve high-value workflow constraints and SQL/product knowledge -> stabilize assistant output quality through reusable prompt primitives |
unknown |
| unknown |
SuperagenticAI/metaharness |
Benchmark-Driven Harness Evolution Toolkit |
propose harness change -> run benchmark matrix -> compare score/runtime/cost -> keep best candidate -> persist ledger |
unknown |
| unknown |
supermemoryai/supermemory |
Open AI Memory Infrastructure |
chat/browser context ingest -> memory indexing -> retrieval scoring -> personalization -> downstream agent loop reuse |
unknown |
| unknown |
suyoumo/ClawProBench |
Live OpenClaw Benchmark Harness |
OpenClaw runtime task -> live scenario execution -> deterministic grading -> structured report -> leaderboard/profile selection |
unknown |
| unknown |
swapedoc/hermes2anti |
Memory and Skill Self-Improvement Toolkit |
task session -> golden path extraction -> skill creation/security scan -> memory recall |
unknown |
| unknown |
swarmclawai/swarmclaw |
Self-Hosted Agent Runtime |
agent runtime -> memory and MCP connectors -> schedules and delegation -> swarm workflows -> self-hosted distribution and release cadence |
unknown |
| unknown |
SWE-bench/SWE-bench |
Agent 评测基准 |
真实 GitHub Issue → 模型生成 Patch → 评估 |
unknown |
| unknown |
syntax-syndicate/OpenHarness-agent-harness |
Open Agent Harness Runtime and Evaluation Workflow |
define standardized harness interfaces for task setup, execution, and judging -> route agent actions through controlled runtime adapters -> log traces and outcomes for comparable replay -> make harness-level protocol changes testable before model-level retraining |
unknown |
| unknown |
Team-Commonly/commonly |
Multi-Agent Swarm Orchestration Runtime and Workflow Infrastructure |
organize agent teams into role-based workflow lanes -> attach memory and control rules per lane -> route tasks across shared context and approval boundaries -> preserve swarm-level learning and execution continuity across long-horizon work |
unknown |
| unknown |
Tencent/TencentDB-Agent-Memory |
Local Long-Term Agent Memory Substrate |
symbolic short-term memory plus layered long-term memory -> plugin-based integration into agent runtimes -> local-first persistence pipeline -> benchmarked token and pass-rate impact reporting |
unknown |
| unknown |
thinkwee/AgentsMeetRL |
Agentic RL and Benchmark Knowledge Index |
aggregate agentic-RL methods and benchmark references into a structured index -> cluster methods by training/evaluation setting -> expose comparison anchors for harness design and evaluator selection -> support downstream pipeline choices with curated benchmark evidence |
unknown |
| unknown |
ThisIsJeron/awesome-openclaw-plugins |
OpenClaw Plugin Catalog and Community Knowledge Index |
organize community plugins by channels, memory, security, observability, and self-improvement roles -> clarify skills versus plugins for execution boundaries -> provide install references and integration routes -> shorten discovery time for building agent capability stacks |
unknown |
| unknown |
THUDM/AgentBench |
Agent 评测基准 |
评估器/打分器 |
unknown |
| unknown |
TIGER-AI-Lab/ClawBench |
Open-Ended Agent Benchmark Harness |
open-ended task generation -> long-horizon agent execution traces -> verifier-guided scoring -> benchmark snapshots for iterative harness improvement |
unknown |
| unknown |
TransformerOptimus/SuperAGI |
自主 Agent 框架 |
自主 Agent → 工具生态 → 任务执行 |
unknown |
| unknown |
UnicomAI/hexagent |
LLM Computer Harness Runtime |
runtime/computer separation -> pluggable local-vm-cloud computer protocol -> middleware hooks and skill injection -> isolated subagent execution with MCP/tool orchestration |
unknown |
| unknown |
vectorize-io/agent-memory-benchmark |
Agent Memory Benchmark |
ingest documents and traces -> retrieve candidate context -> generate agent answer -> judge accuracy and cost -> compare memory strategies across datasets and modes |
unknown |
| unknown |
VectorSpaceLab/general-agentic-memory |
General Agentic Memory Framework with Cross-Task Reuse |
design generalized memory abstractions for diverse agent tasks -> persist and retrieve context under shared memory interfaces -> reuse memory operations across domains and workflows -> scale agent continuity without per-task memory rewrites |
unknown |
| unknown |
Versatly/clawvault |
Persistent Memory Runtime for OpenClaw-Style AI Agents |
store structured memory for autonomous agent workflows -> expose retrieval and evaluation surfaces to long-horizon tasks -> integrate memory into OpenClaw-style execution loops -> treat persistence as a runtime subsystem rather than a prompt-only trick |
unknown |
| unknown |
vilmire/adhdev |
Coding-Agent Control Plane |
coding-agent session -> local dashboard/control plane -> approval, status, history and continuation |
unknown |
| unknown |
voltagent/awesome-agent-skills |
Agent Skills Resource Index |
collect official and community skill packs -> normalize reader entry by tool and provider -> expose reusable procedures as installable skills -> keep the engineering skill ecosystem searchable and comparable |
unknown |
| unknown |
VoltAgent/awesome-openclaw-skills |
OpenClaw Skill and Agent Workflow Index |
collect OpenClaw skills and tools -> categorize by use case and domain -> provide fast lookup and install references -> support reusable skill adoption |
unknown |
| unknown |
VRSEN/agency-swarm |
OpenAI Agents SDK Swarm Orchestrator |
define agents and directional communication flows -> attach function tools and persistence callbacks -> route work through agency-level orchestration -> reuse terminal/web demos and docs as reproducible multi-agent operating patterns |
unknown |
| unknown |
wanxingai/LightAgent |
Memory/MCP Skill Agent Framework |
compose lightweight agents with tools, MCP, and memory -> add native skills and optional trace observability -> delegate via LightSwarm -> chain deterministic multi-step flows with LightFlow -> keep self-learning behavior grounded in runtime memory and reusable tool plans |
unknown |
| unknown |
wazionapps/nexo |
NEXO Agent Memory Runtime |
conversation/session traces -> cognitive memory extraction -> semantic/temporal retrieval -> trust/forgetting gates -> proactive context packets |
unknown |
| unknown |
weaviate/query-agent-benchmarking |
Agent Benchmark Toolkit for Query/Retrieval Evaluation |
package benchmark scenarios for query-agent evaluation -> measure retrieval and answer quality across controlled tasks -> make evaluation pipelines reusable and comparable -> provide practical evidence surface for agent benchmark governance |
unknown |
| unknown |
web-arena-x/webarena |
Agent 评测基准 |
Web 环境 → Agent 浏览 → 任务完成评估 |
unknown |
| unknown |
webmaxru/Agent-Skills |
Reviewed Web API Agent Skills |
Web API source material -> skill authoring -> validation/remediation -> install verification |
unknown |
| unknown |
webzler/agentMemory |
Benchmark Framework for Agent Memory Evaluation and Hallucination Testing |
define memory-capability evaluation tasks -> execute benchmark cases across recall and hallucination dimensions -> report scorecards for different agent memory strategies -> provide reproducible baseline harness for memory quality claims |
unknown |
| unknown |
wuxingyu-ai/LLM4EC |
LLM+EC 交叉综述 |
文献综述 |
unknown |
| unknown |
xai-liacs/LLaMEA |
LLM 驱动算法自动发现 |
进化/搜索循环 → 评估器/打分器 |
unknown |
| unknown |
xiaofangxd/LLM_EA |
LLM+EA 交叉综述 |
文献综述 |
unknown |
| unknown |
xlang-ai/OpenAgents |
Agent 工具使用 |
工具调用 → 函数选择 → 代码执行 |
unknown |
| unknown |
xlang-ai/OSWorld |
Agent 评测基准 |
桌面 OS 环境 → Agent 操作 → 任务评估 |
unknown |
| unknown |
XMUDeepLIT/Awesome-Self-Evolving-Agents |
自进化 Agent 综述 |
综述索引 → 自进化 Agent 论文集合 |
unknown |
| unknown |
XSkill-Agent/XSkill |
Continual Experience and Skill Learning Paper Code |
collect multimodal trajectories -> summarize and critique experiences -> consolidate reusable skill documents and experience entries -> retrieve relevant memory for new tasks -> evaluate transfer on benchmark suites |
unknown |
| unknown |
yennning/awesome-code-as-agent-harness-papers |
Code-As-Agent-Harness Survey Index |
collect code-centric agent papers -> regroup them by harness interface, mechanism, and scaling pattern -> expose memory, tool, debugging, and multi-agent topology lanes -> provide a navigable harness taxonomy |
unknown |
| unknown |
yoheinakajima/babyagi |
自主 Agent 框架 |
目标 → 任务分解 → 优先级 → 执行 → 学习 |
unknown |
| unknown |
yoloshii/ClawMem |
On-Device Memory Layer and Retrieval Runtime for Agents |
index local documents and session artifacts into a persistent memory substrate -> combine hybrid retrieval, hooks, and MCP tooling to surface relevant context automatically -> preserve decisions and handoffs across sessions and agents -> enable compounding memory quality through repeated retrieval and feedback loops |
unknown |
| unknown |
YoungDubbyDu/LLM-Agent-Optimization |
LLM Agent 优化综述 |
文献综述 |
unknown |
| unknown |
YuanchenBei/Mem-Gallery |
Long-Term Memory Benchmark Suite |
assemble memory-intensive tasks and temporal-context datasets -> run agents with different memory strategies -> score recall/consistency/retrieval behavior -> compare long-term memory robustness across setups |
unknown |
| unknown |
YunjueTech/Yunjue-Agent |
In-Situ Self-Evolving Agent System |
open-ended task stream -> tool evolution -> reusable capabilities -> trace/reproduction audit |
unknown |
| unknown |
ZeroLu/awesome-openclaw |
OpenClaw Community Landscape and Resources |
aggregate OpenClaw links and toolkits -> map onboarding resources and examples -> curate ecosystem entry points for rapid adoption |
unknown |
| unknown |
Zesearch/self-improvement-llm |
LLM 自改进综述 |
文献综述 |
unknown |
| unknown |
zhang677/accelopt |
Self-Improving Accelerator Kernel Optimization Agent |
generate candidate kernel -> consult optimization memory -> profile on NKIBench or FlashInfer-Bench -> compare slow-fast kernel pairs -> keep stronger optimization traces |
unknown |
| unknown |
zhangfengcdt/memoir |
Git-like Agent Auto-Memory |
agent activity -> hierarchical memory paths -> Git-like commits/branches -> recoverable continuity |
unknown |
| unknown |
Zijian-Ni/awesome-ai-agents-2026 |
Agent 研究综述 |
2026 Agent 追踪 → 实时更新 |
unknown |
| unknown |
zikuicai/aegisllm |
Self-Reflective Multi-Agent Defense System |
coordinate orchestrator-deflector-responder-evaluator roles -> evaluate adversarial and unlearning threats -> optimize prompts with DSPy loops -> improve runtime defense quality without model retraining |
unknown |
| unknown |
zjunlp/SkillX |
Automated Agent Skill KB Construction |
collect trajectories -> extract multi-level skills -> refine and filter skill library -> expand via exploration -> transfer to other agents |
unknown |
| unknown |
zocomputer/skills |
Open Agent Skills Registry and Distribution Layer |
maintain a curated multi-source skills registry -> validate skill package structure -> sync external skill feeds into a manifest -> distribute reusable skills into agent runtimes with consistent metadata |
unknown |
| unknown |
zorazrw/agent-workflow-memory |
Agent Workflow Memory Runtime with Knowledge Graph Integration |
link memory manager and knowledge graph into workflow execution -> route task context through persistent memory nodes -> query historical traces to stabilize next-step planning -> reduce drift across multi-step agent workflows |
unknown |