Deploying and managing production-grade etcd clusters on cloud providers: failure recovery, disaster recovery, backups and resizing.
-
Updated
Sep 26, 2025 - Go
Deploying and managing production-grade etcd clusters on cloud providers: failure recovery, disaster recovery, backups and resizing.
Production style autonomous agent framework in Python + React that plans multistep tasks, persists state with SQLite, and recovers from failures using exponential backoff retry logic. Includes real time monitoring dashboard, deterministic recovery demo, and API for task orchestration.
A programmatic autonomous LLM agent
🎛️🔁🚀 A tiny, zero‑dependencies retry helper with exponential backoff + jitter—usable for KV, HTTP, Durable Objects, or any async function.
Implementation of a distributed database with multi-version concurrency control, replication, deadlock detection and failurerecovery.
证据约束的开源研究 Agent:只读 GitHub 适配器、计划修订、严格取消与引用审查。
Semantic failure memory for AI agents — record reasoning failures, search past failures before acting, stop rediscovering the same mistakes.
Failure-aware Python runtime for tool-using LLM workflows with bounded recovery, output verification, tracing, and reproducible evaluation.
Evidence-checked survey and data atlas of automatic reset, failure recovery, safety, and sustained autonomy in real-world robot RL.
Distributed RAG ingestion pipeline with idempotent vector indexing, DLQ-backed retries, and deterministic failure recovery for production AI systems
Built a 6-agent Python pipeline where agents detect their own reasoning failures, classify them, retrieve similar past failures from memory using RAG, and recover from the exact step where things went wrong. Uses a Bayesian estimator to decide whether recovery is worth retrying. Powered by Gemini 2.5 Flash.
MultiModel AI Agentic Orchestration Framework - Project J.A.R.V.I.S ★ Echo Mac is engineered to be a state-of-the-art, fully autonomous agentic wrapper for your operating system. It operates at the intersection of streaming multimodal AI and deep OS integration.
专为 OpenClaw 打造的 Browser Operations Platform:intelligence、action policy、handoff、autopilot、failure recovery 全进内核
证据驱动的 Browser Agent 运行时:Plan–Act–Verify–Recover、有界恢复与确定性 Playwright 演示。
Fault-tolerant distributed task processor in Java 21 with custom TCP messaging, concurrent workers, failure detection, retries, time-bounded leases, cooperative cancellation, bounded admission control, stale-result rejection, SQLite recovery, and measured benchmarks.
面向工具调用型 AI Agent 的可靠性评测框架,覆盖故障注入、恢复行为、轨迹验证与结果契约。
A lightweight Java finite state machine (FSM) library for distributed systems. Built-in failure recovery, automatic retry, snapshot persistence, per-execution locking, and Spring Boot + JDBC autoconfiguration for PostgreSQL, MySQL, and more.
A practical workflow to classify failed social posts, reconcile ambiguous outcomes, prevent duplicates, and retry safely.
Verifies provider state before AI agents retry or continue ambiguous side effects.
Multimodal browser reliability runtime for failure diagnosis, policy-aware recovery, recovery execution, and verification.
To associate your repository with the failure-recovery topic, visit your repo's landing page and select "manage topics."