Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
-
Updated
Aug 19, 2026 - Python
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
Multi-agent department framework for long-form complex tasks, fighting AI hallucination, validated on academic research. 共识管线:多智能体部门长线任务解决框架,对抗AI幻觉,以学术研究为验证场景。
Official code repo for NeurIPS 2025 Spotlight paper, "Debate or Vote: Which Yields Better Decisions in Multi-Agent LLMs?"
Framework: Multi-Agent LLMs For Conversational Task-Solving (MALLM)
Research-backed methodology for multi-AI collaborative decision-making with structured debate, consensus synthesis, and bias reduction
Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
Human-in-the-loop adversarial workflows for high-stakes research audit: from ChatGPT-Gemini duels to 4-model MAD.
Code for "Multiple LLM Agents Debate for Equitable Cultural Alignment" [ACL 2025 Oral]
An adversarial AI expert workshop that stress-tests a research paper (rival-tradition referees argue; every comment quote-grounded and independently re-verified) and then rebuilds it: tracked-changes redline, clean version, your code re-run under a provenance wall, and a replication package. A Claude Code skill.
Code review, but with 5 models arguing first.
Three Claude Code skills for working with Codex CLI: codex-bridge (one-shot Codex calls), mad-build (Claude+Codex collaboration with cross-review), and mad-research (three-stream adversarial audit of papers, grants, reports with anonymized cross-critique and fresh-Codex synthesis).
Multi-model deliberative design review — a Claude Code skill that runs structured debate between Gemini and GPT to surface blind spots in architecture decisions.
Claude Code plugin for second opinions: iterative adversarial debates between Claude and Codex over any subject, from specs, designs, and code to non-technical positions, ending in genuine agreement or a documented dispute.
Control Stream Deck buttons for Claude Code terminal sessions with live status, tap-to-focus, hold-to-dictate, and snap-to-grid window layout on macOS
A brutally fault-tolerant Mixture-of-Agents (MoA) pipeline built in pure Python. Designed to orchestrate chaotic, round-robin LLM proxy endpoints through a rigorous 4-stage Agentic Workflow (Generate ➔ Cross-Critique ➔ Rebuttal ➔ Judge). Built to eradicate hallucination and guarantee absolute accuracy in complex, multi-step reasoning tasks.
Broadcast one prompt to 2-6 AI models side-by-side - or convene them as a deliberative panel: an AI-only Habermas Machine with blind drafts, anonymous peer review, explicit convergence, and a minority report. Local-first, bring your own keys.
CLAIR-Fin is a nine-agent framework for faithful, cited QA over long multimodal financial documents where prose, tables, and charts disagree. It splits each question into typed claims in a Financial Claim Ledger, weights evidence by claim type, checks grounding mid-pipeline, and routes contested claims to adversarial debate before a final audit.
Visual multi-agent debate pipeline builder, live token streaming playground, and scoped Model Context Protocol (MCP) tool security gateway for Haize Labs Verdict.
Claude Code plugin: open a review topic and AI agents (Claude, Codex) debate it round after round — design, attack, rebuttal — fully hands-free until they deliver a reasoned decision.md. Multi-topic priority queue, human gate only on code changes, git/non-git, claude-solo fallback.
Multi-LLM debate orchestrator that drives ChatGPT, Claude, and DeepSeek web UIs (no API keys) through a 5-phase loop: propose → critique → revise → synthesize → ratify-or-veto. Editorial dark UI.
To associate your repository with the multi-agent-debate topic, visit your repo's landing page and select "manage topics."