About this project

ai-jury is a command-line tool that orchestrates multiple AI agents into a review panel for code changes. Rather than calling models only at the API level, it drives native coding-agent CLIs (Claude Code, Codex, Antigravity), hosted API providers (Anthropic, OpenAI, Gemini, xAI Grok, OpenRouter, DeepSeek, Groq, and other OpenAI-compatible endpoints), free local open-weight models via Ollama, llama.cpp, vLLM, or LM Studio, and arbitrary CLI agents such as Aider. Each reviewer runs headless in its own environment; the orchestrator controls the round structure. The workflow is a two-round deliberation: reviewers examine the same diff, PR, or issue independently in parallel, then rebut each other's findings, after which a chair agent verifies and synthesizes a single verdict (APPROVE / COMMENT / REQUEST CHANGES) or the panel decides by vote. It supports reviewing GitHub pull requests, local diffs, and issues (with READY / NEEDS-INFO / UNCLEAR outcomes). Notable capabilities include a `jury init` setup that detects installed agents and local models, presets (offline, fast, balanced, thorough), live and full-transcript output, an animated "theater" deliberation view in flat or pixel-art styles, replay of saved runs without re-spending tokens, incremental review, suggested patches, large-diff chunking, and CI gating. Output formats include markdown, JSON, SARIF, and a per-reviewer keel-reviews array. A pre-commit hook and a GitHub Action are provided. The project emphasizes security by default: reviewers run sandboxed or read-only, and the runtime is standard-library-only (Python 3.11+). It ships a mock mode for offline demos, an opt-in live smoke-test suite, and a small directional benchmark comparing solo models against a panel. Installation is available via Homebrew, a curl installer, or pipx, and the tool can also be installed as a plugin into Claude Code, Codex, Antigravity, or Cursor.