Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
-
Updated
Jul 3, 2026 - Python
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress testing, network resilience, and ensemble coordination analysis with statistical validation.
A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey 'LLM Agents: A Survey'.
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents. Works with Claude Code, OpenClaw, Cursor, Windsurf. No fine-tuning required.
Deterministic runtime for agent evaluation
Pit AI coding agents against the same bug. Score them on tests, diff, cost, and time — pick the winning patch.
A curated collection of the world’s most advanced benchmark datasets for evaluating Large Language Model (LLM) Agents.
Measure how DSPy prompt optimization affects the prompt-injection robustness of agentic LLM programs, using AgentDojo's attack suite.
Scores whether an autonomous AI agent actually did the work or just hallucinated its report, by checking every claim against recorded API calls. The core engine behind NoHalu.
Silicon Pantheon - Tactics game played by AI agents coached by human
Release repository for agent benchmark evidence-reporting artifacts and reproduction workflows.
Variance-aware benchmark for AI coding agents. Same agent + same task can swing 70 points — we publish min/max, not just averages. Claude Code · Gemini CLI · Codex CLI · Aider · 10 tasks · Docker sandbox · MIT.
Deterministic evaluation environment for AI code reviewers covering bugs, security (OWASP), and architecture via FastAPI + OpenEnv.
Prediction-market agent arena for AI agent evaluation, paper trading, practice rounds, contests, and leaderboard-based battle testing.
A Pokémon battle arena where any agent can play — human, deterministic game-tree AI, or LLM — over an open WebSocket protocol (MCP, CLI, or your own client). One leaderboard, ranked by who plays best.
🧠 Discover and evaluate advanced benchmark datasets for Large Language Model agents to enhance performance assessment in real-world tasks.
An evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.
Open-source runtime for auditable multi-agent evaluation, starting with professional capital-market investment agents.
Add a description, image, and links to the agent-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the agent-benchmark topic, visit your repo's landing page and select "manage topics."