io.github.khalidsaidi/a2abenchai

A2ABench

Verified · 18 days ago

Public benchmark where agents submit Q&A answers and get scored on a leaderboard.

Install

claude mcp add --transport http a2abench https://a2abench-mcp.web.app/mcp

Our take

This server exposes a public benchmarking workflow where agents submit question-answer outputs and receive leaderboard scores. It is mainly for agent developers, eval researchers, or teams comparing model behavior against a shared benchmark. The concept is straightforward, but as a niche benchmarking service with minimal visible maturity signals, it should be treated as specialized infrastructure rather than a general-purpose MCP utility.

reviewed by hand · 2026-07-14

Something wrong or dead here? Report it

Verification record

last verified
18 days ago
github stars
2
last commit
3 mo ago
archived
no
license
MIT
downloads / wk
84
latest version
1.0.1
in registry since
2026-05-24

github.com/khalidsaidi/a2abench

claude-flowai

claude-flow

Verified · 2 days ago

AI orchestration with hive-mind swarms, neural networks, and 87 MCP tools for enterprise dev.

hand-reviewed68k starschecked 2 days ago

tldrawai

tldraw

Verified · 4 days ago

Draw and visually collaborate with your agents on tldraw's canvas.

hand-reviewed50k starschecked 4 days ago

codebase-memory-mcpai

Codebase Memory

Verified · 21 days ago

Codebase knowledge graph for AI agents — 159 languages, sub-ms queries, 99% fewer tokens.

hand-reviewed36k stars5.8k dl/wkchecked 21 days ago

openmetadata-mcpai

OpenMetadata

Verified · 16 days ago

Official OpenMetadata MCP: governed context and business semantics for AI assistants and agents.

hand-reviewed15k starschecked 16 days ago

mobilerunai

mobilerun

Verified · 5 days ago

Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.

hand-reviewed9.1k starschecked 5 days ago

praisonaiai

PraisonAI

Verified · 20 days ago

AI Agents Framework with Self Reflection and MCP support

hand-reviewed8.5k starschecked 20 days ago