Skip to content
AgentBench logo

AgentBench

Benchmark your AI agent across 40 real-world tasks.

Bootstrapped
LLMAI agentbenchmarkevaluation
Share:

About

AgentBench is a benchmarking suite designed to evaluate the performance of LLMs acting as agents across various interactive environments and tasks. It assesses reasoning and decision-making in multi-turn open-ended settings, with rule-based scoring and 3-layer automated evaluation.

This description came from a bulk import and has not yet been checked against the vendor's own page. Treat it as unverified.

Enterprise Readiness

from cataloged data

Not scored. Nobody has checked this entry against AgentBench’s own site yet, so the signals below are unverified and we will not turn them into a rating. The breakdown shows what our record holds.

Security & ComplianceNot disclosed

No compliance certifications listed

Deployment FlexibilityStrong

Flexible deployment: Cloud, Self-Hosted

Integration DepthPartial

2 documented integrations

Proven AdoptionPartial

5 documented use cases

Company MaturityPartial

Founded 2023

Composite of public compliance, deployment, integration, adoption, and company signals. “Not disclosed” reflects gaps in available data, not a vendor deficiency.

Independently sourced

Pulled directly from regulator filings and platform APIs. No vendor or editorial input.

Activity classifier
Active
last GitHub commit 89d ago
GitHub API2026-05-27
agentbench
Open source product· 9· MIT· active 6mo ago
Shell
cooling · +2★ (+40%) 30d
OSV vulnerability feed2026-08-24
No known advisories

Enterprise Use Cases

Evaluating LLM agents
Benchmarking AI models
Assessing reasoning and decision-making
Testing multi-step workflows
Measuring tool efficiency

Integrations

Claude CodeOpenClaw

How AgentBench compares

ToolReadinessPricingDeploymentComplianceIntegrations
AgentBenchNot scoredFreeCloud, Self-Hosted2
Exa (formerly Metaphor)Not scoredFreemiumCloud, Self-Hosted8
Paperpal AINot scoredFreemiumCloud4
Research Rabbit AI51 · EstablishedFreemiumCloud3

Peers from the same category, ranked by Enterprise Readiness. Readiness is a composite of cataloged compliance, deployment, integration, adoption, and company signals.

Frequently Asked Questions

What is AgentBench used for?

AgentBench is benchmark your AI agent across 40 real-world tasks. It is commonly used for evaluating llm agents, benchmarking ai models, assessing reasoning and decision-making, and testing multi-step workflows.

Is AgentBench free, and how is it priced?

AgentBench is free to use.

How can AgentBench be deployed?

AgentBench supports Cloud and Self-Hosted deployment. On-premise and self-hosted options support data-residency and air-gapped requirements.

What does AgentBench integrate with?

AgentBench documents 2 integrations, including Claude Code and OpenClaw.

When was AgentBench founded?

AgentBench was founded in 2023.

Is AgentBench enterprise-ready?

Based on cataloged data, AgentBench has an Enterprise Readiness score of 42/100 (Emerging tier), derived from its compliance, deployment, integration, adoption, and company-maturity signals.

Quick Facts

PricingFree
DeploymentCloud, Self-Hosted
Founded2023

Procurement

Evaluating vendors like this one?

The Enterprise AI RFI/RFP asks the security, governance, and lock-in questions sales decks skip.

See the template — RFI $299 / RFP $699

Is this your product?

Claiming is free and lets you correct your listing's data. Upgrade to Enhanced ($199/yr) for a vendor-maintained mark, a richer profile, and lead routing.

·