CommerceBench
Standardised evaluationRealistic, comparable benchmarks for commercial capabilities.
- Structured tasks and datasets
- Comparable performance metrics
- Regular benchmark releases
We build evaluation environments for commerce AI — from simulated markets to live pilots — to understand how systems perform in complex, high-stakes, real-world conditions.
From targeted benchmarks to open-ended environments, we evaluate commerce AI across complementary platforms.
Realistic, comparable benchmarks for commercial capabilities.
Diverse, realistic environments for testing and learning.
Study interactions between people, companies, agents and institutions.
We evaluate commerce AI across a range of outcome metrics that reflect real-world value, risk and operational impact.
Task completion and factual accuracy
Total cost and resource efficiency
Time to complete and responsiveness
Policy, compliance and safety outcomes
Need for human intervention
End-to-end success in live or realistic markets
We go beyond pass/fail to understand how systems behave, recover, and what each outcome means for real-world deployment.
Completes task correctly on first attempt.
Recovers from an error without human help.
Requires human intervention to succeed.
Appears correct but contains material errors.
Cannot complete task even with assistance.
Novel methods for evaluating commerce AI systems.
Synthetic market environments for testing and analysis.
Effective human oversight and hybrid decision-making.
Measuring commercial, societal and market impact.
Stress testing for risk, robustness and resilience.
We release selected tasks, data and methods to enable independent research and meaningful comparison. Private answer keys, proprietary data and adversarial suites remain protected.
Read our approachSelected tasks, datasets and evaluation methods.
Private answer keys, proprietary data and adversarial suites.
Enabling progress while safeguarding commercial integrity.