IMPONEXPO · INTELLIGENT WORLD COMMERCE
impoNexpoBecome a Trade Partner
Research Area 05

Evaluation &
Experimental Economies

Measuring not only correctness, but consequence, recoverability, autonomy, human oversight and real-world commercial performance.

We build evaluation environments for commerce AI — from simulated markets to live pilots — to understand how systems perform in complex, high-stakes, real-world conditions.

Our Evaluation Platforms

Three environments. A more complete picture.

From targeted benchmarks to open-ended environments, we evaluate commerce AI across complementary platforms.

Explore all platforms

CommerceBench

Standardised evaluation

Realistic, comparable benchmarks for commercial capabilities.

  • Structured tasks and datasets
  • Comparable performance metrics
  • Regular benchmark releases
Explore CommerceBench

CommerceGym

Dynamic training environments

Diverse, realistic environments for testing and learning.

  • Multi-step, open-ended scenarios
  • Dynamic market conditions
  • Supports multi-agent interactions
Explore CommerceGym

CommerceArena

Multi-party experiments

Study interactions between people, companies, agents and institutions.

  • Multi-stakeholder simulations
  • Market design and mechanisms
  • Real-world inspired environments
Explore CommerceArena
Key Evaluation Dimensions

Performance in the real world is multi-dimensional.

We evaluate commerce AI across a range of outcome metrics that reflect real-world value, risk and operational impact.

Learn about our methodology

Correctness

Task completion and factual accuracy

Cost

Total cost and resource efficiency

Delay

Time to complete and responsiveness

Risk

Policy, compliance and safety outcomes

Human burden

Need for human intervention

Real transaction success

End-to-end success in live or realistic markets

From success to failure

Not all outcomes are the same.

We go beyond pass/fail to understand how systems behave, recover, and what each outcome means for real-world deployment.

First-pass success

Completes task correctly on first attempt.

Autonomous repair

Recovers from an error without human help.

Human-assisted

Requires human intervention to succeed.

!

False pass

Appears correct but contains material errors.

Blocked failure

Cannot complete task even with assistance.

Public Research Programs

Five programs. Real-world impact.

View all research programs

Evaluation Methods

Novel methods for evaluating commerce AI systems.

Simulated Markets

Synthetic market environments for testing and analysis.

Human-AI Collaboration

Effective human oversight and hybrid decision-making.

Economic Impact Analysis

Measuring commercial, societal and market impact.

Safety, Robustness & Risk

Stress testing for risk, robustness and resilience.

Reproducibility & Benchmark Boundary

Open research, with clear boundaries.

We release selected tasks, data and methods to enable independent research and meaningful comparison. Private answer keys, proprietary data and adversarial suites remain protected.

Read our approach

Public releases

Selected tasks, datasets and evaluation methods.

Protected assets

Private answer keys, proprietary data and adversarial suites.

Trusted ecosystem

Enabling progress while safeguarding commercial integrity.