Customized Simulation & Real-Robot Evaluation Platform

X‑Eval

Powered by generative physical simulation as its core engine, the platform delivers end-to-end solutions for robotics AI researchers and engineering teams — from data generation and scientific evaluation to scene asset management.

Comprehensive Task Design for Trustworthy Evaluation

X-Eval supports designing evaluation tasks across key capability dimensions, with customizable combinations tailored to customer business needs. Typical evaluation areas include:

Large-Scale Parallel Simulation Evaluation for Faster Optimization Cycles

Run evaluations for multiple models simultaneously in large-scale scenarios, significantly boosting evaluation throughput and enabling teams to accelerate version comparison, regression testing, and parameter tuning cycles.

From Simulation Assessment to One-Click Real-Robot Evaluation

Eliminates evaluation pipeline fragmentation and bridges the sim-to-real gap, helping customers streamline the journey from post-training validation to pre-deployment acceptance while reducing deployment risks.

Real-robot evaluation media 1/4

01

Sim-to-Real Continuity

Unified policy and protocol to shorten the path from simulation results to real-robot validation

02

Standardized Evaluation Environment

Black background with soft lighting to prevent overexposure and ensure consistent visual conditions

03

High-Precision Scene Replication

Accurate object placement references for improved real-world scene reproduction

04

Customizable Tasks & Scenarios

Task content, robot platforms, and scene configurations tailored to acceptance criteria

End-to-End Capability Loop

Every stage can be customized according to project requirements: standard packages enable rapid deployment, while dedicated solutions align with specific business needs.

  1. 01Data Generation
  2. 02Scene Asset Management
  3. 03Scientific Evaluation in Simulation
  4. 04One-Click Real-Robot Evaluation
  5. 05Acceptance & Iteration

Leaderboards, Data Management, and Result Analysis at a Glance

Evaluation Leaderboard

Model rankings across tasks and capability dimensions

Data Management

Centralized repository for evaluation data, versions, and task assets

Result Analysis

Multi-dimensional comparison to support model selection and iterative improvement.

RANK BY
Overall
ENVIRONMENT
All
TASK SET
50 Tasks
EVALUATION
Clean + Randomized
UPDATED
2026.07.22
11 Policies · 50 Tasks · 100% Complete
Model rankings across tasks and capability dimensions
RankPolicyClean SRRandomized SRRetentionPerformance gapCompletionOverall
01Pi_0570.7%46.0%65.1%24.7pp100/10055.9
02X_VLA68.0%20.9%30.7%47.1pp100/10039.7
03X_WAM61.3%22.7%37.0%38.6pp100/10038.1
04Abot_M057.4%22.9%39.9%34.5pp100/10036.7
05Spatial_Forcing77.2%9.5%12.3%67.7pp100/10036.6
06Xiaomi_Robotics_062.9%18.2%28.9%44.7pp100/10036.1
07EventVLA65.6%15.7%23.9%49.9pp100/10035.7
08FastWAM77.8%1.9%2.4%75.9pp100/10032.3
09GalaxeaVLA62.7%9.1%14.5%53.6pp100/10030.5
10AHA_WAM64.3%3.2%5.0%61.1pp100/10027.6
11starVLA33.1%2.8%8.5%30.3pp100/10014.9

Overall = 40% Clean SR + 60% Randomized SR · Retention = Randomized SR ÷ Clean SR · pp means percentage points

Ecosystem

Integrated with 40+ state-of-the-art models, the platform provides the industry's largest collection of reproduced models and continues to expand its ecosystem partnerships, enabling customers to seamlessly integrate and compare models on a unified platform.

GR00T-N1.7

π Series

OpenVLA-OFT

SmolVLA

MolmoACT2

RDT-1B

Xiaomi-Robotics-0

Hy-Embodied-0.5-VLA

InternVLA-A1

  • A1
  • ACT
  • Aha-WAM
  • Abot-M0
  • Being_H05
  • DP
  • Dexbotic-DM0
  • Dexora-1B
  • DreamZero
  • EventVLA
  • Fast-WAM
  • GO-1
  • GalaxeaVLA
  • GigaWorld-Policy-0
  • H-RDT
  • InternVLA-A1
  • LDA-1B
  • LingBot-VA
  • LingBot-VLA
  • Mem0
  • Spatial Forcing
  • Spirit v15
  • TinyVLA
  • X-VLA
  • X-WAM
  • Xiaomi-Robotics-0
  • Hy-Embodied-0.5-VLA
  • StarVLA-α
  • RISE