01
Sim-to-Real Continuity
Unified policy and protocol to shorten the path from simulation results to real-robot validation
Customized Simulation & Real-Robot Evaluation Platform
Powered by generative physical simulation as its core engine, the platform delivers end-to-end solutions for robotics AI researchers and engineering teams — from data generation and scientific evaluation to scene asset management.
X-Eval supports designing evaluation tasks across key capability dimensions, with customizable combinations tailored to customer business needs. Typical evaluation areas include:
Run evaluations for multiple models simultaneously in large-scale scenarios, significantly boosting evaluation throughput and enabling teams to accelerate version comparison, regression testing, and parameter tuning cycles.
Eliminates evaluation pipeline fragmentation and bridges the sim-to-real gap, helping customers streamline the journey from post-training validation to pre-deployment acceptance while reducing deployment risks.
Real-robot evaluation media 1/4
01
Unified policy and protocol to shorten the path from simulation results to real-robot validation
02
Black background with soft lighting to prevent overexposure and ensure consistent visual conditions
03
Accurate object placement references for improved real-world scene reproduction
04
Task content, robot platforms, and scene configurations tailored to acceptance criteria
Every stage can be customized according to project requirements: standard packages enable rapid deployment, while dedicated solutions align with specific business needs.
Model rankings across tasks and capability dimensions
Centralized repository for evaluation data, versions, and task assets
Multi-dimensional comparison to support model selection and iterative improvement.
| Rank | Policy | Clean SR | Randomized SR | Retention | Performance gap | Completion | Overall |
|---|---|---|---|---|---|---|---|
| 01 | Pi_05 | 70.7% | 46.0% | 65.1% | 24.7pp | 100/100 | 55.9 |
| 02 | X_VLA | 68.0% | 20.9% | 30.7% | 47.1pp | 100/100 | 39.7 |
| 03 | X_WAM | 61.3% | 22.7% | 37.0% | 38.6pp | 100/100 | 38.1 |
| 04 | Abot_M0 | 57.4% | 22.9% | 39.9% | 34.5pp | 100/100 | 36.7 |
| 05 | Spatial_Forcing | 77.2% | 9.5% | 12.3% | 67.7pp | 100/100 | 36.6 |
| 06 | Xiaomi_Robotics_0 | 62.9% | 18.2% | 28.9% | 44.7pp | 100/100 | 36.1 |
| 07 | EventVLA | 65.6% | 15.7% | 23.9% | 49.9pp | 100/100 | 35.7 |
| 08 | FastWAM | 77.8% | 1.9% | 2.4% | 75.9pp | 100/100 | 32.3 |
| 09 | GalaxeaVLA | 62.7% | 9.1% | 14.5% | 53.6pp | 100/100 | 30.5 |
| 10 | AHA_WAM | 64.3% | 3.2% | 5.0% | 61.1pp | 100/100 | 27.6 |
| 11 | starVLA | 33.1% | 2.8% | 8.5% | 30.3pp | 100/100 | 14.9 |
Overall = 40% Clean SR + 60% Randomized SR · Retention = Randomized SR ÷ Clean SR · pp means percentage points
Integrated with 40+ state-of-the-art models, the platform provides the industry's largest collection of reproduced models and continues to expand its ecosystem partnerships, enabling customers to seamlessly integrate and compare models on a unified platform.