Vision
Primary View + Dual Wrist Views: 1080p @ 60 FPS
Hand Kinematics
Hand Joint Angles: 25 DoF per hand Fingertip Pose: 6 DoF per finger
Tactile
Touch points (per hand): 400+, minimum actuation force: 30 g
Spatial Pose
Wrist + Head Relative Pose
Real Human Operation Multimodal Data Acquisition Platform
Our platform captures high-quality multimodal human operation data by synchronizing first-person visuals, dual wrist-mounted camera feeds, hand kinematics, and tactile signals with precise spatiotemporal alignment.
Primary View + Dual Wrist Views: 1080p @ 60 FPS
Hand Joint Angles: 25 DoF per hand Fingertip Pose: 6 DoF per finger
Touch points (per hand): 400+, minimum actuation force: 30 g
Wrist + Head Relative Pose
High-precision hand skeletal tracking and head/wrist pose estimation enable accurate reconstruction of fine-grained hand manipulation.
Real Human Operation Data Collection Platform
Our platform captures high-fidelity first-person vision to generate realistic, continuous, and reproducible human operation data, fully equipped with automated semantic labeling.
3840×2880 @ 30 FPS monocular vision, with full camera intrinsics (enabling distortion correction)
frame-synchronized IMU motion pose data and natural audio
VLM-powered automated data preprocessing pipeline — semantic annotation + hand pose (camera coordinate frame)
complete causal relationships preserved in temporal sequences
Task List
Multimodal Synthetic Data Platform
Leveraging generative physical simulation technology, our platform automatically generates high-diversity synthetic data for arbitrary dual-arm manipulation tasks — including grasping, assembly, folding, and articulated object manipulation.
Our generative simulation engine enables fast synthetic manipulation data generation with broad scenario coverage — allowing flexible changes to objects, robot embodiments, and environments to produce more realistic and complex simulation data.
100,000+ rigid-body assets with manipulation annotations — including grasp points, manipulable parts, and physical properties — ready for task planning and policy training.
Multi-domain scene assets covering industrial manufacturing, modern logistics, and service applications, with on-demand customization to align with real-world business environments.
30+ mainstream robot embodiments — including tabletop dual-arm, mobile manipulation, and humanoid robots — with cross-embodiment data synthesis and evaluation, and freely swappable configurations on demand.
High-fidelity physical simulation is the core capability of X-Sim. By accurately reproducing key physical parameters such as material properties, lighting conditions, and friction, X-Sim ensures that simulation data can effectively transfer to real-world deployment. Based on this capability, X-Sim can infinitely generate diverse, fully annotated training data for challenging scenarios such as deformable object manipulation and fluid dynamics interactions, where traditional data collection is difficult, supporting policy transfer for complex tasks including grasping, assembly, folding, and articulation.
X-Sim supports the simulation and collection of multimodal perception data, including vision, tactile sensing, force sensing, and language. It covers the training input formats required by mainstream embodied AI models, providing comprehensive information for policy learning.
For personalized requirements beyond the coverage of standardized data products, we provide deep customization of high-fidelity scenarios.
AIGC-Powered Data Platform
Driven by AIGC-based video generation with an Agentic closed-loop quality assurance mechanism, our platform provides high-fidelity, high-diversity synthetic video data on demand for various physical manipulation tasks — covering both human and robotic operation modes.
Explore the X-Gen video generation platform now
Try It Now