HumanTouch: A Multimodal System for Scalable Human-Hand Tactile Acquisition
Motion shows what a hand does; touch reveals how the physical world responds.
Overview
Motion shows what a hand does; touch reveals how the physical world responds. This connection is essential for robots to understand contact-rich interaction rather than merely imitate trajectories. Yet tactile data is useful only when real contact can be distinguished from sensor artifacts and drift. HumanTouch combines touch, hand motion, and vision within a calibrated and traceable acquisition system. It treats scalability not simply as recording more hours, but as collecting contact data that remains interpretable and trustworthy.
Why Robots Need Tactile Sensation
When humans manipulate objects, vision and touch provide distinct yet complementary information. Vision captures object appearance, spatial relationships, and global scene geometry, whereas touch directly measures the physical interaction between the hand and the object. Visual perception alone may not reliably determine whether, where, or how contact has been established, particularly when the contact region is occluded by the hand or when objects are transparent, reflective, compliant, or deformable. Moreover, visually similar hand poses can correspond to fundamentally different contact states, as illustrated in the accompanying video.
Touch is therefore not merely an auxiliary channel that supplements vision; it is a direct observation modality for physical interaction. The absence of tactile feedback limits their ability to assess contact establishment, interaction stability, and manipulation safety.
Nevertheless, equipping a robot with an arbitrary tactile sensor does not, by itself, resolve this limitation. Existing robotic tactile sensors vary substantially in sensing principles, form factors, spatial coverage, mounting configurations, and signal representations. Collecting tactile data directly on robots is also constrained by hardware cost, execution speed, operational safety, and the diversity of actions that a robot can reliably perform.
Humans, by contrast, can naturally and efficiently execute a broad range of contact-rich manipulation tasks. If human-hand tactile interactions can be captured reliably and transformed into consistent, robot-compatible representations, human demonstrations could provide a scalable and valuable source of data for robot tactile learning.
Touch, Force, and Proprioception Are Distinct
Touch, force, and proprioception describe related but distinct aspects of physical interaction. Cutaneous sensation arises from receptors in the skin and conveys information associated with contact, local pressure, vibration, skin deformation, and temperature. Proprioception, by contrast, provides information about the configuration and movement of the body through signals associated with muscles, tendons, and joints.
Human somatosensation comprises multiple sensory pathways. In this work, HumanTouch focuses on the contact-related aspects of hand–object interaction. It captures spatially distributed tactile responses across the hand surface, enabling the analysis of when and where contact occurs and how the relative intensity of interaction evolves over time.
These considerations raise two central questions:
Question 1: How can natural human-hand tactile interactions be captured without substantially constraining hand motion?
The sensing system must provide broad coverage of the fingers and palm while minimizing added weight, rigid components, and interference with natural hand movement. Tactile responses must also be synchronized with finger articulation and wrist pose so that contact-related activity can be localized on the moving hand and situated within the three-dimensional workspace.
Question 2: How can tactile data acquisition be scaled without compromising data quality?
Scalability requires more than accumulating recording hours. Data collected across devices, operators, tasks, sessions, and acquisition dates must remain valid, traceable, and interpretable under consistent criteria. Achieving this requires standardized data representations, repeatable calibration and synchronization procedures, continuous device-health monitoring, and explicit quality-screening rules.
Acquiring Tactile Data during Natural Hand Use
Tactile Sensing Technologies
Tactile sensors rely on different transduction principles, each offering distinct trade-offs in spatial resolution, form factor, mechanical compliance, and operating conditions.
Visuotactile sensors typically use an internal camera to observe markers, surface deformation, or illumination changes beneath a compliant contact surface. They can capture detailed local contact geometry at high spatial resolution. However, the camera, optical path, illumination system, and elastomer require substantial volume, making this approach difficult to extend over the entire surface of a human hand. Magnetic tactile sensors infer deformation from changes in the relative position or orientation of embedded magnets and magnetic-field sensors. They can provide sensitive multi-axis responses, but their performance depends on the mechanical structure, magnet–sensor arrangement, calibration procedure, and surrounding electromagnetic environment. Piezoresistive sensors measure changes in electrical resistance caused by mechanical pressure or deformation. Their thin, flexible construction allows them to be integrated into soft, large-area arrays, making them particularly suitable for wearable tactile gloves. Their signals, however, are affected by nonlinearity, hysteresis, drift, posture-dependent deformation, material aging, and variations in glove fit. Capacitive sensors detect contact-induced changes in capacitance, typically arising from variations in electrode spacing, effective area, or dielectric properties. They can offer high sensitivity and good repeatability, but large-area flexible implementations require careful management of electrode routing, parasitic capacitance, shielding, and manufacturing complexity.
HumanTouch adopts a flexible piezoresistive sensing architecture because it provides a practical balance among hand-surface coverage, mechanical compliance, wearability, and scalability. Its limitations motivate the posture- and history-aware calibration and quality-control procedures introduced later in this work.
Trade-offs between Flexibility and Rigidity
Any sensing layer placed between the human hand and an object inevitably alters the original interaction. The relevant question is therefore not whether the device introduces interference, but whether that interference is sufficiently limited and can be quantified experimentally.
Rigid fingertip sensors alter the geometry, friction, and compliance of the contact surface. Sparse finger caps preserve more of the uncovered hand surface but cannot capture contact on the palm, the sides of the fingers, or multiple uncovered regions. Handheld controllers measure interactions between the operator and the controller rather than direct hand–object interactions. Similarly, teleoperated manipulators introduce an embodiment gap between the human and robotic hands, as well as motion mapping, viewpoint transformation, and control latency. Exoskeletons and force-feedback gloves can provide active mechanical feedback, but their actuators, cables, and supporting structures add weight and may constrain natural joint motion.
Flexible tactile gloves offer a different trade-off. Their thin and compliant construction enables broad coverage of the fingers and palm while reducing the need for rigid structures. However, the glove remains part of the contact interface and may alter surface friction, compliance, tactile sensitivity, and object-handling strategies.
For systems designed to capture natural human manipulation, active force feedback is not essential because the objective is to observe interaction rather than mechanically guide the operator. Flexible tactile gloves conform to the fingers and palm, provide broad hand-surface coverage, and reduce the need for rigid structures that may restrict movement. Their soft and lightweight construction also makes them suitable for extended acquisition sessions.
Flexibility, however, introduces additional measurement challenges. Bending, stretching, wrinkling, and variations in glove fit can produce posture-dependent artifacts and reduce signal consistency.
Integrating Tactile Sensing with Hand Pose
Sensors embedded in a flexible glove move as the fingers bend and the fabric stretches, shifts, or wrinkles. Knowing only that “sensor 17 is activated” does not reliably indicate where the corresponding contact region lies in three-dimensional space. Nor does the tactile response alone provide the spatial orientation of that region. Tactile measurements must therefore be synchronized with finger-joint states and wrist pose to situate contact-related activity on the moving hand.
Vision-based hand tracking is a common solution, but tactile gloves obscure skin texture and alter the visual appearance of the hand. Fabric thickness, wrinkles, self-occlusion, and surface reflections can further reduce tracking reliability. As shown in Video 3, the VR hand-tracking system performs poorly when the operator wears a black glove. A skin-toned glove improves the qualitative tracking result by reducing the appearance difference from a bare hand, but its performance remains less reliable than bare-hand tracking. Changing glove color alone therefore cannot eliminate the effects of occlusion, altered hand geometry, and fabric deformation.
Wearable hand-pose tracking can be implemented using several sensing principles, among which stretch sensing, inertial measurement units (IMUs), and electromagnetic field (EMF) tracking are three representative approaches. Stretch sensors infer joint bending from deformation-dependent electrical changes, such as variations in resistance or capacitance. They can be thin and lightweight, but their outputs are sensitive to sensor placement, glove fit, material hysteresis, and repeated deformation. IMU-based systems estimate segment orientation and motion from measurements of angular velocity, linear acceleration, and, in some systems, the local magnetic field. They provide rapid dynamic responses but may accumulate orientation and position errors over time. EMF tracking estimates the relative position and orientation of sensors within a generated electromagnetic field. It does not rely on inertial integration and therefore avoids cumulative inertial drift, although nearby metallic objects and electromagnetic interference can distort its measurements.
HumanTouch uses a MANUS EMF-based hand-tracking system to record finger-joint states during manipulation. To reduce field distortion, the acquisition environment and task setup were arranged to keep metallic objects away from the tracking volume whenever practicable.
Scaling Tactile Data Acquisition
Multimodal Acquisition System
Tactile measurements must be interpreted together with hand motion and visual context. Each HumanTouch acquisition unit integrates bilateral flexible tactile gloves, bilateral finger-joint tracking, six-degree-of-freedom wrist tracking, left and right wrist-mounted cameras, a head-mounted scene camera, device-status monitoring, calibration procedures, and multimodal synchronization.
The wrist trackers provide the position and orientation of each hand within the tracking workspace, complementing the limited depth information available from first-person video. The wrist-mounted cameras capture close-range hand–object interactions that may be occluded or too small to observe clearly from the global viewpoint. The head-mounted camera provides a broader view of the workspace, objects, and task context. Camera calibration and tracking-system transformations establish spatial relationships among these viewpoints and reconstruct hands.
Table 1 — Comparison of HumanTouch with existing tactile data acquisition systems and datasets.
| System | Platform | Sensor Modalities | Hand Pose Tracking | Tactile Coverage | Multi-view Vision | Hours | Tasks | Temporal Resolution | Force Calibration | Natural Manipulation | Multi-site Deployment |
|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenTouch | Human hand | Tactile + Hand Skeleton + RGB | 7 × 6DoF sensors | Palm-side fingers and palm; 169 points/hand | Head RGB | 5.1 | - | 30 Hz | ✗ | ✓ | ✗ |
| World In Your Hands | Human hand | Tactile + Hand Skeleton + Wrist Pose + RGB | 6 × 6DoF sensors | Fingertips; 5 points/hand | Chest + wrist RGB | 1045 | 100 | - | ✗ | ✓ | ✗ |
| FreeTacMan | Robot gripper | Tactile + TCP Pose + Gripper Pose + RGB | TCP 6DoF | Two gripper contact surfaces | Wrist fisheye RGB | - | 50 | 30 Hz | ✗ | ✗ | ✗ |
| Touch in the Wild | Robot gripper | Tactile + TCP Pose + Gripper Pose + RGB | TCP 6DoF | Two gripper contact surfaces | Wrist fisheye RGB | - | 43 | 23 Hz | ✗ | ✗ | ✗ |
| HumanTouch (Ours) | Human hand | Tactile + Hand Skeleton + Wrist Pose + RGB | 25 joints | Fingers and palm; 360 points/hand | Head RGB + wrist RGB | 1000+ | 100+ | 60 Hz | ✓ | ✓ | ✓ |
Tactile-Assisted Hand-Pose Calibration
Stable joint-angle tracking alone does not guarantee that a reconstructed hand matches its wearer. Differences in palm size, finger-segment lengths, and glove fit can misalign the digital hand model with the physical hand. This is especially visible during self-contact: fingertips that touch in reality may remain separated in the model, or appear to intersect.
Tactile self-contact adds a useful geometric constraint. In a controlled calibration gesture, simultaneous contact detections on the thumb pad and another fingertip indicate that the corresponding regions of the hand should meet. The system collects multiple thumb-to-finger contact events and uses them to refine finger-segment lengths, joint offsets, and alignment between tactile sensing regions and the hand model.
Calibration uses contact onset rather than peak signal magnitude, so tactile signals do not need to be converted into physically calibrated force values. Regularization limits parameter updates and keeps the resulting hand geometry anatomically plausible, preventing the model from overfitting to a small set of contact observations.
Calibration and Representation of Flexible Tactile Gloves
Rigid tactile sensors have fixed sensing locations and surface geometry. Under controlled loading, their signals can often be calibrated against an applied force. Flexible tactile gloves are harder to interpret: fabric bends, stretches, shifts, and wrinkles as the hand moves, changing each sensing unit’s position and orientation. Finger motion can also deform piezoresistive material and create posture-dependent signals without external contact. These effects overlap with genuine contact signals.
Flexible piezoresistive sensors also exhibit hysteresis and history dependence. The same input can produce different responses while pressure is applied versus released, and recovery after sustained contact may be slow. Repeated use, temperature, glove fit, and wearing tension can shift the baseline. A static per-sensor calibration curve therefore cannot reliably provide physical force estimates during long, natural interactions.
Rather than mapping every sensing unit to one permanent 3D point, HumanTouch groups adjacent valid units into anatomically meaningful tactile patches. For each patch, it estimates contact confidence, the relative distribution of activity, and an approximate response center. This regional representation tolerates small fabric shifts better than pointwise mapping and transfers more readily across glove and robotic-hand configurations. The hand-pose model provides each patch’s current position, orientation, and surface normal. This converts hardware-specific channel readings into hand-region activity, approximate contact locations, relative response levels, and local surface geometry.
The calibration model combines three information sources: current tactile response, current hand pose and pose velocity, and recent tactile history. Pose features help identify signals caused by glove deformation without external contact. Temporal history helps distinguish rising, sustained, falling, and recovery phases, including effects of recent peaks and contact duration. Its outputs include contact confidence, an estimated response center within each patch, normalized relative contact magnitude, and estimation uncertainty. These values describe contact-related activity only. They are not calibrated pressure, force, or object-load measurements.
After tactile and pose calibration, contact-related activity can be projected onto the dynamic hand model. Contact regions can appear as points, local heat maps, or highlighted patches. Colors and numerical values show normalized relative response, while arrows—when used—show local surface normals, not measured force vectors. Video 6 illustrates how relative contact activity changes across the hand during ball pinching and object rotation.
Glove Lifecycle Monitoring and Quality Control
Flexible piezoresistive gloves may undergo irreversible signal drift over repeated use. In addition, the textile substrate and embedded conductive traces are subject to mechanical fatigue and breakage, potentially resulting in unstable or inactive sensing channels. A single calibration performed at the beginning of the project is therefore insufficient for long-term tactile data acquisition.
To track these changes throughout the glove lifecycle, each glove pair is assigned a unique identifier that links its calibration history, inspection results, recording sessions, and operational status. Before each acquisition day, every glove undergoes a standardized inspection that evaluates both channel availability and response consistency. This procedure enables the system to detect missing or unstable signals and prevents recordings from damaged gloves from entering the quality-qualified dataset.
Under the standardized inspection protocol, an value of at least 0.90 is one of the criteria for admission to the quality-qualified acquisition tier. A glove enters this tier only after passing the complete inspection procedure, including the channel-availability check. Failure to meet the threshold does not necessarily make the glove unusable for exploratory analysis or model development. However, recordings collected with that glove are excluded from the high-quality subset used in this work.
Across the observed glove lifecycle, inspection results do not follow a simple, monotonic trend with acquisition time. Short-term variation is more strongly associated with the tasks performed and the operators wearing the gloves. Tasks differ in contact location, response range, and repetition pattern; operators differ in glove fit, hand posture, contact strategy, and manipulation intensity. As a result, changes in inspection metrics or response statistics cannot be attributed to material degradation or cumulative use alone. Glove condition should be assessed through repeated standardized inspections, rather than inferred from recording duration.
Multi-Modal Data Synchronization
HumanTouch employs an offline multimodal synchronization framework built on a unified 60 Hz reference timeline. Because the sensing modalities operate at different native sampling rates and exhibit different temporal characteristics, all recorded timestamps are first converted into a common clock domain. For each acquisition session, the valid synchronization interval is defined as the temporal intersection of all required modalities. Samples outside this shared interval are excluded, preventing extrapolation beyond the observed data.
For each camera stream, the relationship between frame index and timestamp is modeled as
where is the recorded frame index, is the nominal frame period, is the temporal phase of the camera stream, and represents timestamp jitter and model residuals. The phase of each camera is estimated by least-squares fitting. A common reference phase is then selected to construct the shared timeline
Each recorded camera frame is associated with the corresponding reference timestamp without synthesizing intermediate images.
After the reference timeline has been established, each modality is resampled according to its data type. Positions, tactile responses, and other Euclidean quantities are linearly interpolated. Rotations are interpolated using quaternion spherical linear interpolation (SLERP), while discrete states and image frames are assigned using the nearest valid observation. Hand-joint states are matched by semantic joint identity rather than device-specific channel ordering, providing a consistent representation across acquisition devices. Streams containing excessive temporal gaps, missing required modalities, or incomplete tactile measurements are rejected rather than filled through long-range interpolation.
To support heterogeneous hardware, each acquisition unit maintains its own device identifiers, calibration records, storage metadata, and operational status. Device drivers expose standardized semantic streams—such as left-hand pose, right-wrist pose, and head-camera video—while the synchronization layer performs modality-independent temporal alignment. This modular design limits the effect of device replacement on downstream processing and supports the integration of heterogeneous sensing hardware.
Tactile Data Analysis
To evaluate the repeatability of task-related tactile patterns, we define the Dynamic Contact Signal-to-Noise Ratio (DcSNR). DcSNR measures the magnitude of a task’s mean dynamic contact pattern relative to the variability observed across repeated trials collected under matched conditions. Let denote the raw tactile time series of valid mapped point in file . Its within-file dynamic response amplitude is defined as
where and are the 90th and 10th temporal percentiles, respectively. This percentile range captures contact-related variation within a file while reducing sensitivity to isolated extreme samples. Normalization is performed independently for each hand. Let denote the hand containing point , and let contain the dynamic amplitudes of all valid mapped points on that hand. The normalized response is
If the hand-specific denominator is zero, the normalized responses of that hand are set to zero.
Using the predefined point-mapping coordinates, valid mapped points are projected onto a grid for each hand. When multiple points fall into the same grid cell, their normalized responses are averaged. Valid cells from the left and right hands are then concatenated to form the file-level dynamic contact vector . All files use the same fixed grid mask and cell ordering. Unmapped raw channels and unoccupied grid cells are excluded. For task , the mean dynamic contact pattern is
where is the set of included files for task , and . The signal magnitude is defined as
To estimate repeated-trial variability, files are grouped into matched condition units. Each group contains at least five files with the same operator, task, device, and canonical acquisition date. The mean pattern of group is
where is the number of files in the group. The pooled within-condition variability is then defined as
Finally, the DcSNR for task is
A higher DcSNR indicates that the task’s mean dynamic contact pattern is more prominent relative to repeated-trial variability under matched acquisition conditions.
Task-wise Dynamic Contact Patterns and DcSNR
A higher DcSNR indicates that a task’s mean dynamic contact pattern is more prominent relative to the matched repeated-trial variability used as the noise baseline. To characterize task-associated tactile responses, we selected approximately 1 h of eligible data for each of the ten canonical tasks.
The upper part of Figure shows the average dynamic contact distribution of both hands r the ten types of tasks. All tasks use a unified color scale, and the color represents the normalized relative contact response, not physical pressure or load. Different tasks show different response distributions in the palm surface, finger root and finger areas, indicating that there are distinguishable spatial contact patterns between tasks. Since each file is summarized as a dynamic contact vector, this figure does not reflect the contact sequence or stage changes within the task. The lower part shows the dynamic contact SNR of each task. The point estimates of the ten types of tasks are 3.61–7.19 dB, all higher than 0 dB, indicating that the tactile data of each task has certain information value.
Cross-Operator Temporal Stability of DcSNR
Tactile responses may vary across operators because of differences in hand geometry, contact location, and manipulation strategy. Consequently, cross-operator consistency should not be defined as pointwise agreement between tactile maps. We instead examine whether the strength of dynamic contact patterns remains stable over an extended acquisition period across operators.
The analysis includes 32 operators. For each operator, we selected one acquisition block with a fixed task, device, and canonical acquisition date. Eligible recordings within that block were ordered by acquisition time and partitioned into 60 one-minute intervals according to cumulative effective recording duration. Each operator contributed 1 h of eligible data, resulting in a total of 32 h for the analysis.
The figure shows the temporal evolution of DcSNR over the 60-min acquisition period. Thin lines represent the minute-level DcSNR trajectories of individual operators, the dark line represents the cross-operator median, and the shaded regions indicate the 25th–75th and 10th–90th percentile intervals. The median remains between 7.07 and 8.00 dB, with a temporal standard deviation of 0.19 dB, and exhibits no sustained downward trend. Under the current screening criteria, the strength of the dynamic contact patterns therefore remains broadly stable over the observed one-hour period.
HumanTouch Dataset
The initial HumanTouch release contains approximately 100 h of multimodal data selected from the ten canonical tasks, with approximately 10 h per task. Each episode was manually reviewed and annotated for task completion, and only episodes meeting the predefined quality criteria were included in the release. As summarized in Figure 8, the resulting subset comprises 13,469 episodes and exhibits substantial variation across tasks in episode count, duration, action complexity, and temporal organization.
And one more thing.
The AIGC edition of HumanTouch is also on its way.
Contributors
Project Leaders — Chuqiao Lyu
Corresponding Authors — Wenbo Ding, Tianxing Chen, Qi Xiong
Core Contributors — Chenze Yu, Eric J Chen, Wenxuan Zhu
Contributors — Youchen Lai, Raymond Lee, Daniel Li, Kai-Chong Lei, Kehe Ye, Jiahao Zhang, Yuxiao Huo, Xiwen Tan, Weining Lu, Gengqiu Yang, Junhao Gong, Chenxin Lang, Hekun Tian, Shaolong Yang.