T-Rex: Tactile-Reactive Dexterous Manipulation Using Robot Hands

T-Rex: Tactile-Reactive Dexterous Manipulation Using Robot Hands

Dantong Niu, Zhuoyang Liu, Zekai Wang, Boning Shao, Zhao-Heng Yin +30 more

5 min readJun 16, 2026

Human dexterity relies on more than vision; it depends fundamentally on the ability to feel and rapidly react to fine-grained tactile signals. While everyday tasks like sliding a thin card into a slot or opening a lock with a key are effortless for humans, they remain challenging for current robot learning policies. Mastering them requires tactile-reactive behaviors: immediate, closed-loop motor responses to tactile signals, far faster than conventional vision-based control loops allow.

Introduction

Human dexterity relies on more than vision; it depends fundamentally on the ability to feel and rapidly react to fine-grained tactile signals. While everyday tasks like sliding a thin card into a slot or opening a lock with a key are effortless for humans, they remain challenging for current robot learning policies. Mastering them requires tactile-reactive behaviors: immediate, closed-loop motor responses to tactile signals, far faster than conventional vision-based control loops allow.

Training Recipe

T-Rex is trained with a three-stage recipe that progressively transfers large-scale human visuomotor priors into tactile-reactive dexterous robot control.

Large-scale Human Egocentric Pre-training

Following EgoScale, we pre-train the latent and action experts on 22,889 hours of egocentric human video. The latent expert learns visual and language representations from head-view observations, while the action expert is trained on retargeted human arm and hand motions represented in a unified action space. This stage provides broad semantic grounding and visuomotor priors for dexterous manipulation without tactile expert.

Tactile Grounded Robot Mid-training

Large-scale human pre-training provides broad visuomotor priors but limited grounding in robot-executable contact dynamics. We bridge this gap with 100 hours of teleoperated bimanual manipulation data with synchronized tactile signals, organized around diverse motor primitives and object interactions for compact coverage of contact-rich behaviors. This stage adapts the action expert to robot multiview observations and executable actions, while training the tactile expert to perform high-frequency denoising as a fine-grained refinement.

Tactile sensor data visualization showing contact patterns during dexterous manipulation

Skill-Specific Post-training

After tactile-grounded mid-training, T-Rex already exhibits zero-shot contact-rich manipulation capabilities. For more complex or task-specific skills, we further fine-tune the model on approximately 100 task demonstrations, enabling it to adapt to specific task requirements while preserving the tactile-reactive behaviors acquired during mid-training.

Experiment Setup

Evaluation Protocol and Metrics

We evaluate all methods on the 12 tactile-reactive tasks defined in App. F. For each task, we test for 16 trials, with object positions and rotations randomized across trials. We report average task success rate, using progress-based rubrics for multi-stage tasks to capture partial completion. Results are averaged across trials and then across tasks.

Conclusion

We enable foundational manipulation policies to achieve scalable, tactile-reactive dexterous control. We introduce T-Rex, a Mixture-of-Transformer-Experts (MoT) model utilizing asynchronous tactile refinement and a dynamic tactile VAE encoding. Our framework leverages general human video pre-training, followed by mid-training on our newly contributed, open-source 100-hour tactile-synchronized dexterous manipulation dataset. Post-trained and evaluated across 12 real-world tactile-reactive tasks, T-Rex outperforms existing dexterous and tactile-aware VLA baselines by an average success rate of 30% and significantly improve data efficiency.

Robot hand performing a precision insertion task with tactile feedback visualization

Limitation and Future Work

While T-Rex demonstrates strong performance and data efficiency, it highlights several avenues for future research. First, for long-horizon tasks with precise contact coordination and tight tolerances where teleoperation is difficult, future work could integrate reinforcement learning or online interaction-based refinement. Second, tactile-reactive manipulation remains bottlenecked by hardware, including sensor distortion, calibration drift across devices, and the absence of dense palm sensing for whole-hand manipulation. Future work may explore unified representations across heterogeneous tactile sensors and richer, whole-hand tactile hardware.

We thank Sharpa for providing maintenance updates for their equipment. We also thank Yusuke Kato from Panasonic for his contributions to the collection of part of the T-Rex dataset. UC Berkeley authors were supported in part by the Berkeley Artificial Intelligence Research Humanoid Intelligence Center (BAIR HIC). Sapienza University acknowledges funding from Panasonic and from the Sapienza grant RG123188B3EF6A80 (CENTS). We thank Alessio Sampieri and Luca Franco (ItalAI S.r.l.) for fruitful discussions.

Frequently Asked Questions

What makes T-Rex different from existing robot manipulation policies? T-Rex uses tactile-reactive control with a Mixture-of-Transformer-Experts model that processes tactile signals faster than conventional vision-based loops, enabling real-time contact-rich manipulation.

How much training data does T-Rex require? T-Rex uses a three-stage training recipe: 22,889 hours of egocentric human video for pre-training, 100 hours of teleoperated bimanual data with tactile signals for mid-training, and approximately 100 task demonstrations for skill-specific fine-tuning.

What types of tasks can T-Rex perform? T-Rex was evaluated on 12 tactile-reactive tasks including precision insertion, key turning, and card sliding, all requiring fine-grained tactile feedback and rapid motor responses.

What are the main limitations of tactile-reactive manipulation? Key limitations include sensor distortion, calibration drift across devices, and the absence of dense palm sensing for whole-hand manipulation, which future work aims to address.

🍪 Cookie preferences

We use cookies to measure performance. Privacy Policy