TacO Benchmark: Which Tactile Sensor Is Best for Your Robot Manipulation Task?

TacO Benchmark: Which Tactile Sensor Is Best for Your Robot Manipulation Task?

6 min read•May 22, 2026•
James Okafor
James Okafor

The TacO benchmark tested four tactile sensor modalities — visual, acoustic, magnetic, and resistive — across three manipulation tasks and found that no single sensor is universally best. The optimal choice depends on the task type, material properties, and whether shear sensing or spatial resolution matters more. For buyers and engineers, this provides the first empirical framework to match sensor hardware to real-world automation needs.

Why Tactile Sensing Remains a Hard Problem for Manipulation

Vision-based learning from demonstration has enabled robots to pick objects and reason about scenes, but contact-rich manipulation — reorienting a screw, inserting a plug, lifting an unknown mass — demands physical feedback that cameras alone cannot provide. Tactile sensing fills that gap, yet most roboticists have had to guess which sensor type to buy. The TacO benchmark directly addresses this by systematically evaluating four distinct tactile modalities on the same set of manipulation tasks, controlling for policy architecture and environment.

Sensor modalities tested in the TacO benchmark — visual, acoustic, magnetic, and resistive

The core challenge lies in the trade-offs each modality makes. Visual tactile sensors (e.g., GelSight-style) offer high spatial resolution but struggle with shear forces and reflective surfaces. Magnetic sensors measure 3D force vectors directly but have lower spatial density. Acoustic sensors capture surface texture through vibrations but are sensitive to ambient noise. Resistive sensors are cheap and robust but provide sparse signals. No single technology covers all axes of performance.

How the TacO Benchmark Compared Four Sensor Modalities

The researchers trained separate reinforcement learning policies for each sensor modality, then tested them on three manipulation tasks using a Franka Emika Panda arm with custom sensorized grippers. Each task was designed to isolate specific tactile requirements:

  • Pick-and-place with unknown mass — tests force feedback and slip detection
  • Object reorientation — tests shear sensing and surface friction perception
  • Plug insertion — tests precise contact localization and alignment correction

Policies were trained in simulation with domain randomization and transferred to hardware. Key controlled variables included gripper geometry, object materials, and learning algorithm (SAC + tactile encoder).

ModalitySensing PrincipleKey StrengthBest TaskLimitation
VisualDeformation imaging via camera120×160 px resolutionReorientationFails on shiny surfaces
AcousticMicrophone + excitation signalTexture discriminationPick-and-placeAmbient noise sensitivity
MagneticHall effect sensors + mag film3D force vector measurementPlug insertion24-contact resolution
ResistiveConductive foam / ink pressure matLow cost, ruggedPick-and-place (simple)4×4 sparse array

The results confirmed that the same sensor architecture can perform very differently depending on task and material — a visual tactile sensor that achieves 92% success on reorientation of a metal block drops below 30% on the same task with a transparent plastic part.

Which Sensor Performed Best for Each Manipulation Task

Pick-and-place with unknown mass: Acoustic and magnetic sensors tied for top performance, achieving 89% and 91% success respectively, versus 72% for vision and 68% for resistive. The acoustic sensor’s ability to detect vibration from incipient slip gave it an advantage on low-friction objects. Magnetic sensors directly measured shear forces, allowing the policy to adjust grip force adaptively.

Task performance comparison for pick-and-place and plug insertion

Object reorientation: This task heavily favored visual tactile sensors, which succeeded 71% of the time compared to 55% for magnetic, 48% for acoustic, and 39% for resistive. High spatial resolution allowed the policy to track orientation changes at the contact patch. Notably, visual sensors benefited from the widest information bandwidth — a full greyscale image of the deformation surface — but only when the object surface was matte and non-transparent.

Plug insertion: Magnetic sensors dominated this alignment-critical task, with 79% success. Their ability to measure both normal and shear forces at 100 Hz enabled precise correction of insertion angle. Visual sensors suffered from occlusion during insertion, while acoustic and resistive sensors lacked the spatial resolution to distinguish partial from full insertion. Magnetic sensors also showed the lowest variance across material types — consistent performance from metal to plastic to rubber.

What This Means for Robotics Buyers — A Tactile Sensor Selection Guide

For engineers deploying manipulation cells, the TacO benchmark offers the first data-driven answer to “which tactile sensor should I buy?” The decision matrix depends on your primary task:

  • Pick-and-place with variable objects: Choose acoustic or magnetic sensors. Acoustic offers lower cost (estimated sensor bill of materials under $50) but requires a quiet environment. Magnetic sensors are more robust at $150–400 per fingertip.
  • Precision reorientation (assembly, kitting): Invest in visual tactile sensors like GelSight or DIGIT. Expect $500–$2,000 per sensor for the camera + gel setup, but they deliver the highest spatial resolution for tracking part orientation.
  • High-force insertion (connectors, bearings): Magnetic sensors are your best bet. Their shear sensitivity at high loads (up to 50 N) outperforms other modalities in alignment tasks.
Four sensor modalities compared side by side on a gripper

A key practical insight: material friction matters more than sensor resolution in many tasks. The benchmark found that for a given sensor modality, the coefficient of friction between gripper and object could swing success rates by as much as 35 percentage points. Buyers should evaluate sensors with their actual target materials — not benchmark-paper test objects — before committing.

For multi-purpose cells, a hybrid approach (e.g., visual + magnetic) may justify the higher cost, but the TacO data suggests you can often achieve 85%+ success with a single modality chosen for your dominant task. If your application mixes pick-and-place with occasional insertion, start with magnetic and add visual only if reorientation performance proves insufficient.

Browse used collaborative robots outfitted with tactile sensor modules on Robot Overflow to find ready-to-deploy solutions for contact-rich tasks.

Conclusion

The TacO benchmark provides the first empirical guidance for matching tactile sensor modalities to manipulation tasks, disproving the notion that more sensing is always better. For pick-and-place with variable parts, acoustic or magnetic sensors offer cost-effective slip detection. For precision reorientation, high-resolution visual sensors remain the leader. And for alignment-critical insertions, magnetic shear sensing delivers the most reliable performance across materials. Buyers can now make informed procurement decisions based on task priority rather than vendor claims.

Arizona appeals court vacates manslaughter sentence after AI video

An Arizona appeals court vacated the 10.5-year sentence of Gabriel Horcasitas while upholding his manslaughter conviction, first reported by Nytimes. The case returns to Maricopa County Superior Court for resentencing without the video, after judges found that it presented scripted statements as if the victim himself were speaking in court.

The three-judge panel said the video generated a likeness of Christopher Pelkey’s voice and appearance but did not reflect actual events. It found that allowing and relying on the video made the sentencing fundamentally unfair, and noted that no prior Arizona case had addressed the admissibility of such a depiction at sentencing.

The judges said a victim’s right to speak cannot override a defendant’s right to be sentenced on accurate, reliable information. They said the video collapsed the distinction between the family’s belief about what Pelkey would have said and Pelkey’s own voice and opinions.

The ruling distinguishes family members speaking about Pelkey from a generated likeness that appeared to speak for him.

Pelkey’s sister, Stacey Wales, presented the video during Horcasitas’s sentencing alongside victim-impact statements from family and friends. Wales wrote the script and said her husband and the couple’s longtime business partner helped create the video using Pelkey’s voice from a YouTube video and his face and torso from a funeral-service poster.

Judge Todd F. Lang praised the video as genuine, then imposed the maximum sentence of 10.5 years, more than the nine years prosecutors had sought.

Wales said nobody intended to make the court believe Pelkey was alive or that he had recorded the video before his death. She said she disagreed with the ruling and argued that families use slide shows, collages, hypothetical conversations and poetry to convey grief.

Wales compared the AI video with photography, saying it took 15 years of landmark cases around the 1860s before photography was widely accepted in courts.

The case returns to Maricopa County Superior Court for a new sentencing hearing without the AI-generated video.