The bottleneck

Robots learn to grip by watching people, and watching misses the force.

Internet video, headset video and teleoperation all capture what a grasp looks like. None of them capture how hard the hand squeezed, or the moment the object began to slip.

Average success on a real-world wiping task, using a robot policy trained only on human demonstrations recorded with a tactile glove. No robot data.

Nobody wears a sensor glove for eight hours a day for fun. People who need a hand wear one all day anyway.

Diego can feel what it holds.

A line drawing of the Diego hand with leader lines to grip force, slip and grasp intent.
Grip force
How hard the hand is gripping.
Slip
The moment an object starts to move in the grip.
Grasp intent
What the wearer asked the hand to do, from the two sensors.

How this compares with teleoperation.

TeleoperationDiego
An operator is paid for every hour.The wearer uses the hand because they need it.
Each rig is bought and set up for the job.The hand is sold to a clinic as a product.
The scene is arranged in advance.It happens wherever the wearer goes, including when a grasp fails.
Phase one

Force and slip, all day.

Phase two

The same, paired with what the wearer sees.

Robot learning wants vision and force from the same moment. We have not seen anyone offer both.

Diego has fewer joints than most robot hands.

Yes. A prosthetic hand has to be light, cheap and safe next to a person, so it has fewer degrees of freedom than a research hand.

Foundation models for manipulation already train across many different robot bodies, and retargeting a human demonstration onto a robot with a different hand is standard practice.

What a robot learns from Diego is contact: how hard the grip was, when the object slipped and what a failed grasp felt like. None of that depends on how many joints the hand has. Each generation of Diego will have more joints than the last.

We want to hear from robotics teams.