The bottleneck
Robots learn to grip by watching people, and watching misses the force.
Internet video, headset video and teleoperation all capture what a grasp looks like. None of them capture how hard the hand squeezed, or the moment the object began to slip.
Average success on a real-world wiping task, using a robot policy trained only on human demonstrations recorded with a tactile glove. No robot data.
Nobody wears a sensor glove for eight hours a day for fun. People who need a hand wear one all day anyway.
Diego can feel what it holds.

- Grip force
- How hard the hand is gripping.
- Slip
- The moment an object starts to move in the grip.
- Grasp intent
- What the wearer asked the hand to do, from the two sensors.
How this compares with teleoperation.
| Teleoperation | Diego |
|---|---|
| An operator is paid for every hour. | The wearer uses the hand because they need it. |
| Each rig is bought and set up for the job. | The hand is sold to a clinic as a product. |
| The scene is arranged in advance. | It happens wherever the wearer goes, including when a grasp fails. |
- Phase one
Force and slip, all day.
- Phase two
The same, paired with what the wearer sees.
Robot learning wants vision and force from the same moment. We have not seen anyone offer both.
Diego has fewer joints than most robot hands.
Yes. A prosthetic hand has to be light, cheap and safe next to a person, so it has fewer degrees of freedom than a research hand.
Foundation models for manipulation already train across many different robot bodies, and retargeting a human demonstration onto a robot with a different hand is standard practice.
What a robot learns from Diego is contact: how hard the grip was, when the object slipped and what a failed grasp felt like. None of that depends on how many joints the hand has. Each generation of Diego will have more joints than the last.

