The first approach of a learned tactile reward model
Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task.
TaRL models reward as task-completion progress, a parameterization that matches how tactile signals evolve through successful manipulation: from initial contact, to stable grasp, to directed force application.
Both successful and failed demonstrations are collected at a single object position: from intermediate RL checkpoints in simulation, or by teleoperation in the real world. Each rollout stores the binary outcome and the full sequence of 3D deformation maps from the two fingertip sensors.
A CNN encodes the tactile deformation maps observed so far, and a causal transformer regresses task progress from the sequence. The prediction is used directly as the reward.
After offline training, TaRL is plugged into RL without fine-tuning. At every step it predicts a progress reward that is added to the environment's basic task reward, and optionally to a visual reward. Policies are trained with PPO online in simulation and with IQL offline in the real world.
Following ReWiND, three sets of training examples are constructed from the demonstrations: (1) original tactile videos from successful demonstrations, regressed to progress \(t/T\); (2) rewound tactile videos, formed by concatenating a forward prefix with a reversed suffix at a random split point \(i\), regressed to \((i-t)/T\); and (3) tactile videos from failed demonstrations, regressed to 0. Rewinding lets the model produce decreasing rewards when the policy undoes progress, even though only successful and failed trajectories were recorded.
After offline training, TaRL is plugged into RL without fine-tuning. Its progress prediction is added to the environment's basic task reward, and optionally to a visual reward, as a dense shaping term:
Four contact-rich tasks in simulation (Franka Panda with TacSL tactile simulation) and two real-world tasks (LeRobot SO-101 with PaXini PX-6AX-GEN1 tactile sensors). Unlike the original benchmarks, every episode starts with the object on the table, so policies must reach, grasp, and then insert, thread, or place.






From a stable grasp and release (box placement) to tight-clearance alignment (peg insertion), multi-point meshing (gear assembly) and helical contact under continuous force (nut threading).
Deformation maps are simulated with TacSL; each arrow is the 3D displacement of one marker on the sensor gel.
TaRL is added as a shaping reward on top of a basic task reward that has no terms for grasping or contact, and policies are evaluated at object positions not used for the demonstrations.




Predicted rewards on held-out successful and failed trajectories, at the demonstration position (ID) and at positions never seen by the reward model (OOD).




TaRL is compared with ReWiND, a visual reward model with the same architecture and training recipe, on peg insertion, gear assembly and nut threading.






TaRL trained on box placement is deployed zero-shot to can placement, a different object with a different contact geometry, without any extra training.


TaRL is compared with two alternatives: feeding tactile observations directly to the actor and critic, and a handcrafted tactile reward based on a grasp detector.



Handcrafted tactile reward. A tactile-based grasp detector gives a reward whenever the object is held.
Such rules are task-specific and coarse: once the object is grasped they cannot tell good behavior from bad, and every new task needs a new rule.
TaRL learns fine-grained, task-specific feedback from demonstrations and reaches high success far earlier.
On a LeRobot SO-101, policies are trained offline with IQL on the same data, labeled with sparse success rewards only or additionally with TaRL's dense progress rewards.



@inproceedings{wu2026tarl,
title = {{TaRL}: Learning General and Physical Rewards from Tactile Demonstrations},
author = {Wu, Po-Yi and Chang, Dao-Jan and Hsiao, Shang-Ya and Chen, Hong-Ming and Su, Yu-Cheng and Ke, Tsung-Wei},
year = {2026}
}