TaRL: Learning General and Physical Rewards from Tactile Demonstrations

The first approach of a learned tactile reward model

Po-Yi Wu1, Dao-Jan Chang1, Shang-Ya Hsiao2, Hong-Ming Chen2, Yu-Cheng Su2, Tsung-Wei Ke1
1National Taiwan University 2Delta Electronics

Tactile rewards generalize where visual rewards collapse. Predicted tactile and visual rewards along successful nut-threading rollouts at a seen object position (left) and an unseen one (right). Both reward models are trained on demonstrations from a single position; TaRL generalize better than visual reward model.

70%→81%
Average success in simulation
w/o TaRL → w/ TaRL
20%→59%
Average success in Real World
w/o TaRL → w/ TaRL
88%vs20%
Generalization Gap
reward kept at unseen positions
2.1×
Faster policy learning
iterations to reach baseline success

Abstract

Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task.

TaRL teaser: tactile observations are invariant to object position, capture local contact phases, and boost real-world RL.
Why touch? Left: tactile observations are invariant to varying object positions. Middle: tactile rewards capture local robot-object interaction that cannot be observed visually, so TaRL produces informative feedback at different contact phases (reach, grasp & move, insertion). Right: TaRL enhances RL efficiency on contact-rich manipulation in the real world.
Method

How TaRL works

TaRL models reward as task-completion progress, a parameterization that matches how tactile signals evolve through successful manipulation: from initial contact, to stable grasp, to directed force application.

Overview of TaRL: demonstrations, the tactile reward model, and downstream RL.
Overview of TaRL. (A) TaRL is trained with successful, rewound and failed tactile demonstrations. (B) TaRL conditions on tactile deformation maps and outputs task-progress rewards. (C) TaRL is applied to downstream RL, offering dense rewards.
A

Tactile demonstrations

Both successful and failed demonstrations are collected at a single object position: from intermediate RL checkpoints in simulation, or by teleoperation in the real world. Each rollout stores the binary outcome and the full sequence of 3D deformation maps from the two fingertip sensors.

B

Tactile reward model

A CNN encodes the tactile deformation maps observed so far, and a causal transformer regresses task progress from the sequence. The prediction is used directly as the reward.

C

Downstream RL

After offline training, TaRL is plugged into RL without fine-tuning. At every step it predicts a progress reward that is added to the environment's basic task reward, and optionally to a visual reward. Policies are trained with PPO online in simulation and with IQL offline in the real world.

Training objective

Following ReWiND, three sets of training examples are constructed from the demonstrations: (1) original tactile videos from successful demonstrations, regressed to progress \(t/T\); (2) rewound tactile videos, formed by concatenating a forward prefix with a reversed suffix at a random split point \(i\), regressed to \((i-t)/T\); and (3) tactile videos from failed demonstrations, regressed to 0. Rewinding lets the model produce decreasing rewards when the policy undoes progress, even though only successful and failed trajectories were recorded.

\[ \mathcal{L}(\theta) = \underbrace{\sum_{t=1}^{T} \tilde{R}_\theta(o^{\text{tac}}_{1:t})^2}_{\text{failed}} \;+\; \underbrace{\sum_{t=1}^{T} \Big(\tilde{R}_\theta(o^{\text{tac}}_{1:t}) - \tfrac{t}{T}\Big)^2}_{\text{successful}} \;+\; \underbrace{\mathbb{E}_{i,\,t}\Big[\Big(\tilde{R}_\theta\big([o^{\text{tac}}_{1:i};\, o^{\text{tac}}_{i-1:i-t}]\big) - \tfrac{i-t}{T}\Big)^2\Big]}_{\text{successful but rewound}} \]

Downstream RL

After offline training, TaRL is plugged into RL without fine-tuning. Its progress prediction is added to the environment's basic task reward, and optionally to a visual reward, as a dense shaping term:

\[ R_{\text{total}}(s_t, a_t, s_{t+1}) \;=\; R(s_t, a_t, s_{t+1}) \;+\; \alpha\, \tilde{R}_\theta\!\left(o^{\text{tac}}_{1:t}\right) \;+\; \beta\, \hat{R}_\phi\!\left(o_{1:t}\right) \]
tactile reward \(\tilde{R}_\theta\) from TaRL  ·  visual reward \(\hat{R}_\phi\) from ReWiND (Zhang et al., CoRL 2025)
Benchmark

Task suites

Four contact-rich tasks in simulation (Franka Panda with TacSL tactile simulation) and two real-world tasks (LeRobot SO-101 with PaXini PX-6AX-GEN1 tactile sensors). Unlike the original benchmarks, every episode starts with the object on the table, so policies must reach, grasp, and then insert, thread, or place.

Simulation

Box placement
Box placement
Peg insertion
Peg insertion
Gear assembly
Gear assembly
Nut threading
Nut threading

Real world

Cube pickup
Cube pickup
Peg insertion
Peg insertion

From a stable grasp and release (box placement) to tight-clearance alignment (peg insertion), multi-point meshing (gear assembly) and helical contact under continuous force (nut threading).

What TaRL sees

Tactile demonstrations

Successful demonstrations

Gear assembly
Nut threading
Peg insertion

Failed demonstrations

Gear assembly
Nut threading
Peg insertion

Deformation maps are simulated with TacSL; each arrow is the 3D displacement of one marker on the sensor gel.

Experiments

Experiment Results

Q1 Does TaRL improve downstream RL?

TaRL is added as a shaping reward on top of a basic task reward that has no terms for grasping or contact, and policies are evaluated at object positions not used for the demonstrations.

Box placement
Box placement
Peg insertion
Peg insertion
Gear assembly
Gear assembly
Nut threading
Nut threading
Downstream RL efficiency. Success rate over training iterations with and without TaRL.

Reward quality at unseen object positions

Predicted rewards on held-out successful and failed trajectories, at the demonstration position (ID) and at positions never seen by the reward model (OOD).

Box placement
Box placement
Peg insertion
Peg insertion
Gear assembly
Gear assembly
Nut threading
Nut threading
Generalization to unseen object positions. Predicted reward at each timestep, averaged over successful and failed trajectories at in-domain (ID) and out-of-distribution (OOD) positions.
Takeaway. TaRL improves both learning efficiency and performance on contact-rich manipulation tasks, and its reward generalizes to unseen object configurations

Q2 How does TaRL compare with a visual reward model?

TaRL is compared with ReWiND, a visual reward model with the same architecture and training recipe, on peg insertion, gear assembly and nut threading.

Tactile vs visual reward along successful rollouts at the training position and an unseen position.
Generalization to unseen object position. Predicted tactile and visual rewards along successful rollouts at the training position (left) and an unseen position (right). Visual model collapses at the unseen position, while TaRL does not.

Why touch generalizes

Peg insertion
Peg insertion
Gear assembly
Gear assembly
Nut threading
Nut threading
Tactile vs. visual features. t-SNE of CNN features from tactile and visual observations. Visual features at ID and OOD positions form distinct distributions, while tactile features share the same distribution: tactile data encodes local contact, not global scene appearance.

Vision and touch are informative at different phases

Visual and tactile rewards along a gear assembly trajectory.
Rewards at individual contact phases. On a successful gear-assembly trajectory, the visual reward plateaus as soon as the end-effector reaches the object, whereas the tactile reward keeps growing until task success.

Combining tactile and visual rewards

Peg insertion
Peg insertion
Gear assembly
Gear assembly
Nut threading
Nut threading
Complementary reward shaping. Downstream RL with the visual reward alone versus TaRL + visual. Combining both accelerates training.
Takeaway. TaRL generalizes to unseen object positions far better than the visual reward, and the two rewards are complementary: combining tactile and visual reward trains policies better.

Q3 Does TaRL generalize to novel object instances?

TaRL trained on box placement is deployed zero-shot to can placement, a different object with a different contact geometry, without any extra training.

Reward quality on can placement
Reward quality
Downstream RL on can placement
Downstream RL
Zero-shot generalization to unseen object instances. Left: predicted rewards still separate successful from failed trajectories. Right: the transferred reward improves downstream RL efficiency on the new task.
Takeaway. TaRL captures tactile patterns shared across object instances, enabling directly generalize to new objects.

Q4 Is a learned tactile reward necessary?

TaRL is compared with two alternatives: feeding tactile observations directly to the actor and critic, and a handcrafted tactile reward based on a grasp detector.

Gear assembly
Gear assembly
Peg insertion
Peg insertion
Nut threading
Nut threading
Tactile-based RL efficiency. Actors and critics take tactile observations as input (tactile in state) or not (baseline). In both cases, adding TaRL as a shaping reward still brings substantial gains to learning efficiency and task success.

Learned vs. handcrafted tactile rewards

Handcrafted tactile reward versus TaRL on nut threading.
Success rate with TaRL, handcrafted tactile reward, and baseline.

Handcrafted tactile reward. A tactile-based grasp detector gives a reward whenever the object is held.

Such rules are task-specific and coarse: once the object is grasped they cannot tell good behavior from bad, and every new task needs a new rule.

TaRL learns fine-grained, task-specific feedback from demonstrations and reaches high success far earlier.

Takeaway. Tactile inputs alone do not yield useful rewards, and handcrafted tactile rewards cannot provide fine-grained, task-specific signals across tasks. A learned tactile reward is both the most general and the most effective way to guide training.

Q5 Does TaRL enable real-world RL?

On a LeRobot SO-101, policies are trained offline with IQL on the same data, labeled with sparse success rewards only or additionally with TaRL's dense progress rewards.

Cube pickup
Cube pickup
Peg pickup
Peg pickup
Peg insertion
Peg insertion
Real-world RL efficiency. Success rate during inference with and without TaRL.
Takeaway. TaRL improves contact-rich manipulation tasks in real world.
Videos

Real-world rollouts

Cube pickup

Without TaRL
With TaRL

Peg insertion

Without TaRL
With TaRL

BibTeX

@inproceedings{wu2026tarl,
  title     = {{TaRL}: Learning General and Physical Rewards from Tactile Demonstrations},
  author    = {Wu, Po-Yi and Chang, Dao-Jan and Hsiao, Shang-Ya and Chen, Hong-Ming and Su, Yu-Cheng and Ke, Tsung-Wei},
  year      = {2026}
}