Learn to correct errors
Perturb successful source commands while keeping the desired motion fixed. This teaches feedback beyond a single nominal control sequence.
Dexterous manipulation · Shared residual learning
Execute new motions with one frozen policy. Adapt once for a whole collection.
ReDexT keeps one policy fixed. Each baseline fits the new trajectory.
No target updates
Frozen execution
Fit each trajectory
Execution after fitting
Fit each trajectory
Execution after fitting
Fit each trajectory
Execution after fitting
* Mean minutes per trajectory for initialization, target fitting and rollout. Offline IK and upstream training are excluded. Training uses H200 GPUs and evaluation uses RTX 4090 GPUs; workflows and concurrency differ.
The animation illustrates fitting; videos show the resulting executions.
Dexterous retargeting often requires fitting robot controls to each human hand–object motion, making large collections expensive to process. Offline inverse kinematics provides inexpensive nominal controls, but cannot correct errors caused by contact and object dynamics. We present ReDexT, a two-stage reinforcement-learning framework that learns one residual policy to execute new trajectories from inverse-kinematics commands. The key is to learn reusable feedback by varying both the commands the policy corrects and the motions it practices. We first perturb successful source commands while keeping their motion targets fixed, so the policy learns to correct varied execution errors. Rollout-guided training then broadens trajectory coverage through fresh on-policy learning, without treating rollout actions as imitation labels. Once trained, the shared policy executes new trajectories with frozen weights and supports optional local or shared adaptation for better tracking. On new trajectories screened for initial stability in single-hand simulation, frozen ReDexT achieves higher success and shorter processing times than the evaluated per-trajectory methods. Shared adaptation further raises success from 55.08% to 80.08% across 256 test trajectories. Ablations show that control perturbations ease transfer to kinematic commands, and expanded IK training broadens success coverage. The training recipe extends to four additional robot hands, each with a separately trained policy.
Inverse kinematics can map a human hand–object motion to a robot hand, but the resulting commands cannot react when the object moves or contact changes. Fitting a controller for each motion can improve execution, at a substantial cost across a large collection.
Our goal is to learn the feedback once and reuse it across new motions. ReDexT shares a residual policy across trajectories, while retaining each trajectory’s inexpensive IK commands.
Perturb successful source commands while keeping the desired motion fixed. This teaches feedback beyond a single nominal control sequence.
Use policy rollouts to expand the training set, then continue on-policy learning around the original IK commands.
Compute IK for a new motion and add the policy’s residual corrections. No target-specific weight updates are required.
Bootstrap → expand → execute
Scroll to explore the diagram →
One shared policy corrects the nominal controls at each step.
Training with control perturbations makes the policy less dependent on the source commands. When those commands are replaced by IK, the success-rate drop shrinks from 30.73 to 3.65 percentage points.
| Metric | Control noise | Source controls | IK controls | Drop (pp) |
|---|---|---|---|---|
| SR | Off | 89.06 | 58.33 | 30.73 |
| SR | On | 90.63 | 86.98 | 3.65 |
| HCR | Off | 82.81 | 40.63 | 42.19 |
| HCR | On | 81.25 | 72.92 | 8.33 |
Control perturbations ease the switch to IK commands, while broader IK training increases successful motion coverage. The paper presents the training details and ablations.
Select a motion to compare frozen ReDexT with the IK reference and trajectory-specific methods. The videos share a common clock; click a video to enlarge it.
Click a video to enlarge. Focus the player, then use Space to play/pause, ←/→ for 1/30 s, Shift + ←/→ for 1 s, and Home to restart. All methods share the IK comparison timestamps.
Loading videos…
ReDexT videos use frozen weights. The IK reference is a kinematic visualization; IK-only in the quantitative results is a physical rollout.
On the fixed 32-trajectory comparison subset, frozen ReDexT reaches 50.00% success, compared with 37.50% for SPIDER. Expensive per-trajectory baselines are evaluated on this smaller subset.
ReDexT takes 0.8 minutes per trajectory for initialization and rollout after IK. Processing times exclude offline IK and upstream training; workflows and concurrency differ across methods.
On the full 256-trajectory test set, shared adaptation raises success from 55.08% to 80.08% and horizon completion from 39.06% to 75.39%.
For a single target motion, adapting ReDexT is more effective than learning from scratch. With the same 5,000 target updates, ReDexT reaches 70.31% success, compared with 22.66% from random initialization.
Each of 256 target trajectories receives a separate policy and 5,000 updates. ReDexT and the bootstrap-only model start from shared pretraining; all three routes use the same target-learning budget. Curves show evaluations from update 501 onward.
Pretraining provides both immediate execution and a useful starting point for further learning. The paper reports per-group results and tracking quality.
Click a video to enlarge. Focus the player, then use Space to play/pause, ←/→ for 1/30 s, Shift + ←/→ for 1 s, and Home to restart. Shorter clips hold their last frame.
Loading videos…
31 motion examples across five hands. All recordings show frozen policies. For the four additional hands, frozen success exceeds physical replay of SPIDER controls by 8.20–13.67 percentage points under each hand’s scoring protocol.
Every hand benefits from shared adaptation, improving both mean-error success and completion of the full motion. The training recipe carries over while each hand learns its own policy.
| Robot hand | Frozen | Shared-adapted | ||
|---|---|---|---|---|
| SR (%) | HCR (%) | SR (%) | HCR (%) | |
| Sharpa | 55.08 | 39.06 | 80.08 | 75.39 |
| Allegro | 51.95 | 41.80 | 78.52 | 75.00 |
| Inspire | 44.92 | 27.73 | 67.19 | 65.63 |
| XHand | 50.00 | 35.55 | 73.83 | 68.36 |
| Shadow | 48.83 | 33.20 | 76.17 | 69.53 |
ReDexT shifts the work toward shared training: a frozen policy handles new motions, and optional adaptation refines a target collection together. This combines reusable feedback with lower processing cost per new trajectory.
The experiments study trajectories screened for initial stability in single-hand simulation, with a separate policy for each robot hand. Shared feedback across embodiments, bimanual manipulation and physical deployment remain open.
Read the full paper@misc{shi2026redext,
title = {ReDexT: Learning a Generalizable Residual Policy for Dexterous Retargeting},
author = {Shi, Jin-Chuan and Xu, Yangjinhui and Li, Liyang and Zhu, Muzhi and Hong, Jiadong and Hu, Yue and Chen, Hao and Shen, Chunhua},
year = {2026}
}