Dexterous manipulation · Shared residual learning

ReDexTLearning a Generalizable Residual Policy
for Dexterous Retargeting

Jin-Chuan Shi*Yangjinhui Xu*Liyang LiMuzhi ZhuJiadong HongYue HuHao ChenChunhua Shen

State Key Lab of CAD & CG, Zhejiang University

* Equal contribution

Execute new motions with one frozen policy. Adapt once for a whole collection.

1 / 5
Compare this trajectory ↓

Shared feedback across trajectories

ReDexT keeps one policy fixed. Each baseline fits the new trajectory.

InputReference trajectory · Seq 01Object, wrist and fingertip motion

ReDexT

No target updates

πθ
Shared residual policySame frozen weights
Processing time*0.8 min/traj
IK + residualUses execution feedback

Frozen execution

SPIDER

Fit each trajectory

Refit for Seq 01
SimulateUpdate controls
Processing time*1.2 min/traj
u01∗Fitted controlsFor Seq 01

Execution after fitting

CHORD

Fit each trajectory

Refit for Seq 01
SimulateUpdate policy
Processing time*840.7 min/traj
π01∗Trained policyFor Seq 01

Execution after fitting

DexMachina

Fit each trajectory

Refit for Seq 01
SimulateUpdate policy
Processing time*2,653.2 min/traj
π01∗Trained policyFor Seq 01

Execution after fitting

Separate baseline fits

* Mean minutes per trajectory for initialization, target fitting and rollout. Offline IK and upstream training are excluded. Training uses H200 GPUs and evaluation uses RTX 4090 GPUs; workflows and concurrency differ.

The animation illustrates fitting; videos show the resulting executions.

One pretrained policy, with offline IK for each trajectory. Method and evaluation details
Full abstract

Dexterous retargeting often requires fitting robot controls to each human hand–object motion, making large collections expensive to process. Offline inverse kinematics provides inexpensive nominal controls, but cannot correct errors caused by contact and object dynamics. We present ReDexT, a two-stage reinforcement-learning framework that learns one residual policy to execute new trajectories from inverse-kinematics commands. The key is to learn reusable feedback by varying both the commands the policy corrects and the motions it practices. We first perturb successful source commands while keeping their motion targets fixed, so the policy learns to correct varied execution errors. Rollout-guided training then broadens trajectory coverage through fresh on-policy learning, without treating rollout actions as imitation labels. Once trained, the shared policy executes new trajectories with frozen weights and supports optional local or shared adaptation for better tracking. On new trajectories screened for initial stability in single-hand simulation, frozen ReDexT achieves higher success and shorter processing times than the evaluated per-trajectory methods. Shared adaptation further raises success from 55.08% to 80.08% across 256 test trajectories. Ablations show that control perturbations ease transfer to kinematic commands, and expanded IK training broadens success coverage. The training recipe extends to four additional robot hands, each with a separately trained policy.

Human motion gives us a target.
Contact makes execution hard.

Inverse kinematics can map a human hand–object motion to a robot hand, but the resulting commands cannot react when the object moves or contact changes. Fitting a controller for each motion can improve execution, at a substantial cost across a large collection.

Our goal is to learn the feedback once and reuse it across new motions. ReDexT shares a residual policy across trajectories, while retaining each trajectory’s inexpensive IK commands.

Method

One residual policy, many motions

ReDexT adds state-dependent corrections to IK commands. The policy sees the execution state and upcoming motion, so it can respond to contact and tracking errors while keeping its weights fixed across trajectories.

01 / Vary the controls

Learn to correct errors

Perturb successful source commands while keeping the desired motion fixed. This teaches feedback beyond a single nominal control sequence.

02 / Broaden the motions

Expand through execution

Use policy rollouts to expand the training set, then continue on-policy learning around the original IK commands.

03 / Reuse the policy

Execute with frozen weights

Compute IK for a new motion and add the policy’s residual corrections. No target-specific weight updates are required.

Training and execution
The complete workflow

Bootstrap → expand → execute

ReDexT: source-control bootstrap, rollout-guided task expansion, and shared residual control Successful SPIDER controls are perturbed for online PPO bootstrap. The policy then evaluates candidates with clean IK; seeds, current successes and up to 64 lowest-error failures define the next task set. Failure error is normalized by the success thresholds. Fresh rollouts and PPO use original IK controls with inherited noise. Solid arrows carry data; dashed blue arrows carry policy updates. Every rollout uses the same feedback controller. Deployment freezes the policy and disables noise and observation delay. Animation is a schematic, not experimental data. a Source-control bootstrap b Rollout-guided IK training c Shared residual controller b̂i,t update weights initialize expansion score warm start the next round at ut simulator state SPIDER seeds bisrc Noise curriculum adaptive control noise Shared rollout (c) sampled actions PPO update Initial policy θ0 Candidate tasks biIK Shared rollout (c) clean IK, mean action Next tasks seeds + current successes + lowest-error failures Fresh rollouts + PPO original IK + noise Updated policy θk+1 Active nominal b̂i,t Reference τi Observation oπi,t Residual policy πθ Bound + scale + Clip Physics

Scroll to explore the diagram →

One shared policy corrects the nominal controls at each step.

Data and control Policy updates. Schematic animation.

Learn feedback that survives a change in controls

Training with control perturbations makes the policy less dependent on the source commands. When those commands are replaced by IK, the success-rate drop shrinks from 30.73 to 3.65 percentage points.

MetricControl noiseSource controlsIK controlsDrop (pp)
SROff89.0658.3330.73
SROn90.6386.983.65
HCROff82.8140.6342.19
HCROn81.2572.928.33
Frozen bootstrap policies on 192 seed trajectories, using SPIDER source controls or IK at the ablation control scale. SR is mean-error success; HCR is horizon completion. Rates are percentages.

Control perturbations ease the switch to IK commands, while broader IK training increases successful motion coverage. The paper presents the training details and ablations.

Results

Execute new trajectories without refitting

Once trained, ReDexT uses the same frozen weights for every new motion. It achieves higher success than the evaluated per-trajectory methods, with less processing time per new trajectory.

Watch the same motions under different methods

Select a motion to compare frozen ReDexT with the IK reference and trajectory-specific methods. The videos share a common clock; click a video to enlarge it.

Higher success on new motions

On the fixed 32-trajectory comparison subset, frozen ReDexT reaches 50.00% success, compared with 37.50% for SPIDER. Expensive per-trajectory baselines are evaluated on this smaller subset.

Mean-error success rate (SR) on the same 32 trajectories. ReDexT requires no target-specific weight updates. ManipTrans† retains its native hand and simulator; baseline adaptations and evaluation protocols are detailed in the paper.

ReDexT takes 0.8 minutes per trajectory for initialization and rollout after IK. Processing times exclude offline IK and upstream training; workflows and concurrency differ across methods.

Adaptation

Adapt once for a collection of motions

Frozen execution provides a first pass. When higher tracking quality is needed, shared adaptation refines one policy on all target motions together.

On the full 256-trajectory test set, shared adaptation raises success from 55.08% to 80.08% and horizon completion from 39.06% to 75.39%.

FrozenShared-adapted
One policy jointly adapted on all 256 targets for 20,000 updates. SR measures mean tracking error; HCR additionally requires completion of the active horizon without an earlier termination signal. Adapted results describe performance on the target collection.

Pretraining speeds up local adaptation

For a single target motion, adapting ReDexT is more effective than learning from scratch. With the same 5,000 target updates, ReDexT reaches 70.31% success, compared with 22.66% from random initialization.

ReDexTBootstrap onlyFrom scratch

Each of 256 target trajectories receives a separate policy and 5,000 updates. ReDexT and the bootstrap-only model start from shared pretraining; all three routes use the same target-learning budget. Curves show evaluations from update 501 onward.

Pretraining provides both immediate execution and a useful starting point for further learning. The paper reports per-group results and tracking quality.

Five hands

The same recipe works across five robot hands

Beyond Sharpa, we apply the training recipe to Allegro, Inspire, XHand and Shadow. Each hand has its own separately trained policy, shared across that hand’s motion trajectories.

Consistent gains across five embodiments

Every hand benefits from shared adaptation, improving both mean-error success and completion of the full motion. The training recipe carries over while each hand learns its own policy.

Robot handFrozenShared-adapted
SR (%)HCR (%)SR (%)HCR (%)
Sharpa55.0839.0680.0875.39
Allegro51.9541.8078.5275.00
Inspire44.9227.7367.1965.63
XHand50.0035.5573.8368.36
Shadow48.8333.2076.1769.53
The same 256 OOD capture IDs are evaluated for each hand, with 20,000 shared adaptation updates. Policies are trained independently with hand-specific configurations, so these results demonstrate applicability rather than rank the hands.

Retarget a collection, not one motion at a time.

ReDexT shifts the work toward shared training: a frozen policy handles new motions, and optional adaptation refines a target collection together. This combines reusable feedback with lower processing cost per new trajectory.

The experiments study trajectories screened for initial stability in single-hand simulation, with a separate policy for each robot hand. Shared feedback across embodiments, bimanual manipulation and physical deployment remain open.

Read the full paper

Execution

Citation

@misc{shi2026redext,
  title = {ReDexT: Learning a Generalizable Residual Policy for Dexterous Retargeting},
  author = {Shi, Jin-Chuan and Xu, Yangjinhui and Li, Liyang and Zhu, Muzhi and Hong, Jiadong and Hu, Yue and Chen, Hao and Shen, Chunhua},
  year = {2026}
}