ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

CoRL 2026, Austin Texas, USA

Yushan Bai1,2,* Boyu Zheng1,2,* Zhiyang Mao4 Hongzheng Sun1,2 Yuchuang Tong1,2,† En Li1,2,3 Zhengtao Zhang1,2,3,† 1CAS Engineering Laboratory for Intelligent Industrial Vision, Institute of Automation, Chinese Academy of Sciences 2School of Artificial Intelligence, University of Chinese Academy of Sciences 3Beijing Zhongke Huiling Robot Technology Co., LTD. 4School of Intelligent Science and Technology, Xinjiang University *Equal contribution.    Corresponding authors.

We propose ProxiDex, a framework that integrates teleoperation with proximity-policy learning. The left panel shows immersive teleoperation with proximity feedback, while the right panel demonstrates robust performance in both simulation and real-world dexterous manipulation.

Abstract

Multi-finger dexterous manipulation relies on stable hand-object interactions, yet these interactions are partially observable in practice. Visual observations are often occluded by the hand, tactile sensors introduce hardware-specific modalities and calibration burdens, and existing policies rarely model how these cues evolve under actions, making them brittle under contact uncertainty. To address these, we present ProxiDex, a dynamics-guided proximity policy framework that treats hand-object proximity as an interaction state for dexterous manipulation. ProxiDex reconstructs interaction point clouds and converts geometric distances into proximity cues, forming a hardware-agnostic contact representation that provides immersive feedback during VR teleoperation. Built on this representation, ProxiDex learns action-conditioned proximity dynamics with a coupled forward–inverse design: future observation latents are predicted from actions, while proximity variations are decoded from latent changes. Leveraging these dynamics, ProxiDex adaptively reweights proximity tokens across manipulation phases and uses dynamics-consistency supervision to guide policy inference, stabilizing action generation under unreliable visual feedback. Simulation and real-world experiments demonstrate improved success rates and robustness over representative baselines across standard, unseen objects, and perturbation scenarios.

1. Framework

2. Teleoperation

Teleoperation in the simulator

Teleoperation in the real-world

3. Interaction Point Cloud Visualization

Benchmark: Adroit

Benchmark: DexArt

Self-Design Tasks in IsaacLab

Real-World

4. Results

Results in Simulation

Simulation results on dexterous manipulation benchmarks. We evaluate all methods on Adroit, DexArt, and three self-designed IsaacLab tasks.
Algorithm\Tasks Adroit DexArt Self-Designed Avg.
Hammer Door Pen Laptop Faucet Toilet Bucket Box Driller Banana
DP3 100.0±0.0 76.7±4.7 56.7±2.6 89.7±0.9 41.7±0.5 79.7±0.9 31.3±0.5 90.3±3.2 74.3±2.6 77.8±1.2 71.8±1.7
ManiFlow 100.0±0.0 80.3±1.2 55.5±5.8 93.0±1.6 45.0±3.6 79.9±3.3 35.3±2.1 94.7±2.1 78.6±1.4 79.8±2.3 74.2±2.3
AFRO 96.0±2.8 83.8±3.4 72.3±2.1 86.3±2.5 42.3±1.8 82.2±3.4 48.0±2.3 97.2±1.3 79.4±1.5 81.4±1.8 76.9±2.3
CordViP 96.0±2.1 84.7±1.7 76.2±3.2 92.6±2.1 56.5±1.3 81.0±2.2 57.0±1.6 96.3±2.0 85.3±1.1 84.7±1.6 81.0±1.9
ProxiDex (Ours) 100.0±0.0 86.5±1.9 84.3±2.4 93.0±1.2 53.4±2.1 84.3±2.3 64.0±3.2 96.7±1.6 89.2±2.0 87.3±1.5 83.9±1.8

Generalization to Unseen Objects

DP3
ManiFlow
AFRO
CordViP
ProxiDex
Pick
Pinch
Sweep

Robust to Clutter Scenarios

DP3
ManiFlow
AFRO
CordViP
ProxiDex
Pick
Pinch
Sweep

Robust to Visual Occlusion

DP3
ManiFlow
AFRO
CordViP
ProxiDex
Pick
Pinch
Sweep

Performance on Contact-Rich Tasks

DP3
ManiFlow
AFRO
CordViP
ProxiDex
Twist Cap
Flip Cap