2026/01/01 by Tian Wang, Benhua Gao, Haofei Ma +3
Computer Science · Engineering · #Grippers #Hand Gesture Recognition Systems #Imitation #Modular Robots and Swarm Intelligence #Robot #Robot Manipulation and Learning #Robotic hand
paper · doi:10.1016/j.cirp.2026.04.055
published in CIRP Annals 75(1), 55-59 (Elsevier BV)
openalex publication_date 2026/01/01 · openalex created_date 2026/05/08 · openalex updated_date 2026/08/01
Increasing complexity and precision in Human–Robot Collaborative Assembly (HRCA) require robots to understand language, vision, and tactile information for contact-rich manipulation. To address the challenge, this paper proposes a vision-language conditioned physics-aware imitation learning approach for bimanual dexterous assembly. Firstly, a Mixed Reality (MR)-based bilateral teleoperation system is designed for multimodal human demo collection. Then, a Vision-Language Model (VLM)-augmented physics-aware diffusion policy is developed for manipulation skill learning. Furthermore, a coarse-to-fine visual Chain-of-Thought (CoT) strategy is integrated for task planning. Finally, the proposed method has been demonstrated on a dual-arm dexterous hand-based robotic platform by performing various HRCA tasks.