Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

IF 4.6 2区计算机科学 Q2 ROBOTICS IEEE Robotics and Automation Letters Pub Date : 2025-02-12 DOI:10.1109/LRA.2025.3541334

Guokang Wang;Hang Li;Shuyuan Zhang;Di Guo;Yanhong Liu;Huaping Liu

引用次数: 0

Abstract

In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this letter, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model. Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best-Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach enables the agent to reposition a third-person camera to actively observe the environment based on the task goal, and subsequently determine the appropriate manipulation actions. We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

求助全文

约1分钟内获得全文去求助

来源期刊

IEEE Robotics and Automation Letters Computer Science-Computer Science Applications

CiteScore

9.60

自引率

15.40%

发文量

1428

期刊介绍： The scope of this journal is to publish peer-reviewed articles that provide a timely and concise account of innovative research ideas and application results, reporting significant theoretical findings and application case studies in areas of robotics and automation.

期刊最新文献

Table of Contents IEEE Robotics and Automation Society Information IEEE Robotics and Automation Letters Information for Authors IEEE Robotics and Automation Society Information Image-Based Visual Servoing for Enhanced Cooperation of Dual-Arm Manipulation