DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception
2026-08-31 • Robotics
RoboticsComputer Vision and Pattern Recognition
AI summaryⓘ
The authors created a special dataset called DARP using two robot arms with cameras to look at objects on a table from different angles. Each arm has sensors that record color images, depth, and infrared data, along with the robot’s positions, to help understand where the cameras are in space. They tested how accurate their 3D scans were by combining different views into detailed models without using AI-based guessing, finding that most points were very close to the real surfaces. This dataset is meant to help researchers study how robots see objects better by using multiple views and sensor types together.
robotic perceptiondual-arm robotsRGB-D sensorseye-in-hand manipulationmulti-view fusionpoint clouds3D reconstructionsensor calibrationdepth imagingactive perception
Authors
Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz
Abstract
Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Robotic Perception) https://doi.org/10.21227/rmv3-be47, a calibrated dual-arm RGB-D-IR dataset for object-centered robotic perception using two independently moving eye-in-hand manipulators positioned on opposite sides of a shared tabletop workspace. Each arm carries an Intel RealSense sensor that continuously records RGB, depth, and stereo infrared data while synchronized robot joint states are logged for pose recovery. Objects are placed without fixed poses or marked locations, and the acquisition procedure performs automatic localization, cross-arm confirmation, adaptive viewpoint generation, and continuous multimodal recording. DARP contains ten unique tabletop objects and preserves the original sensor recordings, robot-state logs, object-level metadata, and calibration information required to reconstruct camera trajectories in a shared metric frame. To evaluate the geometric consistency of the acquisition, we implement a deterministic multi-view fusion pipeline that converts calibrated RGB-D observations into complementary partial point clouds and measured surface meshes without using learned or generative completion methods. Evaluation on 224 held-out RGB-D keyframes comprising 1,563,466 three-dimensional query points yields a median point-to-mesh distance of 2.13~mm and an RMSE of 4.04~mm, with 96.56\% of points within 10~mm of the measured-surface mesh. DARP is intended as a reusable resource for multi-view reconstruction, collaborative robotic perception, multimodal fusion, active perception, and future learning-based reasoning over partial object observations.