
Bimanual Manipulation Explained: Why Two Arms Are Harder Than One
I use ALOHA's cup-opening task to explain why bimanual manipulation demands coordinated motion, timing, force, training data, and feedback.
Browse humanoid robot and embodied AI training datasets. Explore source-backed articles and glossary definitions covering robot learning, hardware, control, and simulation.
Get new datasets, articles, and glossary entries in your inbox.

I use ALOHA's cup-opening task to explain why bimanual manipulation demands coordinated motion, timing, force, training data, and feedback.

A VLA can tell a robot what to do. One shoelace attempt shows why prediction before action, feedback during it, and learning after failure must work together.

Physical AI training data needs more than video: humanoid robots also need actions, touch, failures, body state, and evidence from the real machine.

An evidence-ranked guide to humanoid robot applications in factories, warehouses, homes, healthcare, services, hazardous work and space.

HiPHI releases 617.5 hours and 200.1 million frames of 90 Hz optical motion-capture data from 132 performers, comprising 308.7 hours of original capture paired with mirrored counterparts. Its standardized 55-joint BVH motions span 22 FrameNet frames and 214 Frame–LU labels with natural-language instructions; the 245.7-hour human-object interaction subset synchronizes body motion with trajectories and high-resolution meshes for 40 real objects across 12 categories.

RekaDaily-10k is an incrementally released corpus targeting 10,312 hours of unscripted first-person daily-life video recorded by paid collectors with head-mounted and handheld phones in homes and workplaces across multiple regions. The raw tier preserves sessions as collected and provides activity, lighting, duration, frame-rate, resolution, frame-count, codec, and salted collector metadata; Reka also announced a processed tier of short machine-captioned clips. Roughly 1,670 hours of the full corpus are native 4K.

AIRoA MoMa 5k contains 1,184,259 successful primitive-action episodes spanning 5,025.1 hours and 180,905,084 frames of teleoperated mobile manipulation with 44 Toyota Human Support Robots across five sites. The public, train-only LeRobot v3.0 release covers 68 main short-horizon task templates and combines head- and hand-camera RGB video with robot state and actions, end-effector pose, wrist force-torque history, low-level servo telemetry, and hierarchical task metadata.

The 10Kh RealOmni-Open Dataset contains more than 13,000 hours and 5 million clips of bimanual human demonstrations captured across over 10,000 real household scenarios. Collected from more than 3,000 contributors with GenDAS grippers, it covers 30 manipulation skills in 10 scenario groups, including clothing, clutter organization, kitchen cleaning, and shoe handling. The 95 TB MCAP release pairs 1600 × 1296 fisheye video at 30 fps with reconstructed end-effector trajectories, gripper state, 6-axis IMU readings, and tactile-array signals.