| Task 1 | Project benchmark Task 1 | high | Cross-Camera Temporal Handoff | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | Synchronized surround-camera video in which one tracked object leaves one camera and later appears in another. | Continue locating the same marked object after it crosses from one camera view into another. | required | existing benchmark protocol | nuScenes | | |
| Task 2 | Project benchmark Task 2 | high | Closed-Loop Object Adjustment | dynamic > active > execution > Manipulation | closed-loop separated egocentric camera streams + robot state | Synchronized head and wrist camera observations updated after every robot action. | Adjust the manipulated object to the target state, using each new multi-camera observation to correct the next action. | required | existing benchmark protocol | RoboTwin 2.0 | | |
| Task 3 | Absolute Depth | high | Absolute Depth | static > passive > perception > Metric Estimation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | What is the optical-axis depth of the red-marked object? Return one value in metres. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 4 | Absolute Direction | high | Absolute Direction | static > passive > perception > Pose Estimation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | In which surround sector is the marked object: front, front-left, front-right, back, back-left or back-right? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 5 | Absolute Size | high | Absolute Size | static > passive > perception > Metric Estimation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Estimate the marked object's length, width and height in metres. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 6 | Absolute Spatial Relation | high | Absolute Spatial Relation | static > passive > reasoning > Spatial Relation Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Which world-frame direction places the red-marked object relative to the blue-marked object: north, east, south or west? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 7 | Absolute Speed | high | Absolute Speed | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | What is the marked object's speed at the final timestamp? Return metres per second. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 8 | Acceleration | high | Acceleration | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | Across the clip, is the marked vehicle speeding up, slowing down or moving at approximately constant speed? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 9 | Active Exploration | high | Active Exploration | dynamic > active > planning > View Planning | closed-loop separated egocentric camera streams + agent state | RoboTwin 2.0: separated head, left-wrist and right-wrist RGB streams plus robot proprioception. The streams remain separated and update after every predicted action; no simulator ground truth is exposed to the model. | Select the next camera or viewpoint movement that will reveal whether the hidden object is behind the left or right partition. | required | simulator episode + state/success oracle | RoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/) | | Automatic scoring is available mainly in simulation; real multi-camera active datasets remain limited. |
| Task 10 | Behind-Camera Inference | high | Behind-Camera Inference | dynamic > passive > reasoning > Occlusion Inference | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | Which marked object is currently behind the wearer? Answer Red or Blue. | required | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 11 | Best-View Selection | high | Best-View Selection | static > passive > planning > View Planning | synchronized egocentric multi-view images | nuPlan: eight synchronized outward RGB cameras from one vehicle at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Which camera gives the least occluded view of the marked object? Answer with one camera name. | required | derived labels with manual geometric validation | nuPlan (https://www.nuscenes.org/nuplan) | nuScenes (https://www.nuscenes.org/nuscenes) | Assembly101 (https://assembly-101.github.io/) | | Ground truth in nuPlan, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 12 | Betweenness Relationship | high | Betweenness Relationship | static > passive > reasoning > Spatial Relation Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Is the green-marked object spatially between the red- and blue-marked objects in 3D? Answer Yes or No. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 13 | Camera Movement Direction | high | Camera Movement Direction | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | Did the wearer move mainly forward, backward, left or right during the clip? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 14 | Camera Orientation Estimation | high | Camera Orientation Estimation | static > passive > perception > Pose Estimation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | What is the FRONT camera's final world-frame heading? Return yaw in degrees. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 15 | Counting | high | Counting | static > passive > perception > Counting | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | How many visible mugs are present across all cameras after removing duplicate views of the same mug? | diagnostic control | source annotation + cross-view shortcut filtering | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 16 | Depth Consistency | high | Depth Consistency | static > passive > reasoning > Consistency Reasoning | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Are the two marked observations assigned a mutually consistent 3D depth? Answer Consistent or Inconsistent. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 17 | Depth Ordering | high | Depth Ordering | static > passive > reasoning > Occlusion Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Which marked object is closer to the wearer, Red or Blue? | required | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 18 | Ego Motion Reasoning | high | Ego Motion Reasoning | dynamic > passive > reasoning > Physical Reasoning | synchronized egocentric multi-view video | Aria Digital Twin: synchronized RGB, left-SLAM and right-SLAM video with camera labels and shared timestamps. A target object is marked at the first timestamp. | Did the marked object move in the world, or did only the wearer move? Answer Object moved or Wearer moved. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HoloSet (https://zenodo.org/records/7200131) | nuScenes (https://www.nuscenes.org/nuscenes) | Task 3: Object Motion vs Wearer Motion | Ground truth in Aria Digital Twin, HoloSet supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 19 | Object Displacement | high | Object Displacement | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | HOT3D: synchronized RGB, left-SLAM and right-SLAM video sampled at 1 fps over the same interval. The target is marked once at the beginning. | From the first to the final timestamp, did the marked object move closer to or farther from the wearer? | beneficial | deterministic derivation from geometry/pose/track annotations | HOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | Task 2: Object Distance Change | Ground truth in HOT3D, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 20 | Object Motion Reasoning | high | Object Motion Reasoning | dynamic > passive > reasoning > Physical Reasoning | synchronized egocentric multi-view video | Aria Digital Twin: synchronized RGB, left-SLAM and right-SLAM video with camera labels and shared timestamps. A target object is marked at the first timestamp. | Is the apparent image motion caused mainly by the marked object or by wearer motion? Answer Object or Wearer. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes) | Task 3: Object Motion vs Wearer Motion | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 21 | Object Rotation Direction | high | Object Rotation Direction | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | Did the marked object rotate clockwise or counter-clockwise in the world frame? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 22 | Physical Contact | high | Physical Contact | static > passive > reasoning > Physical Reasoning | synchronized egocentric multi-view images | HOT3D: three synchronized headset views (RGB, left monochrome SLAM and right monochrome SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Is the bottle physically touching the plate at the final timestamp? Answer Yes or No. | beneficial | derived labels with manual geometric validation | HOT3D (https://facebookresearch.github.io/hot3d/) | Assembly101 (https://assembly-101.github.io/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | | Ground truth in HOT3D, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 23 | Point Tracking | high | Point Tracking | dynamic > passive > perception > Motion Tracking | synchronized egocentric multi-view video | nuScenes: an 8-second six-camera surround-view clip sampled at 1 fps. A red point identifies the target surface point only at the first timestamp. | A red dot marks one surface point at the first timestamp. Return the camera name and pixel coordinate of the same 3D point at the final timestamp. | required | deterministic derivation from geometry/pose/track annotations | nuScenes (https://www.nuscenes.org/nuscenes) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | Task 1: Cross-Camera Temporal Handoff Localization | Ground truth in nuScenes, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 24 | Relative Depth | high | Relative Depth | static > passive > perception > Metric Estimation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Which marked object has greater optical depth from the FRONT camera, Red or Blue? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 25 | Relative Direction | high | Relative Direction | static > passive > reasoning > Spatial Relation Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Is the red-marked object left, right, above or below the blue-marked object in the wearer frame? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 26 | Relative Distance | high | Relative Distance | static > passive > perception > Metric Estimation | synchronized egocentric multi-view images | HOT3D: synchronized RGB, left-SLAM and right-SLAM video sampled at 1 fps over the same interval. The queried mug is marked in the first timestamp only. | At the final timestamp, is the marked mug closer to or farther from the wearer than at the first timestamp? Answer Closer or Farther. | beneficial | deterministic derivation from geometry/pose/track annotations | HOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | Task 2: Object Distance Change | Ground truth in HOT3D, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 27 | Revisit Counting | high | Revisit Counting | dynamic > passive > memory > Count Memory | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | How many distinct times does the wearer return to the marked place during the sequence? | required | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | HD-EPIC (https://hd-epic.github.io/site/) | | Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 28 | Rotation Angle | high | Rotation Angle | dynamic > passive > perception > Pose Estimation | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | Through what angle did the marked object rotate between the first and final timestamp? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131) | | Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 29 | Spatial Localization | high | Spatial Localization | static > passive > perception > Grounding | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Which camera and surround sector contain the marked object? | required | source annotation + cross-view shortcut filtering | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 30 | Travelled Path Length | high | Travelled Path Length | dynamic > passive > perception > Metric Estimation | synchronized egocentric multi-view video | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced. | How many metres did the wearer travel from the first to the final timestamp? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 31 | Vertical Spatial Relationships | high | Vertical Spatial Relationships | static > passive > reasoning > Spatial Relation Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | Is the red-marked object above, below or level with the blue-marked object in 3D? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 32 | Vertical World Position | high | Vertical World Position | static > passive > reasoning > Spatial Relation Inference | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | What is the marked object's height above the ground plane? Return metres. | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/) | | Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 33 | Viewpoint Apparent Size | high | Viewpoint Apparent Size | static > passive > reasoning > Perspective Transformation | synchronized egocentric multi-view images | Aria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden. | From which camera does the marked object appear largest? | beneficial | deterministic derivation from geometry/pose/track annotations | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | | Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review. |
| Task 34 | Viewpoint Control | high | Viewpoint Control | dynamic > active > execution > Viewpoint Control | closed-loop separated egocentric camera streams + agent state | RoboTwin 2.0: separated head, left-wrist and right-wrist RGB streams plus robot proprioception. The streams remain separated and update after every predicted action; no simulator ground truth is exposed to the model. | Rotate or move the agent until the marked object's front face is visible; output the next viewpoint-control action each step. | required | simulator episode + state/success oracle | RoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/) | | Automatic scoring is available mainly in simulation; real multi-camera active datasets remain limited. |