Egocentric Multi-View Benchmark Task Taxonomy

  • Egocentric Multi-View Spatial Benchmark
    • dynamic
      • active
        • execution
          • Manipulation
            • Task 2 · Closed-Loop Object Adjustment
              Datasets: RoboTwin 2.0
          • Viewpoint Control
            • Task 34 · Viewpoint Control
              Datasets: RoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/)
        • planning
          • View Planning
            • Task 9 · Active Exploration
              Datasets: RoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/)
      • passive
        • memory
          • Count Memory
            • Task 27 · Revisit Counting
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | HD-EPIC (https://hd-epic.github.io/site/)
        • perception
          • Metric Estimation
            • Task 30 · Travelled Path Length
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
          • Motion Tracking
            • Task 1 · Cross-Camera Temporal Handoff
              Datasets: nuScenes
            • Task 13 · Camera Movement Direction
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)
            • Task 19 · Object Displacement
              Datasets: HOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes)
            • Task 21 · Object Rotation Direction
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)
            • Task 23 · Point Tracking
              Datasets: nuScenes (https://www.nuscenes.org/nuscenes) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
            • Task 7 · Absolute Speed
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)
            • Task 8 · Acceleration
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)
          • Pose Estimation
            • Task 28 · Rotation Angle
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)
        • reasoning
          • Occlusion Inference
            • Task 10 · Behind-Camera Inference
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)
          • Physical Reasoning
            • Task 18 · Ego Motion Reasoning
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HoloSet (https://zenodo.org/records/7200131) | nuScenes (https://www.nuscenes.org/nuscenes)
            • Task 20 · Object Motion Reasoning
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)
    • static
      • passive
        • perception
          • Counting
            • Task 15 · Counting
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
          • Grounding
            • Task 29 · Spatial Localization
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
          • Metric Estimation
            • Task 24 · Relative Depth
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
            • Task 26 · Relative Distance
              Datasets: HOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
            • Task 3 · Absolute Depth
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
            • Task 5 · Absolute Size
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
          • Pose Estimation
            • Task 14 · Camera Orientation Estimation
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)
            • Task 4 · Absolute Direction
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)
        • planning
          • View Planning
            • Task 11 · Best-View Selection
              Datasets: nuPlan (https://www.nuscenes.org/nuplan) | nuScenes (https://www.nuscenes.org/nuscenes) | Assembly101 (https://assembly-101.github.io/)
        • reasoning
          • Consistency Reasoning
            • Task 16 · Depth Consistency
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
          • Occlusion Inference
            • Task 17 · Depth Ordering
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)
          • Perspective Transformation
            • Task 33 · Viewpoint Apparent Size
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)
          • Physical Reasoning
            • Task 22 · Physical Contact
              Datasets: HOT3D (https://facebookresearch.github.io/hot3d/) | Assembly101 (https://assembly-101.github.io/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/)
          • Spatial Relation Inference
            • Task 12 · Betweenness Relationship
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
            • Task 25 · Relative Direction
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
            • Task 31 · Vertical Spatial Relationships
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
            • Task 32 · Vertical World Position
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)
            • Task 6 · Absolute Spatial Relation
              Datasets: Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)

Approved tasks

IDSource normalized taskVerdictProposed taskTarget pathInput formExample model inputExample prompt / instructionMulti-view roleConstructionCandidate datasetsCurrent task linkLimitation
Task 1Project benchmark Task 1highCross-Camera Temporal Handoffdynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoSynchronized surround-camera video in which one tracked object leaves one camera and later appears in another.Continue locating the same marked object after it crosses from one camera view into another.requiredexisting benchmark protocolnuScenes
Task 2Project benchmark Task 2highClosed-Loop Object Adjustmentdynamic > active > execution > Manipulationclosed-loop separated egocentric camera streams + robot stateSynchronized head and wrist camera observations updated after every robot action.Adjust the manipulated object to the target state, using each new multi-camera observation to correct the next action.requiredexisting benchmark protocolRoboTwin 2.0
Task 3Absolute DepthhighAbsolute Depthstatic > passive > perception > Metric Estimationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.What is the optical-axis depth of the red-marked object? Return one value in metres.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 4Absolute DirectionhighAbsolute Directionstatic > passive > perception > Pose Estimationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.In which surround sector is the marked object: front, front-left, front-right, back, back-left or back-right?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review.
Task 5Absolute SizehighAbsolute Sizestatic > passive > perception > Metric Estimationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Estimate the marked object's length, width and height in metres.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 6Absolute Spatial RelationhighAbsolute Spatial Relationstatic > passive > reasoning > Spatial Relation Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Which world-frame direction places the red-marked object relative to the blue-marked object: north, east, south or west?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 7Absolute SpeedhighAbsolute Speeddynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.What is the marked object's speed at the final timestamp? Return metres per second.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 8AccelerationhighAccelerationdynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.Across the clip, is the marked vehicle speeding up, slowing down or moving at approximately constant speed?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 9Active ExplorationhighActive Explorationdynamic > active > planning > View Planningclosed-loop separated egocentric camera streams + agent stateRoboTwin 2.0: separated head, left-wrist and right-wrist RGB streams plus robot proprioception. The streams remain separated and update after every predicted action; no simulator ground truth is exposed to the model.Select the next camera or viewpoint movement that will reveal whether the hidden object is behind the left or right partition.requiredsimulator episode + state/success oracleRoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/)Automatic scoring is available mainly in simulation; real multi-camera active datasets remain limited.
Task 10Behind-Camera InferencehighBehind-Camera Inferencedynamic > passive > reasoning > Occlusion Inferencesynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.Which marked object is currently behind the wearer? Answer Red or Blue.requireddeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 11Best-View SelectionhighBest-View Selectionstatic > passive > planning > View Planningsynchronized egocentric multi-view imagesnuPlan: eight synchronized outward RGB cameras from one vehicle at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Which camera gives the least occluded view of the marked object? Answer with one camera name.requiredderived labels with manual geometric validationnuPlan (https://www.nuscenes.org/nuplan) | nuScenes (https://www.nuscenes.org/nuscenes) | Assembly101 (https://assembly-101.github.io/)Ground truth in nuPlan, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 12Betweenness RelationshiphighBetweenness Relationshipstatic > passive > reasoning > Spatial Relation Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Is the green-marked object spatially between the red- and blue-marked objects in 3D? Answer Yes or No.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 13Camera Movement DirectionhighCamera Movement Directiondynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.Did the wearer move mainly forward, backward, left or right during the clip?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 14Camera Orientation EstimationhighCamera Orientation Estimationstatic > passive > perception > Pose Estimationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.What is the FRONT camera's final world-frame heading? Return yaw in degrees.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review.
Task 15CountinghighCountingstatic > passive > perception > Countingsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.How many visible mugs are present across all cameras after removing duplicate views of the same mug?diagnostic controlsource annotation + cross-view shortcut filteringAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review.
Task 16Depth ConsistencyhighDepth Consistencystatic > passive > reasoning > Consistency Reasoningsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Are the two marked observations assigned a mutually consistent 3D depth? Answer Consistent or Inconsistent.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 17Depth OrderinghighDepth Orderingstatic > passive > reasoning > Occlusion Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Which marked object is closer to the wearer, Red or Blue?requireddeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 18Ego Motion ReasoninghighEgo Motion Reasoningdynamic > passive > reasoning > Physical Reasoningsynchronized egocentric multi-view videoAria Digital Twin: synchronized RGB, left-SLAM and right-SLAM video with camera labels and shared timestamps. A target object is marked at the first timestamp.Did the marked object move in the world, or did only the wearer move? Answer Object moved or Wearer moved.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HoloSet (https://zenodo.org/records/7200131) | nuScenes (https://www.nuscenes.org/nuscenes)Task 3: Object Motion vs Wearer MotionGround truth in Aria Digital Twin, HoloSet supports automatic generation, but samples still need single-view-necessity and wording review.
Task 19Object DisplacementhighObject Displacementdynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoHOT3D: synchronized RGB, left-SLAM and right-SLAM video sampled at 1 fps over the same interval. The target is marked once at the beginning.From the first to the final timestamp, did the marked object move closer to or farther from the wearer?beneficialdeterministic derivation from geometry/pose/track annotationsHOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes)Task 2: Object Distance ChangeGround truth in HOT3D, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review.
Task 20Object Motion ReasoninghighObject Motion Reasoningdynamic > passive > reasoning > Physical Reasoningsynchronized egocentric multi-view videoAria Digital Twin: synchronized RGB, left-SLAM and right-SLAM video with camera labels and shared timestamps. A target object is marked at the first timestamp.Is the apparent image motion caused mainly by the marked object or by wearer motion? Answer Object or Wearer.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | nuScenes (https://www.nuscenes.org/nuscenes)Task 3: Object Motion vs Wearer MotionGround truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 21Object Rotation DirectionhighObject Rotation Directiondynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.Did the marked object rotate clockwise or counter-clockwise in the world frame?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 22Physical ContacthighPhysical Contactstatic > passive > reasoning > Physical Reasoningsynchronized egocentric multi-view imagesHOT3D: three synchronized headset views (RGB, left monochrome SLAM and right monochrome SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Is the bottle physically touching the plate at the final timestamp? Answer Yes or No.beneficialderived labels with manual geometric validationHOT3D (https://facebookresearch.github.io/hot3d/) | Assembly101 (https://assembly-101.github.io/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/)Ground truth in HOT3D, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review.
Task 23Point TrackinghighPoint Trackingdynamic > passive > perception > Motion Trackingsynchronized egocentric multi-view videonuScenes: an 8-second six-camera surround-view clip sampled at 1 fps. A red point identifies the target surface point only at the first timestamp.A red dot marks one surface point at the first timestamp. Return the camera name and pixel coordinate of the same 3D point at the final timestamp.requireddeterministic derivation from geometry/pose/track annotationsnuScenes (https://www.nuscenes.org/nuscenes) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Task 1: Cross-Camera Temporal Handoff LocalizationGround truth in nuScenes, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review.
Task 24Relative DepthhighRelative Depthstatic > passive > perception > Metric Estimationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Which marked object has greater optical depth from the FRONT camera, Red or Blue?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 25Relative DirectionhighRelative Directionstatic > passive > reasoning > Spatial Relation Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Is the red-marked object left, right, above or below the blue-marked object in the wearer frame?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 26Relative DistancehighRelative Distancestatic > passive > perception > Metric Estimationsynchronized egocentric multi-view imagesHOT3D: synchronized RGB, left-SLAM and right-SLAM video sampled at 1 fps over the same interval. The queried mug is marked in the first timestamp only.At the final timestamp, is the marked mug closer to or farther from the wearer than at the first timestamp? Answer Closer or Farther.beneficialdeterministic derivation from geometry/pose/track annotationsHOT3D (https://facebookresearch.github.io/hot3d/) | Aria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Task 2: Object Distance ChangeGround truth in HOT3D, Aria Digital Twin supports automatic generation, but samples still need single-view-necessity and wording review.
Task 27Revisit CountinghighRevisit Countingdynamic > passive > memory > Count Memorysynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.How many distinct times does the wearer return to the marked place during the sequence?requireddeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | HD-EPIC (https://hd-epic.github.io/site/)Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review.
Task 28Rotation AnglehighRotation Angledynamic > passive > perception > Pose Estimationsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.Through what angle did the marked object rotate between the first and final timestamp?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K) | HoloSet (https://zenodo.org/records/7200131)Ground truth in Aria Digital Twin, Ego-1K supports automatic generation, but samples still need single-view-necessity and wording review.
Task 29Spatial LocalizationhighSpatial Localizationstatic > passive > perception > Groundingsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Which camera and surround sector contain the marked object?requiredsource annotation + cross-view shortcut filteringAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | Assembly101 (https://assembly-101.github.io/) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, Assembly101 supports automatic generation, but samples still need single-view-necessity and wording review.
Task 30Travelled Path LengthhighTravelled Path Lengthdynamic > passive > perception > Metric Estimationsynchronized egocentric multi-view videoAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM), sampled over one 8-second interval at 1 fps. Camera names and shared timestamps are shown; a queried target is marked only when first introduced.How many metres did the wearer travel from the first to the final timestamp?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 31Vertical Spatial RelationshipshighVertical Spatial Relationshipsstatic > passive > reasoning > Spatial Relation Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.Is the red-marked object above, below or level with the blue-marked object in 3D?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 32Vertical World PositionhighVertical World Positionstatic > passive > reasoning > Spatial Relation Inferencesynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.What is the marked object's height above the ground plane? Return metres.beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | nuScenes (https://www.nuscenes.org/nuscenes) | JRDB (https://jrdb.erc.monash.edu.au/dataset/)Ground truth in Aria Digital Twin, nuScenes supports automatic generation, but samples still need single-view-necessity and wording review.
Task 33Viewpoint Apparent SizehighViewpoint Apparent Sizestatic > passive > reasoning > Perspective Transformationsynchronized egocentric multi-view imagesAria Digital Twin: three synchronized Aria views (RGB, left SLAM and right SLAM) at one shared timestamp. Each view is labeled; queried objects or points are color-marked, while calibration-derived answer labels remain hidden.From which camera does the marked object appear largest?beneficialdeterministic derivation from geometry/pose/track annotationsAria Digital Twin (https://facebookresearch.github.io/projectaria_tools/docs/open_datasets/aria_digital_twin_dataset/) | HOT3D (https://facebookresearch.github.io/hot3d/) | Ego-1K (https://huggingface.co/datasets/jaeyounglee/Ego-1K)Ground truth in Aria Digital Twin, HOT3D supports automatic generation, but samples still need single-view-necessity and wording review.
Task 34Viewpoint ControlhighViewpoint Controldynamic > active > execution > Viewpoint Controlclosed-loop separated egocentric camera streams + agent stateRoboTwin 2.0: separated head, left-wrist and right-wrist RGB streams plus robot proprioception. The streams remain separated and update after every predicted action; no simulator ground truth is exposed to the model.Rotate or move the agent until the marked object's front face is visible; output the next viewpoint-control action each step.requiredsimulator episode + state/success oracleRoboTwin 2.0 (https://robotwin-platform.github.io/) | Habitat + HM3D (https://aihabitat.org/datasets/hm3d/)Automatic scoring is available mainly in simulation; real multi-camera active datasets remain limited.