Egocentric collection
First-person and wearable-camera recordings of agreed tasks. Define consent, camera placement, visibility, environment coverage and acceptance criteria before collection.
Human demonstrations made useful for machine learning. We scope, collect, annotate and review real-world task data, connecting what people do with the instructions, interactions and outcomes behind each action.
For teams developing systems that perceive and act in physical environments. Each program starts with the task taxonomy, capture conditions and annotation schema your learning objective needs.
First-person and wearable-camera recordings of agreed tasks. Define consent, camera placement, visibility, environment coverage and acceptance criteria before collection.
Record how people handle objects, use tools and move through task sequences. Track the instruction, task setup and demonstrated result.
Break continuous recordings into meaningful actions and task steps using an agreed vocabulary, including transitions and ambiguous boundaries.
Mark action starts and ends, task boundaries, transitions and events. Review timing against source video and project-specific tolerance rules.
Label visible hands, manipulated objects, contact, interaction type and object state. Link actions to the relevant hand and object where the footage supports it.
Identify tools, objects, locations and relevant environment context. Capture observable state changes and flag occlusion or insufficient visual evidence.
Describe the action, manipulated object and changes over time. Keep captions grounded in visible evidence and aligned to the annotated interval.
Turn complex demonstrations into structured steps with dependencies, intermediate states and completion conditions agreed with your team.
Connect the requested instruction to the demonstration, action sequence and observed result. Flag mismatched, incomplete or ambiguous instructions.
A robot-camera frame shows grippers, objects and target markers. The example below describes how a complete pick-and-place sequence could be annotated; outcomes and timing must be verified from the source video.

Every demonstration is labelled at three levels of detail, so a robot-learning model can read it as atomic motions, object-centric task segments, or a complete episode. Human reviewers apply the same schema consistently, frame by frame.
The foundation layer: each individual action or interaction, captured on the exact frames it starts and ends.
Atomic actions on the same object or sub-task are grouped into a segment, with boundaries that align to the underlying actions.
A single summary of the entire demonstration — what was achieved, and in what environment.
Capabilities in this program
Where the task allows it, label the observed outcome against explicit criteria. Identify failure points, incomplete attempts and cases where the evidence is insufficient.
Check label consistency, temporal accuracy, captions, task steps and hand-object relationships. Route disagreements to senior review and adjudication.
Use reference tasks, annotation guidance and edge-case examples. Feed recurring errors back into contributor calibration and the project taxonomy.

Collection hardware, environments, labeling granularity and output format are confirmed during scoping. Robotics hardware operation, robot control, 3D reconstruction and sensor-derived ground truth require separate validation; they are not assumed from video annotation.
Agree tasks, instructions, capture protocol, consent, taxonomy and review criteria. Run a pilot to establish what the footage can support.
Qualify contributors, collect or ingest approved video, annotate in agreed tooling, and review difficult cases before release.
Deliver agreed structured annotations with clip references, temporal boundaries, labels, captions, outcomes and review status. Track dataset and guideline versions.
Share your intended use, demonstration environment, available footage and annotation requirements.