Robotics data
Reading time 8 min readPerceptron

How Perceptron Segments Egocentric Robot Demonstration Videos

The system follows both hands through first-person video to identify manipulation events, temporal boundaries and task labels.

By TechniaHQRobot

Perceptron reports 0.280 semantic end-to-end F1 on WGO-Bench. We clearly explain hand tracking, task segmentation, baselines and benchmark limits.

WGO-Bench contains 100 manually annotated robot and egocentric episodes with 743 gold manipulation segments.

Perceptron reports 0.280 semantic end-to-end F1 for its Egocentric pipeline.

Macrodata’s public benchmark page reports 0.158 for one-pass Gemini segmentation labels and 0.168 for a seeded relabeling pipeline.

No independent reproduction of Perceptron’s 0.280 result was located at publication time.

Original X post

Open on X

Egocentric video records the task from the worker’s viewpoint

Egocentric video is captured from a camera worn on the head, chest or body, so the frame follows the person completing the task. Hands, tools and manipulated objects dominate the view. That perspective is valuable for robot learning because it records the sequence of physical decisions in ordinary human environments.

Raw footage is not yet a training-ready demonstration. A ninety-second clip may contain reaching, pauses, two-handed coordination, occlusion and repeated corrections. A useful annotation system must identify when a meaningful manipulation starts, when it ends and what changed in the scene.

Why robot-learning teams need temporal structure

Task segments can support imitation learning, policy retrieval, reward modeling and dataset filtering. They let a team search for every example of opening a lid or placing an object instead of reviewing complete videos manually.

WGO-Bench tests boundaries and action labels

WGO-Bench, short for What’s Going On Benchmark, is a public dataset from Macrodata Labs. Its current dataset card lists 100 episodes, 743 gold subtask segments and 63 unique task instructions drawn from human egocentric and robot-camera sources.

The benchmark evaluates two linked problems. Boundary detection asks where one completed manipulation event ends and the next begins. Subtask labeling asks the model to describe the completed event. The end-to-end setting requires both the time window and the semantic label to match the human annotation.

The annotations follow object-state changes

The benchmark policy treats events such as picking up, placing, opening, closing, pouring and wiping as meaningful boundaries. Small pose corrections or camera motion are not supposed to become separate subtasks unless the physical state changes.

Technical details

System
Perceptron Egocentric
Input
First-person or robot demonstration video
Output
Temporal manipulation segments, hand tracks, keypoints and semantic action labels
Reported score
0.280 semantic end-to-end F1 on WGO-Bench
Public comparison
0.158 one-pass Gemini labels and 0.168 seeded relabeling in Macrodata’s published evaluation
Reproduction status
Company-reported result, not independently reproduced

What the 0.280 semantic end-to-end F1 score measures

Perceptron reports a semantic end-to-end F1 score of 0.280 on WGO-Bench. In this setting, a prediction must overlap the correct temporal segment and pass the semantic label test. Precision penalizes extra or incorrectly labeled segments, recall penalizes missed gold segments and F1 balances the two.

The number is not a robot task-success rate. It measures agreement with a small human-annotated benchmark under a specific matching protocol. A score of 0.280 also shows that most events are not yet captured perfectly, especially when actions are brief, hands are occluded or the label depends on context outside the current frame.

The 0.158 Gemini comparison needs precise context

The Perceptron launch compares 0.280 with 0.158 for a Gemini-based pipeline. Macrodata’s public end-to-end table identifies 0.158 as its one-pass segmentation-and-labeling result. The same table reports 0.168 when predicted segments are passed through a separate seeded relabeling stage.

Those figures make the comparison reproducible at the baseline level, but they also show why “strongest Gemini pipeline” can be ambiguous unless the exact recipe is named. The article therefore reports 0.158 as the referenced one-pass Gemini baseline and notes the higher two-stage result separately.

Continuous hand tracking can locate grasps and releases

Perceptron describes tracking left and right hands throughout the video with persistent identities, boxes and keypoints. A grasp can indicate that an object has become controlled by the hand, while release can mark the completion of a placement or transfer. Those events provide physically meaningful candidates for segment boundaries.

Tracking the entire sequence avoids reducing the video to a contact sheet of isolated frames. It can preserve motion order and short interactions that fall between sampled images. It can also represent bimanual work, where one hand stabilizes a container while the other opens, pours or inserts an object.

Hand contact is useful evidence, not complete semantics

The same grasp can belong to carrying, pouring, cleaning or placing. The system still needs visual context, object identity and temporal history to assign a useful task label.

Automation can reduce annotation work without removing review

A pipeline that proposes boundaries and labels can move human annotators from drawing every segment to reviewing uncertain cases. That can make large robotics datasets more practical, especially when one minute of complex egocentric video takes many minutes to label carefully.

Quality control remains necessary. Hands disappear behind objects, left and right identities can switch after occlusion and short actions may last only a few frames. Dataset teams need confidence scores, audit samples and clear rules for correcting uncertain labels before the output becomes policy-training data.

The benchmark result remains company-reported

The WGO-Bench dataset and Macrodata baseline pipeline are public. Perceptron’s product description and launch material report the 0.280 result, but no independent reproduction or third-party evaluation was located at publication time.

WGO-Bench is also intentionally small. It covers 100 episodes from three source families, so performance may change on longer tasks, different camera placements, severe occlusion, workshops, outdoor scenes or demonstrations with unfamiliar tools. A broader evaluation would need hidden test data, more environments and published inference settings.

Verification notes

  • Perceptron’s 0.280 result is reported by the company and was not independently reproduced during publication.
  • The 0.158 comparison is the one-pass Gemini end-to-end result in Macrodata’s public table; the separate seeded relabeling pipeline reaches 0.168.
  • The public WGO-Bench dataset card currently lists 63 unique task instructions, while an earlier Macrodata article summary lists 62. The article uses the current dataset card for the dataset count.

Explore the technical topic

Share this article

Share the current TechniaHQRobot article page.

Sources

Editor

Editor : @techniahqrobot

TechniaHQRobot editorial coverage on AI, robotics, automation and Physical AI.

@TECHNIAHQROBOT

FollowTechniaHQRobot

Independent coverage of humanoid robots, Physical AI, industrial robotics, robot hardware and emerging automation systems.

Follow our daily updates or explore the latest robotics coverage.

service@techniahqservice.com