How LingBot-VLA 2.0 Filters 90,000 Raw Robot Hours
The useful number is not only dataset size. It is how the pipeline removes jerky trajectories, static episodes, blurred video and state-to-image misalignment.
By TechniaHQRobot
Robbyant collected about 90,000 hours of physical robot operation and retained 50,000 hours of high-quality robot trajectories after filtering. The official paper lists 20 robot configurations spanning arms, mobile manipulators, half-humanoids and humanoids.
The official paper says approximately 90,000 hours of raw robotic data were filtered into 50,000 hours of high-quality robot trajectories.
The paper and repository enumerate 20 robot configurations; the TechniaHQRobot source post described coverage as more than 20, so this article uses the official enumerated count.
The configurations include Astribot S1 and Leju KUAVO 4 Pro alongside single-arm, dual-arm, mobile and humanoid systems.
Dataset scale and embodiment diversity do not by themselves prove reliable deployment in every task or environment.
Original X post
Open on XThe pipeline removed 40,000 hours instead of treating all data as equal
Large robot datasets contain failed demonstrations, duplicated idle time, camera blur, dropped frames and state logs that no longer align with the recorded image. Training on every recorded second can teach a policy to imitate noise or associate the wrong action with the wrong visual state.
Robbyant’s paper says it began with about 90,000 hours of robotic data and retained 50,000 hours after a redesigned cleaning process. That 44% reduction is a reminder that recorded volume and usable training signal are different quantities.
Motion filters inspect velocity, acceleration and jerk
The pipeline computes derivatives of action and state signals to detect trajectories with abnormal velocity, acceleration or jerk. Thresholds are set separately for each embodiment because a compact arm, mobile manipulator and humanoid do not share the same normal motion range.
Episodes are also removed when state and action signals remain almost unchanged for more than 95% of the recording. This prevents long idle segments from dominating the dataset while contributing little information about manipulation.
Technical details
- Raw robotic pool
- Approximately 90,000 hours
- Retained robot data
- 50,000 hours after filtering
- Robot configurations
- 20 enumerated in the official paper and repository
- Platform coverage
- Single-arm, dual-arm, half-humanoid, humanoid and mobile manipulation systems
- Named examples
- Astribot S1 and Leju KUAVO 4 Pro
- Additional pretraining stream
- 10,000 hours of filtered egocentric human video, discussed separately in the release article
Video and robot state must describe the same physical motion
A trajectory can look smooth numerically and still be unusable when timestamps, camera frames and joint states are misaligned. Robbyant projects the robot model into the image using the recorded state and URDF, then uses human review to find mismatches between the projected body and the video.
Human annotators also remove blur, severe occlusion, dropped frames and misalignment across camera views. These checks matter because a VLA policy learns associations between what the camera sees and what the robot does next.
Hardware diversity changes the action space
The official table covers arms with grippers, dual-arm platforms, mobile bases, half-humanoids and humanoids with dexterous hands. Named systems include Astribot S1 and Leju KUAVO 4 Pro. Other configurations contribute different arm, body, hand, head, waist and mobility signals.
Cross-embodiment training can expose the model to shared manipulation patterns while forcing it to separate hardware-specific dynamics. A grasping instruction may be semantically similar across robots, but the joint layout, reachable workspace, control frequency and end effector differ.
Why cleaned cross-embodiment data matters
A generalist robot policy needs repeated evidence that links language, visual state and action across many objects and bodies. Cleaning reduces contradictory supervision. Hardware diversity reduces the risk that the policy memorizes one camera geometry or one arm kinematic chain.
Neither step guarantees generalization. A model can still fail when the lighting, object friction, calibration, payload, camera placement or task sequence moves outside the training distribution.
What the demonstration proves
The official paper and repository document a concrete data-processing pipeline and enumerate 20 robot configurations. They also report evaluations on GM-100 and long-horizon mobile manipulation across two physical platforms.
This supports the claim that LingBot-VLA 2.0 was trained from a large, heterogeneous and explicitly filtered corpus. It does not reveal every source episode or independently verify the quality label assigned to each retained trajectory.
What remains unproven
The 50,000-hour scale does not prove reliable operation on every one of the 20 configurations, nor does it establish performance in factories, homes or public spaces. Dataset coverage is an input to training, not a deployment certificate.
Independent replication would require access to the data composition, filtering thresholds, rejected-sample distribution, benchmark protocol and model checkpoints used for each reported result. The public release provides code and weights, but not the complete 50,000-hour proprietary robot corpus.
The two dataset totals refer to different scopes
The 50,000-hour figure refers to cleaned robot trajectories. The 60,000-hour pretraining total adds 10,000 hours of filtered egocentric human-operation video. They are not competing estimates of the same collection.
Keeping those scopes separate prevents the raw 90,000-hour robot pool, the cleaned 50,000-hour robot subset and the combined 60,000-hour pretraining corpus from being merged into one misleading number.
Verification notes
- The official paper and repository enumerate 20 robot configurations. The source post described more than 20; the official enumerated count is used in the article body.
- The raw 90,000-hour robot pool, cleaned 50,000-hour robot subset and combined 60,000-hour pretraining corpus refer to different stages and scopes.
- The complete cleaned robot dataset is not presented as a public download in the sources reviewed.
Frequently asked questions
Did LingBot-VLA 2.0 train on all 90,000 robot hours?
No. Robbyant reports collecting about 90,000 hours of robotic data and retaining 50,000 hours after filtering.
Why does another LingBot-VLA 2.0 figure say 60,000 hours?
The 60,000-hour pretraining total combines 50,000 hours of cleaned robot trajectories with 10,000 hours of filtered egocentric human video.
Does coverage of 20 robot configurations prove deployment on all 20?
No. Training coverage and reported evaluations do not establish reliable production deployment on every configuration or in every environment.
Explore the technical topic
Share this article
Share the current TechniaHQRobot article page.
Sources
Editor : @techniahqrobot
TechniaHQRobot editorial coverage on AI, robotics, automation and Physical AI.
Related articles
Original technical source
Follow Robbyant on X:
https://x.com/robbyant_brain