Robot foundation model data
Reading time 8 min readLingBot-VLA 2.0

How LingBot-VLA 2.0 Filters 90,000 Raw Robot Hours

The useful number is not only dataset size. It is how the pipeline removes jerky trajectories, static episodes, blurred video and state-to-image misalignment.

By TechniaHQRobot

Robbyant collected about 90,000 hours of physical robot operation and retained 50,000 hours of high-quality robot trajectories after filtering. The official paper lists 20 robot configurations spanning arms, mobile manipulators, half-humanoids and humanoids.

The official paper says approximately 90,000 hours of raw robotic data were filtered into 50,000 hours of high-quality robot trajectories.

The paper and repository enumerate 20 robot configurations; the TechniaHQRobot source post described coverage as more than 20, so this article uses the official enumerated count.

The configurations include Astribot S1 and Leju KUAVO 4 Pro alongside single-arm, dual-arm, mobile and humanoid systems.

Dataset scale and embodiment diversity do not by themselves prove reliable deployment in every task or environment.

Original X post

Open on X

The pipeline removed 40,000 hours instead of treating all data as equal

Large robot datasets contain failed demonstrations, duplicated idle time, camera blur, dropped frames and state logs that no longer align with the recorded image. Training on every recorded second can teach a policy to imitate noise or associate the wrong action with the wrong visual state.

Robbyant’s paper says it began with about 90,000 hours of robotic data and retained 50,000 hours after a redesigned cleaning process. That 44% reduction is a reminder that recorded volume and usable training signal are different quantities.

Motion filters inspect velocity, acceleration and jerk

The pipeline computes derivatives of action and state signals to detect trajectories with abnormal velocity, acceleration or jerk. Thresholds are set separately for each embodiment because a compact arm, mobile manipulator and humanoid do not share the same normal motion range.

Episodes are also removed when state and action signals remain almost unchanged for more than 95% of the recording. This prevents long idle segments from dominating the dataset while contributing little information about manipulation.

Technical details

Raw robotic pool
Approximately 90,000 hours
Retained robot data
50,000 hours after filtering
Robot configurations
20 enumerated in the official paper and repository
Platform coverage
Single-arm, dual-arm, half-humanoid, humanoid and mobile manipulation systems
Named examples
Astribot S1 and Leju KUAVO 4 Pro
Additional pretraining stream
10,000 hours of filtered egocentric human video, discussed separately in the release article

Video and robot state must describe the same physical motion

A trajectory can look smooth numerically and still be unusable when timestamps, camera frames and joint states are misaligned. Robbyant projects the robot model into the image using the recorded state and URDF, then uses human review to find mismatches between the projected body and the video.

Human annotators also remove blur, severe occlusion, dropped frames and misalignment across camera views. These checks matter because a VLA policy learns associations between what the camera sees and what the robot does next.

Hardware diversity changes the action space

The official table covers arms with grippers, dual-arm platforms, mobile bases, half-humanoids and humanoids with dexterous hands. Named systems include Astribot S1 and Leju KUAVO 4 Pro. Other configurations contribute different arm, body, hand, head, waist and mobility signals.

Cross-embodiment training can expose the model to shared manipulation patterns while forcing it to separate hardware-specific dynamics. A grasping instruction may be semantically similar across robots, but the joint layout, reachable workspace, control frequency and end effector differ.

Why cleaned cross-embodiment data matters

A generalist robot policy needs repeated evidence that links language, visual state and action across many objects and bodies. Cleaning reduces contradictory supervision. Hardware diversity reduces the risk that the policy memorizes one camera geometry or one arm kinematic chain.

Neither step guarantees generalization. A model can still fail when the lighting, object friction, calibration, payload, camera placement or task sequence moves outside the training distribution.

What the demonstration proves

The official paper and repository document a concrete data-processing pipeline and enumerate 20 robot configurations. They also report evaluations on GM-100 and long-horizon mobile manipulation across two physical platforms.

This supports the claim that LingBot-VLA 2.0 was trained from a large, heterogeneous and explicitly filtered corpus. It does not reveal every source episode or independently verify the quality label assigned to each retained trajectory.

What remains unproven

The 50,000-hour scale does not prove reliable operation on every one of the 20 configurations, nor does it establish performance in factories, homes or public spaces. Dataset coverage is an input to training, not a deployment certificate.

Independent replication would require access to the data composition, filtering thresholds, rejected-sample distribution, benchmark protocol and model checkpoints used for each reported result. The public release provides code and weights, but not the complete 50,000-hour proprietary robot corpus.

The two dataset totals refer to different scopes

The 50,000-hour figure refers to cleaned robot trajectories. The 60,000-hour pretraining total adds 10,000 hours of filtered egocentric human-operation video. They are not competing estimates of the same collection.

Keeping those scopes separate prevents the raw 90,000-hour robot pool, the cleaned 50,000-hour robot subset and the combined 60,000-hour pretraining corpus from being merged into one misleading number.

Verification notes

  • The official paper and repository enumerate 20 robot configurations. The source post described more than 20; the official enumerated count is used in the article body.
  • The raw 90,000-hour robot pool, cleaned 50,000-hour robot subset and combined 60,000-hour pretraining corpus refer to different stages and scopes.
  • The complete cleaned robot dataset is not presented as a public download in the sources reviewed.

Frequently asked questions

Did LingBot-VLA 2.0 train on all 90,000 robot hours?

No. Robbyant reports collecting about 90,000 hours of robotic data and retaining 50,000 hours after filtering.

Why does another LingBot-VLA 2.0 figure say 60,000 hours?

The 60,000-hour pretraining total combines 50,000 hours of cleaned robot trajectories with 10,000 hours of filtered egocentric human video.

Does coverage of 20 robot configurations prove deployment on all 20?

No. Training coverage and reported evaluations do not establish reliable production deployment on every configuration or in every environment.

Explore the technical topic

Share this article

Share the current TechniaHQRobot article page.

Sources

Editor

Editor : @techniahqrobot

TechniaHQRobot editorial coverage on AI, robotics, automation and Physical AI.

Original technical source

Follow Robbyant on X:

https://x.com/robbyant_brain

@TECHNIAHQROBOT

FollowTechniaHQRobot

Independent coverage of humanoid robots, Physical AI, industrial robotics, robot hardware and emerging automation systems.

Follow our daily updates or explore the latest robotics coverage.

service@techniahqservice.com