Open-source robot foundation models
Reading time 8 min readLingBot-VLA 2.0

LingBot-VLA 2.0 Open Source Release and Robot Coverage

The release exposes model code, training and deployment paths and pretrained weights, while the full proprietary pretraining corpus is not presented as an open dataset.

By TechniaHQRobot

Robbyant has published the LingBot-VLA 2.0 codebase and a 6B pretrained checkpoint. The official pretraining total is about 60,000 hours: 50,000 hours of cleaned robot trajectories plus 10,000 hours of filtered egocentric human video.

The official GitHub repository is public under the Apache-2.0 license and provides training, evaluation and deployment instructions.

Robbyant publishes a 6B LingBot-VLA 2.0 pretrained checkpoint through Hugging Face and ModelScope links.

The 60,000-hour pretraining corpus combines 50,000 hours of cleaned robot trajectories with 10,000 hours of filtered egocentric human video.

The TechniaHQRobot announcement says 17 robot brands; the official paper separately enumerates 20 configurations and does not publish a distinct brand-count field.

Original X post

Open on X

The release includes code and a pretrained 6B checkpoint

The public repository contains installation instructions, model download commands, training configuration, open-loop evaluation, RoboTwin deployment and a real-robot policy server. Robbyant links a 6B pretrained native-depth checkpoint from the repository.

The Apache-2.0 license makes the codebase usable for research and development under the license terms. Public code does not mean the full training dataset, every internal tool or every evaluation environment has been released.

The 60,000-hour total combines two different data streams

Robbyant’s paper defines the pretraining corpus as 50,000 hours of cleaned robot trajectories across 20 configurations plus 10,000 hours of filtered egocentric human-operation video. The human video stream is used to supply manipulation-centric motion and semantic priors after reconstruction and quality control.

This resolves the apparent difference between 50,000 and 60,000 hours. The smaller number is the retained robotic subset. The larger number is the combined pretraining corpus.

Technical details

Public repository
Robbyant/lingbot-vla-v2
License
Apache-2.0
Released checkpoint
LingBot-VLA 2.0 6B native-depth model
Pretraining total
Approximately 60,000 hours
Data composition
50,000 hours robot trajectories + 10,000 hours egocentric human video
Embodiment coverage
20 robot configurations in the official paper

Twenty configurations are documented; the 17-brand count has a narrower source

The official paper and repository describe 20 robot configurations. They include single-arm and dual-arm systems, half-humanoids, humanoids and mobile manipulators, with named platforms such as Astribot S1 and Leju KUAVO 4 Pro.

The TechniaHQRobot announcement associated with the release states 17 robot brands. The official paper does not expose a separate brand-count field, so this article preserves 17 as a source-post claim rather than deriving it from the configuration table.

A unified action vector connects different robot bodies

LingBot-VLA 2.0 maps heterogeneous platforms into a 55-dimensional canonical state and action vector. It allocates fields for arm joints, end-effector pose, grippers, dexterous hands, waist, head and mobility signals, with padding where a body lacks a component.

This representation lets one model ingest trajectories from robots with different shapes. It does not make their dynamics identical. Post-training, normalization and configuration mapping are still required for a target platform.

Open weights lower the entry barrier but do not remove deployment work

The repository provides a path from model download to post-training and deployment. Real hardware still needs camera calibration, action mapping, safety limits, latency measurement, low-level control and a dataset that matches the target task.

Robbyant reports one inference call taking about 130 milliseconds on an NVIDIA GeForce RTX 4090D with 10 denoising steps for the documented deployment path. That number belongs to the stated configuration and should not be generalized to other GPUs or complete robot response time.

What the demonstration proves

The public repository, license and model links prove that a usable LingBot-VLA 2.0 codebase and pretrained checkpoint have been released. The paper documents the intended data composition, unified action representation and evaluation tasks.

The accompanying demonstrations show the model operating on the authors’ selected robot platforms and tasks. They support reproducible inspection of the released model, but they do not expose the full proprietary pretraining corpus.

What remains unproven

Open-source availability does not prove that the model will reproduce Robbyant’s results on different hardware, cameras, data quality or task definitions. Users still need to measure success, latency, interventions and safety on their own system.

The release does not establish universal cross-embodiment control, unattended production deployment or reliable operation in every home and factory. The 20-configuration pretraining coverage is broader training evidence, not a guarantee for each platform.

The practical value is inspectability

Researchers can inspect the architecture, load the checkpoint, prepare a LeRobot-format dataset and test post-training on a defined task. That is more useful than a video-only announcement because failures can be measured against the same code path.

The next strong evidence will come from independent reproductions that publish hardware, dataset size, trial counts, success criteria, failure modes and end-to-end latency.

Verification notes

  • The 60,000-hour total is composed of 50,000 hours of cleaned robot trajectories and 10,000 hours of filtered egocentric human video.
  • The official paper documents 20 robot configurations. The 17-brand figure is retained only as a claim from the associated TechniaHQRobot announcement.
  • The code and checkpoint are public; the complete proprietary pretraining corpus is not presented as a downloadable dataset in the reviewed sources.

Frequently asked questions

Is LingBot-VLA 2.0 open source?

Yes. Robbyant publishes the code under Apache-2.0 and links a 6B pretrained checkpoint. The complete 60,000-hour pretraining corpus is not presented as a public dataset.

Why are both 50,000 and 60,000 hours reported?

The 50,000 hours are cleaned robot trajectories. The 60,000-hour total adds 10,000 hours of filtered egocentric human video.

Does the model support 17 brands or 20 configurations?

The TechniaHQRobot announcement states 17 brands. The official Robbyant paper separately enumerates 20 robot configurations and does not publish a distinct brand-count field.

Explore the technical topic

Share this article

Share the current TechniaHQRobot article page.

Sources

Editor

Editor : @techniahqrobot

TechniaHQRobot editorial coverage on AI, robotics, automation and Physical AI.

Original technical source

Follow Robbyant on X:

https://x.com/robbyant_brain

@TECHNIAHQROBOT

FollowTechniaHQRobot

Independent coverage of humanoid robots, Physical AI, industrial robotics, robot hardware and emerging automation systems.

Follow our daily updates or explore the latest robotics coverage.

service@techniahqservice.com