Perceive
Vision and sensors feed scene understanding into the World Model. Optional YOLO backend for pretrained detection.
A MODULAR VLA STACK
SIM-FIRST · MOCK + MUJOCO
Lightweight modular Vision–Language–Action for robotics. Not one giant transformer—specialized modules talk through a World Model.
World Model is the integration bus. Language never emits joint commands. Safety lives in the motion controller. LLM and planner run on events, not every tick.
THE PIPELINE
Vision and sensors feed scene understanding into the World Model. Optional YOLO backend for pretrained detection.
Language and planner turn instructions into subtasks—event-driven, not every control tick. Optional BitNet for local reasoning.
Policy → control → robot. BC-trained policy, damped-least-squares IK, and optional LeRobot/SmolVLA pretrained checkpoints.
train-bc for behavior cloning, multi-task eval across pick and place, and vla-bench for systems benchmarks.



Benchmarks
Latency and literature comparison with open VLAs—honest about what Kinetic is and isn't. Systems efficiency on the control loop, plus a gaps roadmap for what still needs real robots, vision, and trained skills.
FIG. 02 — SYSTEMS LATENCY