An end-to-end, hierarchical VTLA model for fine manipulation, enabling native, anthropomorphic last-millimeter interaction.
VTLA
Language
Vision
Tactile & Proprioception
(Force, torque, etc.)
~1Hz
System 2
Reasoning Brain
VLM
~10Hz
System 1
Motion Brain
~100Hz
System 0
Interaction Brain
Fine Actions
Progress Vector
Semantic Intent
Physical Intent
Coarse Actions
State Vector
CraftNet is the combination of System 0 and System 1, forming a VTLA (Vision-Tactile-Language-Action) model that translates high-level, multimodal inputs into continuous, fine-grained actions executed by dexterous humanoid robots — at a level of precision and speed previously only achievable by humans.