AI · Background · 6 min read
Physical AI
A language model can describe a doorway. A robot has to pass through it without taking the frame along.
Most people met AI as text. Physical AI is the branch that has to survive contact. It includes vision, control, prediction of what a object will do when touched, and the refusal to fall over. The output is not a paragraph. It is a force, a wheel speed, or a decision to stop.
Why the demo lies
A scripted clip hides the long tail: glare, a wet floor, a lead that moved, a child, a pallet in the wrong place. Models that look general in a video are often narrow, heavily supervised, or running on a computer that will never ship inside the chest. When you read a robotics-AI announcement, separate four things that press releases glue together.
- The policy: what action the model chooses.
- The data: whose warehouses, whose hands, whose failures it learned from.
- The computer: on the robot, in the building, or in a far data centre.
- The envelope: the places and tasks where the team will actually let it run.
Language still matters
Vision-language-action models try to take an instruction in words and turn it into motion. That is genuinely new, and it is also where people over-claim. Following “put the mug in the sink” in a tidy kitchen is not the same skill as a night shift in a cluttered back room. Treat the sentence as a research direction with early products, not as a finished worker.