Embedl is now part of AMD's Open Robotics Ecosystem, featured at Advancing AI 2026. Our contribution: we optimized Physical Intelligence's π0.5, a state-of-the-art vision-language-action model, to run on the AMD Ryzen AI Embedded iGPU with robot reasoning latency under 100 milliseconds and no loss in accuracy.
That number now sits on AMD's own product pages. The benchmark behind the Kria AI platform's "robot reasoning in under 100 ms" claim is Embedl's work, measured on GPU VLA inference with π0.5.
Why 100 milliseconds matters
Vision-language-action models are how modern robots reason: they look at a scene, understand an instruction, and decide what to do next. They are also large, and robots don't get datacenters. A robot arm deciding its next grasp, or a mobile robot rerouting around a person, needs that decision on the device, within the time the physical world allows. Below 100 milliseconds, VLA reasoning stops being a lab demo and becomes something you can build a product on.
What Embedl did
π0.5 was not built for an embedded iGPU. Getting it there is Embedl's craft: we optimize the model and its deployment for the specific silicon it will run on, squeezing out latency at every level of the stack while holding accuracy where the original model set it. The result on the Ryzen AI Embedded iGPU is reasoning fast enough for a robot to act on, with the same answers the full-size deployment would give.
The point is not one benchmark. The same optimization work applies to whatever model our customers ship, LLMs, VLMs, or VLAs, on whatever edge hardware they've chosen. π0.5 under 100 milliseconds is what it looks like when that process meets serious silicon.
That silicon is AMD's Ryzen AI Embedded X100 series, the device at the heart of the new AMD Kria™ AI SOM: real-time CPU control, iGPU and NPU inference, and 128 GB of unified memory on one production-ready module, backed by an open software stack built on ROCm™ and ROS 2. An open, partner-driven ecosystem for robotics is the right way to build this industry, and we're glad to stand in it alongside the leaders in AI models, robot hardware, and silicon.
That number now sits on AMD's own product pages. The benchmark behind the Kria AI platform's "robot reasoning in under 100 ms" claim is Embedl's work, measured on GPU VLA inference with π0.5.
Why 100 milliseconds matters
Vision-language-action models are how modern robots reason: they look at a scene, understand an instruction, and decide what to do next. They are also large, and robots don't get datacenters. A robot arm deciding its next grasp, or a mobile robot rerouting around a person, needs that decision on the device, within the time the physical world allows. Below 100 milliseconds, VLA reasoning stops being a lab demo and becomes something you can build a product on.
What Embedl did
π0.5 was not built for an embedded iGPU. Getting it there is Embedl's craft: we optimize the model and its deployment for the specific silicon it will run on, squeezing out latency at every level of the stack while holding accuracy where the original model set it. The result on the Ryzen AI Embedded iGPU is reasoning fast enough for a robot to act on, with the same answers the full-size deployment would give.
The point is not one benchmark. The same optimization work applies to whatever model our customers ship, LLMs, VLMs, or VLAs, on whatever edge hardware they've chosen. π0.5 under 100 milliseconds is what it looks like when that process meets serious silicon.
That silicon is AMD's Ryzen AI Embedded X100 series, the device at the heart of the new AMD Kria™ AI SOM: real-time CPU control, iGPU and NPU inference, and 128 GB of unified memory on one production-ready module, backed by an open software stack built on ROCm™ and ROS 2. An open, partner-driven ecosystem for robotics is the right way to build this industry, and we're glad to stand in it alongside the leaders in AI models, robot hardware, and silicon.
"Bringing large VLA and robotics models to efficient edge hardware is what Embedl's technology enables. AMD's platform lets us do it without compromising on accuracy, and π0.5 at under 100 milliseconds shows what that combination unlocks."
Hans Salomonsson, CEO and Co-Founder, Embedl
What this means for your team
If you're putting VLMs or VLAs on real machines, the question is rarely whether the model works. It's whether it works on your hardware, at your latency budget, with accuracy you can defend. That's the work we did here, and it's the work we do for customers. Bring us your model and your target, and we'll tell you what's possible.
Read more about the platform on AMD's Kria AI pages and Advancing AI 2026.