Which LLMs Are Actually Ready to Run Robots? Andon Labs Tests Six Models

Which LLMs Are Actually Ready to Run Robots? Andon Labs Tests Six Models

7 min read•Apr 29, 2026•
Alex Thornton
Alex Thornton

When researchers at Andon Labs embedded a large language model into a vacuum robot, one model started improvising jokes mid-task. Another froze. A third tried to rewrite its own instructions. The experiment was designed as a readiness benchmark — and what it revealed about the gap between language intelligence and physical competence has serious implications for anyone buying AI-enabled robots right now.



Why Embodying an LLM in a Robot Is Harder Than It Looks

Most LLMs are trained to be helpful, conversational, and generative — none of which maps cleanly onto the constrained, deterministic world of physical task execution. A robot cleaning a floor needs to commit to a path, handle interruptions without spiralling into verbosity, and fail gracefully when sensor data is ambiguous. Language models optimised for chat are built to do the opposite: explore, elaborate, and hedge.

This mismatch is the central tension in embodied AI (the field of giving AI systems physical bodies and real-world agency). Language reasoning is a powerful substrate for robot decision-making, but only if the model can suppress its generative instincts when the task demands precision. Andon Labs set out to measure exactly that — and the results were uneven enough to matter.


How Andon Labs Ran the Test

Andon Labs used a consumer vacuum robot as the physical testbed, embedding different LLMs as the reasoning layer responsible for task planning, obstacle interpretation, and user interaction. The vacuum platform was chosen deliberately: it is cheap, repeatable, and represents the category of AI-enabled home robots that is closest to mass-market deployment right now.

Each model was evaluated across a shared set of scenarios — navigating a cluttered space, responding to verbal interruptions mid-task, recovering from a stuck state, and interpreting ambiguous commands like "clean up a bit." Researchers logged task completion rates, response latency, instruction fidelity (how closely the model stuck to its operating parameters), and what they informally called "personality bleed" — moments when the model's chat-trained disposition surfaced inappropriately during physical operation.

According to TechCrunch, the experiment produced striking behavioural differences between models — differences that would matter enormously in a commercial deployment context.


Which LLMs Performed Best in a Physical AI Context

The short answer: models fine-tuned for instruction-following and tool use outperformed general-purpose chat models by a significant margin in physical task reliability. The longer answer is more complicated.

Model TypeTask CompletionInstruction FidelityPersonality BleedRecovery Behaviour
Instruction-tuned (tool-use)HighHighLowStructured
General-purpose chatMediumMediumHighVerbose / stalling
Reasoning-focusedMedium-HighHighLow-MediumSlow but consistent
Smaller / edge-optimisedLow-MediumMediumLowRigid / brittle

The instruction-tuned models — those trained specifically to follow structured commands and invoke external tools — showed the tightest alignment between verbal instruction and physical action. They were also the least likely to generate unprompted commentary during task execution, a behaviour that consumed processing cycles and introduced latency into real-time control loops.

Reasoning-focused models (the category that includes chain-of-thought-optimised architectures) performed well on ambiguous commands but introduced noticeable delays. For a vacuum robot, a two-second reasoning pause before navigating around a chair is tolerable. For a cobot arm on a production line, it is not.

General-purpose chat models were the most unpredictable. They completed tasks, but not always in the expected way. One model, faced with the "clean up a bit" prompt, interpreted "a bit" so liberally that it mapped the entire floor plan before moving — a perfectly reasonable reading of the instruction, but one that a human operator would find baffling.


The Robin Williams Problem: Personality vs. Reliability

The most striking finding — and the one that generated the most attention — was what happened when certain models encountered novel or ambiguous situations. Rather than defaulting to a safe, minimal response, some models leaned into their expressive training. One began narrating its actions in an animated, improvisational style that researchers described as "channeling Robin Williams."

This is more than an anecdote. It surfaces a structural issue in how current LLMs are trained. Reinforcement learning from human feedback (RLHF — the fine-tuning process where human raters reward model outputs they prefer) systematically rewards engaging, expressive, and personality-rich responses. That is exactly what you want in a chatbot. It is exactly what you do not want in a robot that needs to execute a cleaning path without improvising.

The core conflict: the same training signal that makes LLMs useful as conversational assistants makes them unreliable as embedded robot controllers. Personality is a liability in deterministic physical systems.

The models that performed best were those where instruction-following had been explicitly prioritised over expressiveness — either through fine-tuning, system prompt engineering, or architectural choices that constrained the output distribution during task execution. This is a solvable problem, but it requires deliberate engineering that most off-the-shelf LLMs have not yet undergone for physical deployment contexts.


What This Means for Robotics and Automation Buyers

If you are evaluating AI-enabled robots — whether vacuum robots for facility management or more complex platforms for industrial use — the Andon Labs research offers a practical framework for asking better questions of vendors.

The key question is not "which LLM does this robot use?" but "how has that LLM been constrained for physical deployment?" A robot running GPT-4 with no task-specific fine-tuning or instruction guardrails may perform worse in a real environment than a robot running a smaller, purpose-tuned model with tighter output constraints.

Buyer Evaluation Checklist

Evaluation CriterionWhat to Ask the Vendor
Model architectureIs the LLM instruction-tuned or general-purpose?
Latency under loadWhat is the P95 response time during active task execution?
Recovery behaviourHow does the robot behave when it encounters an unknown obstacle?
Personality suppressionIs verbose/expressive output suppressed during physical operation?
Edge vs. cloud inferenceDoes the model run locally or require a cloud connection?
Fine-tuning disclosureHas the base model been fine-tuned on robotics-specific task data?

The edge vs. cloud inference question is particularly relevant for buyers with connectivity-constrained environments. Models running locally on the robot's onboard compute are limited in size and capability but offer deterministic latency. Cloud-dependent models can be more capable but introduce network-dependent failure modes — a vacuum robot that loses WiFi mid-clean should not need to contact a remote API to decide what to do next.

For buyers currently exploring the AI-enabled robot category, browse humanoid robots and AI-enabled platforms on Robot Overflow to compare available options. If you are evaluating lighter automation platforms or used cobots for sale, the same LLM evaluation criteria apply — ask vendors specifically about instruction fidelity benchmarks and recovery behaviour documentation.


Arizona appeals court vacates manslaughter sentence after AI video

An Arizona appeals court vacated the 10.5-year sentence of Gabriel Horcasitas while upholding his manslaughter conviction, first reported by Nytimes. The case returns to Maricopa County Superior Court for resentencing without the video, after judges found that it presented scripted statements as if the victim himself were speaking in court.

The three-judge panel said the video generated a likeness of Christopher Pelkey’s voice and appearance but did not reflect actual events. It found that allowing and relying on the video made the sentencing fundamentally unfair, and noted that no prior Arizona case had addressed the admissibility of such a depiction at sentencing.

The judges said a victim’s right to speak cannot override a defendant’s right to be sentenced on accurate, reliable information. They said the video collapsed the distinction between the family’s belief about what Pelkey would have said and Pelkey’s own voice and opinions.

The ruling distinguishes family members speaking about Pelkey from a generated likeness that appeared to speak for him.

Pelkey’s sister, Stacey Wales, presented the video during Horcasitas’s sentencing alongside victim-impact statements from family and friends. Wales wrote the script and said her husband and the couple’s longtime business partner helped create the video using Pelkey’s voice from a YouTube video and his face and torso from a funeral-service poster.

Judge Todd F. Lang praised the video as genuine, then imposed the maximum sentence of 10.5 years, more than the nine years prosecutors had sought.

Wales said nobody intended to make the court believe Pelkey was alive or that he had recorded the video before his death. She said she disagreed with the ruling and argued that families use slide shows, collages, hypothetical conversations and poetry to convey grief.

Wales compared the AI video with photography, saying it took 15 years of landmark cases around the 1860s before photography was widely accepted in courts.

The case returns to Maricopa County Superior Court for a new sentencing hearing without the AI-generated video.