The AI Learning Revolution That's Powering the Humanoid Boom

The AI Learning Revolution That's Powering the Humanoid Boom

7 min read•Apr 24, 2026•
Carlos Mendez
Carlos Mendez

Investors poured $6.1 billion into humanoid robots in a single recent year — four times the prior year's total. That capital surge didn't come from better motors or cheaper actuators. It came from a fundamental breakthrough in how robots learn, one that's been quietly building since 2015 and has now made the science-fiction robot a plausible engineering target.



Why Robot Learning Changed Everything After 2015

For most of robotics history, intelligence meant rules — thousands of hand-coded instructions written by engineers to cover every foreseeable situation. A robot arm folding laundry needed explicit logic for sleeve orientation, fabric stiffness, collar detection, and dozens of edge cases. The ruleset exploded in complexity before it ever became reliable.

That approach produced dependable industrial robots for structured environments — welding lines, pick-and-place cells, conveyor systems — but it couldn't generalise. Move the same arm to a different context, change the lighting, introduce a new object shape, and performance collapsed immediately.

The gap between what robots could do and what researchers dreamed they could do remained stubbornly wide. Then, around 2015, the methodology shifted.

According to MIT Technology Review's deep-dive into the contemporary history of robot learning, the pivotal change was moving from rule-encoding to data-driven trial and error — and then, after 2022, to AI foundation models that learned from internet-scale data rather than hand-crafted simulations alone.


From Rules to Reinforcement: The Simulation Era

Around 2015, leading robotics labs began replacing hand-written rules with reinforcement learning (RL) — a training method where an AI agent receives reward signals for successful actions and penalty signals for failures, then iterates millions of times to discover its own strategies.

OpenAI's Dactyl project, a five-fingered robotic hand trained entirely in simulation, demonstrated both the power and the core limitation of this approach. Dactyl learned to manipulate small cubes by training inside digital environments — essentially a virtual physics engine — before being deployed on real hardware. The problem: even minor discrepancies between the simulated world and physical reality caused performance to degrade sharply.

The engineering solution was domain randomisation — deliberately introducing random variation across millions of simulated training environments. Friction coefficients, lighting conditions, object colours, and surface textures were all varied randomly so that the trained policy would be robust enough to handle the messiness of the real world. The technique worked well enough that Dactyl eventually solved Rubik's Cubes — though only 60% of the time on standard scrambles, dropping to 20% on harder configurations.

Those numbers matter for understanding where the field stood at the time. Simulation-trained RL produced genuinely impressive dexterity, but reliability was insufficient for commercial deployment. OpenAI shuttered its robotics division in 2021, reflecting the ceiling the technique had hit.

The Simulation-to-Reality Gap: Key Technical Challenges

ChallengeDescriptionMitigation Used
Visual mismatchColours and textures differ from simulationDomain randomisation
Physical propertiesFriction, deformation not perfectly modelledRandomised physics parameters
Sensor noiseReal sensors introduce latency and errorNoise injection in training
Mechanical wearActuators degrade over timeNot solved by sim-to-real alone

How Foundation Models Gave Robots Common Sense

The arrival of large language models changed robotics more profoundly than any hardware advance of the past decade. The key insight was architectural: LLMs learn by predicting what token (word, sub-word, or character) comes next in a sequence, ingesting massive corpora of text to build rich internal representations of language and world knowledge. Roboticists asked an obvious but transformative question — could the same architecture work if the tokens were sensor readings, camera frames, and joint positions instead of words?

Google DeepMind's answer was RT-1 and its successor RT-2 (Robotic Transformer). RT-1 was trained on 17 months of teleoperation data covering 700 distinct tasks, receiving robot camera views and arm joint states as inputs and generating motor commands as outputs. On tasks it had seen during training, it achieved 97% success. On entirely novel instructions, it still managed 76% — a dramatic improvement over anything simulation-only approaches had achieved.

RT-2 went further by incorporating internet-scale image and text data, giving the robot a form of common sense grounded in the broader visual world rather than just the robotics lab. This is the key conceptual leap: instead of programming robots with rules, or training them purely on robot-specific data, researchers discovered that general world knowledge — the kind baked into vision-language models during web-scale pretraining — transferred surprisingly well to physical manipulation tasks.

The practical implication is significant. A robot that has seen millions of images of kitchens, drawers, and cups during pretraining arrives with contextual understanding that rule-based systems could never acquire. It isn't certain which cup the human wants, but it has a reasonable prior. That prior dramatically reduces the amount of robot-specific training data required to reach useful performance levels.


The Limits That Still Hold the Industry Back

The current excitement is real, but it's worth mapping what remains genuinely unsolved. Foundation models for robotics face a data problem that doesn't exist for language models in the same form. Text data is abundant, cheap, and easily scraped from the web. High-quality robot demonstration data — diverse, physically grounded, and accurately labelled — is expensive to collect, hardware-dependent, and difficult to transfer between robot morphologies.

Early social robots illustrate a different limitation: capability without reliability. Jibo, the MIT-developed home social robot that raised $3.7 million in crowdfunding and retailed at $749, had compelling vision but was ultimately undermined by the pre-LLM language technology of its era. Its conversations relied on scripted response snippets that quickly felt repetitive and shallow. Today's voice AI would transform what Jibo could have been — but the new generation of AI-powered toys introduces the opposite risk. Scripted systems couldn't go off-rails; generative AI systems absolutely can, as documented cases of AI companions giving children dangerous guidance have demonstrated.

The field has traded one set of limitations (rigidity, brittleness) for another (unpredictability, safety uncertainty). Neither problem is fully solved. What has changed is that the trajectory of improvement is now measurably steeper.


What This Means for Robotics Buyers and the Hardware Market

The AI learning revolution isn't just an academic story — it's already reshaping hardware valuations in ways that matter to buyers and operators right now.

Robots whose capabilities were locked to their original programming depreciate quickly in the current market. Second-generation industrial arms with fixed motion programs have declining resale value as buyers increasingly expect adaptability. Meanwhile, hardware platforms designed to run learning-based software — with accessible compute, open APIs, and sufficient sensor payloads — are holding value more robustly.

For buyers evaluating purchases today, several implications stand out:

  • Platform extensibility matters as much as current capability. A cobot that runs modern ML inference locally will have a longer useful life than one locked to vendor-specific programming environments.
  • Used hardware pricing reflects AI readiness. Robots from platforms that have received major learning-based software updates retain value; those left behind by their manufacturers are discounting significantly.
  • Data infrastructure is the new differentiator. Buyers deploying multiple units should plan for teleoperation data collection from day one — that demonstration data becomes the training corpus for improved performance.

For operators considering entry-level deployment, the current used industrial robot market offers access to capable hardware at reduced cost, though buyers should assess software update roadmaps carefully. Similarly, the growing cobot category is particularly well-positioned to benefit from foundation model deployment, given cobots' inherently flexible, human-adjacent operating contexts.


Boston Dynamics names former Amazon AI executive Rohit Prasad CEO

Boston Dynamics has named former Amazon executive Rohit Prasad as CEO, effective tomorrow, nearly nine months after former CEO Robert Playter stepped down, first reported by Therobotreport. Prasad will replace interim CEO Amanda McMaster, as Boston Dynamics says his appointment will accelerate its physical AI strategy of combining robotics and advanced AI to commercialize intelligent machines at scale.

McMaster took over after Playter left in February. Prasad is the company’s third CEO; founder Marc Raibert led it from its creation in 1992 until 2020.

Before joining Boston Dynamics, Prasad was Amazon’s senior vice president and head scientist for Alexa and artificial general intelligence. During 12 years at Amazon, he helped build Alexa from its earliest days and later led development of the Amazon Nova foundation model family used by enterprises. Before Amazon, he spent nearly 14 years at Raytheon BBN Technologies, leading machine-learning research and its real-world application for U.S. government and commercial use.

Prasad said he plans to productize intelligent robotic systems to improve safety, productivity and operational efficiency across industrial and commercial environments. His background spans consumer AI and enterprise foundation models, while Boston Dynamics says its strategy combines advanced AI with robotics to commercialize intelligent machines.

Jaehoon Chang, Hyundai vice chair and chair of Boston Dynamics’ board, said the company’s robotics, Prasad’s AI product experience, and Hyundai Motor Group’s manufacturing, logistics and mobility capabilities provide a foundation to build and scale physical AI. Hyundai acquired a controlling stake in Boston Dynamics from SoftBank Group in 2021.

Subject to the relevant approval process, Prasad is also expected to join the company’s board.