Watch a modern humanoid robot walk and you are watching something strange: a machine moving with the loose, heel-striking confidence of a person, controlled not by equations written by an engineer but by a neural network that taught itself. Nobody hand-programs a humanoid's walk anymore. The walk is learned, and the classroom is a simulation where the robot falls down millions of times before it ever touches a real floor.

The method behind this is reinforcement learning, and in the last few years it has become the standard way to teach humanoids to move. As Toyota researcher Takahiro Ito puts it, reinforcement learning is a type of machine learning in which an AI repeatedly attempts certain actions in an environment and autonomously discovers strategies that maximize a predefined reward. For walking, his team sets rewards such as a bonus for walking near a target speed and a penalty when a foot slips. Then the robot tries, fails, and tries again.

Trial, error, repeat, at superhuman speed

The reason this works now, and did not work a decade ago, is compute. Training happens in GPU-accelerated physics simulators where thousands of virtual humanoids run in parallel, each with slightly different physical parameters. Figure, the humanoid company behind Figure 02, says it simulates years of walking data in only a few hours. Toyota's team reports that within about one to two hours of training, their simulated robots learned to balance and walk.

The parallelism matters as much as the speed. Each virtual robot encounters different terrains, changes in actuator dynamics, and surprise events like trips, slips, and shoves. A single neural network policy learns to handle all of them at once. This is how you get robustness: not by programming responses to every possible push, but by letting one policy experience a lifetime of pushes in an afternoon.

The reward design is where the engineering taste lives. A naive reward, like "move forward," produces robots that lurch, shuffle, or flail in ways that technically satisfy the goal. Figure injects a preference for human-like walking by rewarding the robot for mimicking human walking reference trajectories, establishing a prior over acceptable walking styles. The policy learns heel-strikes, toe-offs, and arm-swing synchronized with leg movement, while additional reward terms optimize for velocity tracking, power consumption, and robustness to perturbations and terrain changes.

A single neural network policy learns to handle a lifetime of pushes in an afternoon.

Crossing the sim-to-real gap

Enjoying this story?

Get the five most important stories in tech, every morning. Free.

Learning to walk in simulation is only half the battle. A simulated robot is, at best, an approximation of a complex electro-mechanical system, and a policy trained in simulation is guaranteed to work only on simulated robots. The differences between the virtual and physical worlds even have a name: the sim-to-real gap. Toyota's Mitsuki Morita describes struggling with it directly, with behaviors differing significantly between simulation and hardware.

The main weapon against the gap is domain randomization. During training, the simulator randomizes the physical properties of every virtual robot: body mass, actuator characteristics, sensor noise, floor friction. Toyota's team adds noise to encoder readings, which measure joint rotation, and IMU readings, which measure orientation and motion, while randomizing floor friction to replicate environmental variability. The policy that survives this chaos is the one that can handle a real robot it has never exactly seen before.

Figure combines domain randomization with high-frequency closed-loop torque control running at kilohertz rates on the robot, which compensates for errors in how the simulator models the actuators. The result, the company says, is zero-shot transfer: policies move from simulation to real hardware with no additional tuning, producing repeatable human-like walking across an entire fleet. Figure has demonstrated ten Figure 02 robots all running the same neural network with no tweaks, and sees that as evidence the approach scales to thousands of robots without per-robot engineering.

What still goes wrong

Teaching a robot to walk, by the numbers

Toyota's training time to stable walking1 to 2 hours
Virtual robots trained in parallelThousands
Simulated experience collectedYears, in hours
Torque control loop rate (Figure)kHz range
Figure 02 robots sharing one policy10, zero tweaks

For all the progress, the researchers are candid that simulation still cheats. Morita notes that a gait can look perfectly stable in simulation while secretly relying on tricks no real hardware can reproduce: oscillatory control commands, shuffling feet, abrupt leg movements. Toyota's workflow is to test every new walking model on the real robot, and when it fails, form hypotheses, apply countermeasures, retrain, and test again. Achieving walking on the real robot, Morita says, turned out to be much more difficult than stable walking in simulation.

Toyota attacks the gap from the other direction too, with a technique called Real2Sim: optimizing the simulator's actuator model using data collected from the real motors, identifying joint parameters like static, dynamic, and viscous friction plus inertia through black-box optimization, so that simulated joint trajectories match the real robot's. It is a reminder that sim-to-real is a two-way street. You randomize the simulation to be robust, and you calibrate the simulation to be honest.

Harder tasks expose the limits faster. Toyota's team also taught a humanoid to dribble a basketball, and found reward design far more difficult than for walking: dribbling requires modeling a constantly moving ball, brief and highly time-constrained contact, and precise launch speed and direction. Hand-designed rewards produced unnatural motions, so the team switched to imitating human motion-capture data, converting recorded joint angles to the robot's skeleton and rewarding the robot as its motion approached the reference. Even then, the biggest sim-to-real problem was perception: in simulation the ball's position is known perfectly, while the real robot must estimate it from a head-mounted camera, with recognition errors and latency breaking behaviors that worked virtually.

Why this matters beyond walking

Visualization of a neural network learning
Modern humanoid walking controllers are neural networks trained with reinforcement learning, not hand-written code.

Walking is the foundation, not the goal. The same pipeline, simulate broadly, randomize aggressively, transfer carefully, now underpins whole-body manipulation, disturbance recovery, and the dexterous tasks humanoid companies actually want to sell. The "Kick Me" demos at conferences like IROS 2026, where robots absorb shoves and keep their balance, are downstream of exactly this training: policies that have already been shoved a million times before breakfast.

The trajectory is clear. Training cycles that took months now take hours. Policies that needed per-robot tuning now transfer zero-shot across fleets. And the robots are starting to move less like machines executing trajectories and more like bodies that have learned, through exhaustive failure, how to stay upright in an unpredictable world. Which, when you think about it, is also how every toddler does it. The robots just get to skip the bruises.