← All articles

A robotic hand has learned to walk on its own fingers. Part 2: how the training works and what produced the gain

Topics: Robots, AI, Science

A translucent robotic hand resting on its fingertips, with a coil spring of differing thickness running from each fingertip to a point on the surface, over a faint coordinate grid

This is a continuation. Part 1 covers what the robot can do: crawl on its fingertips, recover from falls, press keys and push a cube to a target, with all the numbers. This part is about how it was taught, and why the authors give so much space to a single term in the reward formula. There will be several formulas, each explained in words.

Three words right away, since they come up constantly below. A policy is a trained behaviour program: it looks at sensor readings and issues the next command. Reinforcement learning is a way of obtaining one: the policy is run many times in a model of the world and rewarded for the desired outcome with a number called the reward. PPO is a widely used method of such training, which changes the policy in small steps so that it does not fall apart from abrupt edits.

There is one source: the preprint arXiv:2609.17172. All the numbers about this work below come from it, marked as to whether they were obtained in simulation or on hardware. The timeframes in the automation example at the end of the article have nothing to do with the preprint: that example is mine.

Why you cannot take the ready-made recipe for four-legged robots

Training gaits for legged robots is long-established. But one familiar technique does not apply here because of how the task is built, and the authors deliberately set aside a second.

Symmetry. A four-legged robot has left-right mirror symmetry, and it is used actively: a penalty for asymmetry is added to training, examples are mirrored, networks are built that behave identically under reflection. All of this requires a mirroring of states and actions that leaves the reward unchanged. A hand with an opposed thumb and four unequal fingers has no such mirror.

A schedule of steps. That one is the authors' choice, not a prohibition. Rewards for legged robots often prescribe timing: which leg should be in the air when, which on the ground, with what phase offsets between them. The authors went a different way: their term anchors place rather than time. Each fingertip is pulled toward its own point, captured from the settled stance, while when and which finger steps is left to the policy.

There is a third difficulty, purely geometric: the palm rests at a tilt in the stance, so the body's own coordinate system is not a level reference for motion commands.

A frame with the tilt removed

The reference stance is obtained like this: in simulation, with the 80-gram payload, fixed joint targets are held for 4 seconds, after which the settled root pose and joint positions are saved.

Then a frame V is introduced. It comes from the body frame through a rotation that levels this stance and aligns the "forward" axis with the horizontal vector from the palm to the centre of the distal-link origins of the four fingers from index to little. Formally it is the current body rotation multiplied by a fixed corrective rotation.

A subtlety the authors note separately: from then on V follows the body's rotation rather than staying gravity-levelled at all times. That is, the palm tilt characteristic of the stance is removed, while body motion is not suppressed.

The reward: virtual springs around the footprint

The main term the authors call the footprint objective. It works as follows.

Each of the five fingertips has its own reference point. Each environment captures it at the first reward evaluation after initialization and then holds it fixed across resets. The fingertip position is taken relative to the body, in that same frame V, so the points travel with the body. For fingertip number j, the difference between its current position and its reference is taken, and the following is computed:

r̄ = Σ (p_j − p_j)ᵀ · W · (p_j − p_j), where W = diag(0.25; 1; 1)**

In words: for each finger the three components of the deviation from its point are taken (along the direction of travel, sideways and vertically), each is squared, multiplied by its own weight, and summed across all five fingers. The weight along the direction of travel is 0.25, sideways and vertically it is one, so lateral and vertical deviations are penalised four times as strongly as fore-aft ones.

This enters the reward with a coefficient of minus 30, that is, as a penalty.

Why these are springs: an expression of the form "a weight multiplied by the square of a displacement" is exactly the formula used for the energy stored in a spring. So each fingertip is in effect tied to its point by three springs of differing stiffness: a soft one along the direction of travel and stiff ones sideways and upwards. The soft spring along the direction of travel makes a step cheap, while the stiff lateral and vertical ones discourage the hand from sprawling.

Two caveats from the authors that are easy to miss. First: these stiffnesses are reward weights, not controller gains, that is, they affect training rather than how the hardware executes a command. Second: the reference points move with the body and prescribe neither ground locations nor a footfall sequence.

Two auxiliary terms

For lifting. This term nudges a stepping frequency that grows with the commanded speed. Internally a phase is tracked, increasing each step by 2π·f·Δt, where Δt is 20 milliseconds, and the frequency f linearly maps a speed from 0.02 to 0.16 metres per second into a frequency from 1 to 3.5 hertz, with clamping at the ends of the range.

The vertical velocity of a fingertip is multiplied by a complex exponential of that phase and averaged by a moving average with a time constant of 0.3 seconds. Strip away the mathematics and this construction measures how well a finger's lifts fall in step with the required frequency, and rewards falling in step. The phase itself is not shown to the policy; it does not see it.

A lift counts only under three conditions: the tip is airborne (force below 0.15 newtons), there has been contact for that finger since the last phase wrap, and the commanded speed is at least 0.02 metres per second.

The weight of this term changes over the course of training: 6.25 for the first 18,000 per-environment steps (that is 750 of the 8,000 training iterations), then it decays geometrically to 0.625 over the next 36,000 steps and stays there.

For direction. The component of a fingertip's velocity opposite to the command is penalised, and only under two conditions at once: the force at the tip is below the threshold (the finger is airborne) and a non-zero speed has been commanded. The point: planted fingers need to move backward relative to the body, otherwise they do not propel it forward, whereas swinging backward through the air is of no use.

A separate detail about the order of training: yaw-rate tracking is not switched on at the start but at that same 18,000-step mark, once a forward gait has formed, so that the policy does not learn stepping and turning at once.

A simulator calibrated against hardware

The policies act through the hand's built-in position controller, so the simulator was calibrated against the measured response to precisely those filtered position commands: frequency sweeps of the joints, loaded fingertip pulls, timed response measurements.

What was measured What came out
Joint stiffness multiplier relative to the uncalibrated model 15 to 22 (18.7 plus or minus 2.5 on average)
Kinetic fingertip friction 0.88 (range 0.80 to 0.97)
Closed-loop delay about 19 ms
Filtered joint speed 2.8 and 2.5 rad/s (flexion and abduction)
Command filter cutoff 3 Hz

The stiffness multiplier of 18.7 relative to the original model deserves attention on its own: it shows how far the calibrated model diverged from the original one. What the transfer would have looked like without this calibration was not tested in the work.

What exactly was trained

The policy is an ordinary fully connected feedforward network trained with PPO. Onboard it runs 50 times a second and outputs not angles but an increment to the joint position targets: the increment is computed as a scale multiplied by the hyperbolic tangent of the network's output, and added to the previous target while respecting the joint limits. The scale differs per task: 0.060 radians for crawling and recovery, 0.028 for the keyboard, 0.040 for pushing.

The inputs are shared across policies: 20 joint angles relative to reference angles, a unit vector for the direction of gravity, angular velocity and the previous policy output, 46 quantities per step in total. Task inputs are added to these: for crawling that is the planar motion command (two numbers) and the yaw-rate command (one). Each quantity retains its last eight values, after which everything is concatenated and normalised using fixed training statistics. In the training table the actor input is listed as 392 with an output of 20, which matches 49 quantities at eight values each.

The training setup for the crawl policy:

  • 4096 parallel environments (that is, 4096 copies of the task computed at once), 20-second episodes;
  • actor hidden layers of 512, 256, 128 with ELU activation;
  • physics at 200 Hz, one policy action per four physics steps;
  • 8,000 PPO iterations plus another 5,000 for adaptation to the payload and contact conditions;
  • in a PPO update, 24 steps per environment, five epochs, four minibatches, a clip ratio of 0.2, an adaptive learning rate starting at 0.001 (this group of settings restrains how big a single edit to the policy can be: clipping reduces the benefit of overly large changes, so that training does not jump around);
  • a discount factor of 0.99, generalized advantage estimation with a coefficient of 0.95, an entropy coefficient of 0.005 (this group sets how far ahead the policy looks when weighing consequences, and how much it is allowed to try different things rather than repeat what it has found);
  • commands during training: forward speed from 0 to 0.16 m/s, lateral from minus 0.16 to 0.16 m/s, turning from minus 0.45 to 0.45 rad/s;
  • an episode terminates on palm contact, roll beyond 35 degrees or pitch beyond 30 degrees.

Two details that matter for understanding the setup. First: two networks take part in this kind of training. The actor is the policy that acts. The critic is a second network that only assesses how well things are going, and is needed only during training. In training the critic additionally sees the base state, joint torques and fingertip contact forces, that is, what the actor does not have on hardware. This is a standard technique, giving the assessing part more information than the acting one. Second: the randomisation during training covers fingertip friction, effort scale, offsets of the payload's centre of mass, palm mass, its centre of mass and the tilt sensor's bias, while the actuator gains vary around the calibrated stiffness.

What turning terms off showed

This is the most careful part of the work. Seven configurations were compared: the authors' own, four with terms removed, and two baselines based on rewards from an off-the-shelf walking task for the four-legged ANYmal-D robot, one raw and one tuned for this hand.

How the comparison was set up: each configuration was trained on the same twelve random seeds, with 4096 environments and 8,000 iterations, after which each policy was evaluated in 256 flat-ground episodes. It is noted separately that the weights of the authors' formulation were fixed by the hardware campaign before the comparison and were not retuned, while the tuned baseline was obtained by screening 24 random reward-weight settings. That is an honest caveat against their own case: the baseline was tuned for this comparison, while the authors' weights were not retuned for it, even though they had earlier been chosen through work on hardware.

The results (all in simulation):

  • Against the tuned baseline, the authors' reward is faster by 0.65 cm/s on average, with a 95% confidence interval from 0.26 to 1.02, and faster in 10 of 12 seeds.
  • Removing the footprint term reduces speed by 0.79 cm/s (interval from 0.37 to 1.19) and lowers five-finger participation.
  • The auxiliary terms for lifting and for direction: the comparison did not establish an independent speed benefit from them, separately or together. Moreover, removing the direction term alone increases five-finger participation by 0.31 (interval from 0.07 to 0.52). The authors keep it in the reported formulation for an honest reason: it is what the policies that went to hardware were trained with.

The confidence interval here should be read literally: if it does not include zero, the gain counts as established; if it does, the effect is not established. So the phrase "did not establish an independent gain" means precisely an absence of proof, not proof of absence.

Why looking at speed alone is not enough

The authors introduce two measures beyond speed.

Five-finger participation: the fraction of episodes in which every fingertip touched down at least three times and spent at least 5% of the episode in contact. In plain terms, a measure of the hand not crawling on three fingers while dragging two.

Contact posture: a touch is classified by the tilt of the pad axis above horizontal toward the nail. The worst-tip planting fraction is the fraction of contact time below 64 degrees for the finger that rests on its nail side most often.

Why this is needed is visible from the comparison: the tuned baseline has higher mean five-finger participation (though that difference is uncertain), but in every run it plants fingers on their nails more. The authors' formulation has higher participation without nail contacts: plus 0.35, with an interval from 0.18 to 0.53. That is, speed alone would have hidden what exactly the hand touches the floor with. Wear on the fingers when resting on the nail side is a possible practical risk rather than a measured result: wear was not measured in the work.

Honesty is kept here too: at the deployed weight the footprint term shifts the mean contact tilt by 8.8 degrees toward the nail side (interval from 1.4 to 15.9). So it improves five-finger participation without improving every aspect of contact posture.

One more result about weights: doubling the footprint weight gives a higher worst-tip planting fraction than doubling the lift weight (a paired difference of 0.22, interval from 0.11 to 0.33, in 10 runs of 12), while the difference in speed between them is uncertain. Above the deployed weight the planting fraction keeps improving, mean five-finger participation peaks at twice the deployed weight and decreases again at 3.3 times. All these higher weights were tested only in simulation.

Recovery and working with objects

The recovery policy was trained by the same method: a 20-second episode starts with the hand lying on its side at a random heading, the reward penalises palm tilt away from level and rewards reaching the height and pose of the crawl stance, with no early termination. Once upright, the policy keeps the hand balanced through continuous joint motion, and then a smooth transition over 1.5 seconds brings it to a static stance. On hardware this transition is started by an upright-state detector: tilt relative to the stance below 10 degrees and angular speed below 4 rad/s for half a second.

The keyboard and pushing policies use the same calibrated simulator, the same history of measurements and the same PPO implementation. The actor has three hidden layers of 512, 256, 128, the critic 256, 128, 64, and during training the critic additionally receives privileged simulation state.

The camera model used in training for pushing deserves separate attention: 60 Hz on a sample-and-hold scheme, a latency of one to three control steps, 3% dropped frames and 2 millimetres of position noise. That is, the imperfections of a real camera were built into training in advance.

What is worth taking away

Three things that apply far beyond robotics.

First: when a familiar technique does not fit because of how the task is built (here, the absence of symmetry), it is worth looking at the framing rather than only at the method of solving. The authors replaced "a schedule of steps" with "a place for each finger" and arrived at a formulation that does not require the parts to be alike.

Second: the gap between the model and reality was closed here by measurement. A stiffness multiplier of 18.7 shows how large that gap was for joint stiffness alone.

Third: it is worth choosing your measure of success in advance, and more than one. Had the authors measured only speed, the difference in what the hand touches the floor with would have gone unnoticed.

How to apply that to choosing your first task for automation. Take a task whose result can be counted as a number, and write down two measures in advance: the main one (for example, how many requests are handled per shift) and a guard measure that catches the way the result is achieved (for example, the share of requests that later brought a complaint or a redo). Then you need a baseline: both measures are counted for the month before the change and the month after, from your own records, with no arguing about definitions. By those two measures the decision counts as a success if the main one grew and the guard one did not get worse. That is not yet a conclusion about payback: alongside it you need to put the cost of the rollout and of ongoing support and compare them with the gain in money. It is worth revisiting the decision in three cases: the guard measure got worse (that is, the gain came at the expense of quality), the main one did not move in two months, or the cost of support turned out to exceed the benefit obtained. And if there is nothing to count the guard measure with at all, it is better to take a different task for your first automation, because otherwise there will be nothing to check against.

I work on AI agents and automation. If you would like to see my projects or discuss your own task, take a look at my portfolio.

Sources

  • Amirhossein Kazemipour, Hehui Zheng, Robert Katzschmann. Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand. Preprint, arXiv:2609.17172, submitted 15.09.2026, sections III, IV and VI-A. https://arxiv.org/abs/2609.17172