← All articles

HiPHI: 617 hours of human movement for humanoid robots and what is behind that number

Topics: Robots, AI, Science

A motion capture studio: a ring of cameras on tripods, a glowing skeleton of dots in a walking pose in the middle, a box and a chair outlined in light next to it, and the silhouette of a humanoid robot behind

Short answer: HiPHI is a dataset of recordings of human movement for training humanoid robots, collected by the company Noitom Robotics. 132 people moved in a studio with an optical motion capture system which, according to the authors, tracks markers on the body with sub-millimetre accuracy, and in part of the recordings they worked with real objects, whose movement was recorded together with the person's. The dataset description gives a figure of 617.5 hours, but this is 308.7 hours of original recordings plus their mirrored copies, in which left and right are swapped. The authors, mostly Noitom employees, showed in their own tests that the humanoid robot Unitree G1 finds it easier to reproduce movements from HiPHI than from several other datasets, though not for every type of movement. The authors themselves name the limitations: there is always one person in the frame, forces and contact were not measured, and everything was filmed in a studio. Access to the dataset is granted on request, and its licence is non-commercial: for research, education and evaluation. The news hook: on October 7, 2026, a white paper on the dataset appeared in the IEEE Spectrum feed; IEEE Spectrum and the publisher Wiley released it with Noitom itself as sponsor. Below is how the dataset is built, what is behind the numbers, what the tests showed and how to read similar claims.

Why robots need recordings of human movement

A humanoid robot needs to keep its balance, walk, change posture, and carry and push objects. It can be taught this from recordings of how a person does it: the robot reproduces the movement in a computer model, and then what it has learned is transferred to the real machine.

In their paper, the HiPHI authors explain why existing data is not enough. Videos from the internet show many different actions, but from video it is hard to determine body position precisely. Laboratory motion recordings are precise but usually cover a narrow set of actions. And in many recordings of interaction with objects, the object itself is missing: the authors give the example of "a person sitting in the air", where the body is recorded but the chair is not.

Optical motion capture, to simplify, means filming a person with markers on the body using many cameras, from which a computer reconstructs the body's position at every moment.

How HiPHI is built

According to the research paper on arXiv (version of September 8, 2026) and the dataset card on Hugging Face:

What How much
Total in the release 617.5 hours, about 200 million frames
Of which original recordings 308.7 hours, the rest are mirrored copies
Movements without objects 371.8 hours
Movements with objects 245.7 hours, 40 objects of 12 kinds weighing from 0.45 to 6.25 kg
Performers 132 people, 76 men and 56 women
Recording rate 90 frames per second
Semantic labels 214 types of movement in 22 groups
File size 215 GB

Movements were selected not by scripts such as "wipe sweat from the forehead" but using FrameNet, a linguistic dictionary of situations and actions. The authors explain why: two different scripts, for example "raise a hand to wipe sweat" and "raise a hand to block the sun", can produce almost the same movement. So they took verbs of movement from FrameNet and varied direction, speed, amplitude, posture and object, to cover different movements rather than different stories. The 50 most frequent types of movement account for 53.7% of the total duration.

The objects in the recordings are real: furniture, containers, cleaning tools and sports equipment. The authors stress that the weight, friction and resistance of a real object change how a person moves, and miming an action without an object does not capture this.

What is behind the 617 hours

The authors state plainly that the 617.5 hours were obtained by left-right mirroring of 308.7 hours of original recordings, as was done in another dataset, BONES-SEED. Every recording has a pair with the suffix "mirror", in which the left and right sides are swapped. A mirrored copy gives the robot the movement "with the other hand", but it is not a new recording of a new person.

This is not hidden: both the paper and the dataset card say so. But the white paper's description gives only 617.5 hours, without mentioning the mirrored copies. When comparing datasets by hours, it is important to look at how each one counts them. According to the authors, all the tests in the paper were run on the original recordings without mirrored copies.

What the tests showed

The tests were run by the authors themselves. Movements were transferred to the humanoid robot Unitree G1 in a computer model, and how well the robot reproduced them was compared against other datasets: AMASS, LAFAN1, Motion-X++ and BONES-SEED.

  • Reproducing body movements. With the same amount of data, 3 and 20 hours, the robot trained on HiPHI, according to the authors, succeeded more often and learned faster.
  • More data, less error. When the amount of HiPHI recordings used for training was increased from 3 to 300 hours, the error in reproducing movements on other datasets decreased steadily.
  • Movements with objects. Here the picture is mixed. In accuracy of body movement, HiPHI is better for most types, but for pushing, HiPHI's mean joint position error, 99.2 mm, is larger than that of the OMOMO dataset, 66.1 mm. In accuracy of the object's movement, HiPHI is better for kicking and pushing, and for carrying only in the object's rotation.
  • A real robot. The skills trained on HiPHI were transferred to a real Unitree G1, which, according to the authors, ran, sat down, crawled, carried a box, did a flip and pulled a suitcase.

On the dataset card, Noitom states that the paper was accepted at CoRL 2026, a conference on robot learning. This is information about publication, not a check of the results: the sources for this article contain no independent confirmation of the test results.

What the dataset does not provide

The authors name the limitations in the paper themselves:

  • each recording contains one person; there is no interaction between people;
  • movement is recorded, not forces or contact: the dataset does not measure pressing force or the feel of an object;
  • everything was filmed in a studio, not in real homes, workshops or shops.

There is also a practical limitation: the dataset is distributed under the ModalityNet Open Research License, which, according to the authors, is non-commercial and intended for research, education and evaluation. For commercial use, Noitom offers a separate licence. The dataset can be downloaded after requesting access on Hugging Face. The recordings are in a skeleton motion format, and the card states directly that they need to be retargeted and checked for a specific robot.

Who is behind it

Most of the paper's authors are from Noitom Robotics, with several more from the National University of Singapore, Tsinghua University, the Hong Kong University of Science and Technology and the University of Hong Kong. The white paper was released by IEEE Spectrum and Wiley, and its sponsor is Noitom Robotics. In other words, the dataset and tests were prepared mainly by Noitom employees together with university co-authors, and Noitom sponsored the white paper about the dataset. This does not make the results wrong, but they are best read as a description of a product by its maker.

On ethics, the authors write that data collection was approved by an ethics committee, all performers were adults, gave written consent, were paid for each capture session and may ask for their data to be removed at any time. In the release, performers are anonymized: each has an anonymized identifier, and the personal details given are height, weight and gender.

Five questions about any "hours of data"

These questions are useful when a robot supplier or researcher boasts about the volume of data. A manager reading a proposal can ask them, and look for the answers in the paper, the dataset card or the contract.

  1. How many of these hours are original recordings? In HiPHI, half the hours are mirrored copies.
    • How to check: the data description gives the number of hours without mirrored copies and without other artificial additions.
  2. What exactly was recorded? Body position, object movement, forces, contact, video.
    • How to check: the description lists the types of data; if you need work involving force, for example assembly or packing, and the data has no forces, that is a limitation for your task.
  3. Where was it recorded? In a studio or in conditions similar to yours.
    • How to check: the description states where the recording took place; for work in a workshop or warehouse, ask whether the robot was tested there.
  4. Who checked the results? The authors themselves or independent researchers.
    • How to check: whether there is a comparison made by someone other than the creator of the data, and where the dataset turned out worse than others.
  5. Can the data be used in a commercial product?
    • How to check: the licence states directly whether commercial use is allowed; HiPHI's main licence is non-commercial, and for commercial use Noitom offers a separate one.

If the answers do not satisfy you: there is no data of the kind you need, there were no tests in conditions similar to yours, or the licence does not allow your use, ask the supplier to demonstrate the robot on your task in test mode, in a fenced-off area and with no people nearby, before buying. If such a demonstration is impossible or the licence does not fit, this data should not be relied on when deciding to buy, and the claimed hours should not be taken as an argument.

I wrote about how robots and AI are changing jobs done with your hands in the article Will AI replace plumbers, electricians and dentists: what is changing in jobs done with your hands.

Summary

HiPHI is a dataset of recordings of the movements of 132 people in a studio, with precise tracking of the body and, in part of the recordings, of real objects. The figure of 617.5 hours includes mirrored copies: there are 308.7 hours of original recordings. According to the authors' own tests, the Unitree G1 robot finds it easier to reproduce movements from HiPHI than from several other datasets, although for pushing the dataset was outperformed by another in accuracy of body movement while being better in accuracy of object movement, and on the real robot the trained skills worked. The authors name the limitations themselves: one person in the frame, no forces or contact, studio only, and the dataset's licence is non-commercial. The dataset and tests were prepared mainly by Noitom employees, and Noitom sponsored the white paper, so the main skill here is a reader's one: ask how the hours were counted, what was recorded, where and who checked it.

I work on AI agents and automation. If you would like to see my projects or discuss your own task, take a look at my portfolio.

Sources