NeurIPS 2026
Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models
Aalto University, Finland
1 Department of Electrical Engineering and Automation 2 Department of Computer Science
The usual score can't tell these apart. SIF can.
In short
Retargeting carries a motion from one body to another. A model can show the right action on the new body without carrying over anything from the clip it was given. We show that when data are sparse and the bodies differ, the standard ways of training cannot tell these two cases apart, and neither can the usual action-level score. Our measure, SIF, can. On animal motion, nine of the fourteen methods we tested score at or near the level of a method that ignores its source.
Read the abstract
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, conditional-mean degeneration: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.
What does retargeting have to get right?
It has to carry over how this particular clip moves, not only which action it is. Here each method gets the same bird attack clip and must turn it into a king cobra attack.
Bird → King cobra, attack
One bird clip on top. Below it, the king cobra motion each method made from that clip. The outputs differ a lot, though every method got the same clip.
Why can't training tell the difference?
The animal data shows which bodies do which actions. It never shows which target motion belongs to which source clip. The two usual ways of training leave that choice open.
Training without pairs
Each dot is one motion. Training sees only how each body's motions are spread out, never which goes with which.
Now turn the target body's dots. Their shape does not change.
The training loss does not change either.
But every source dot now lands on a different target dot.
Many different maps fit the training data equally well.
gauge non-identifiability
Training with pairs matched only by action
One source clip. Several target clips of the same action would all be valid.
Each training step pairs it with one of them at random.
The model's answer is pulled toward the middle.
It settles on the average, the same answer for every source clip.
In a test where the correct answer is known, going from 2 to 50 training pairs per group does not help. The model keeps about 2% of the variation it should.
conditional-mean degeneration
How do we measure it?
Take a few source clips of the same action and look at their outputs on one target body. Source clips that differ more should give outputs that differ more.
Three source clips: same body, same action.
A line joins every pair. The thicker the line, the more the two clips differ.
A faithful model: the output lines match the source lines.
A model that ignores the source: its output lines no longer follow the source lines.
SIF asks whether outputs differ from each other the way their sources do.
We scored fourteen methods on 1,891 combinations of source animal, target animal and action. Nine score at or near the floor: the range of scores a method gets when it ignores its source.
source-blind floor
SIF on 1,891 animal test cases
- Above the floor
- Near the floor
- At the floor
Raw scores with 95% intervals. Methods that cannot handle a case are scored on the rest (at least 1,413 cases). The paper also reports scores with clip length controlled.
The methods above the floor keep only part of the source. The two highest in this raw view barely change their outputs when the source changes, and they fall to or below the next two when clip length is controlled.
What losing the source looks like
Three different bird clips go in on the left. Each row shows one method's king cobra motions on the right. The top two rows sit at the floor: their outputs move alike whatever the source. The bottom two rise above it but keep only part of the source. The fourth keeps the order of the sources, but its outputs barely differ.
Why aren't action-level scores enough?
An action-level score checks that the target does the right kind of thing. It cannot check that the motion came from the source.
Action-level score on animals left out of training
held-out AUC
Lookup of stored target clips ANCHOR
Random clip, same action group
Random clip, same exact action
A random clip with the right label scores as well.
What happens when the true answer is known?
For human motion moved to humanoid robots, a public retargeting tool gives a robot version of every human clip, which we treat as the true answer. Trained on those true pairs, a model keeps what makes each clip different. Trained without pairs, or with pairs matched only by the routine, it loses most of it.
These pairs come from a retargeting tool, not animators, and the robots are all humanoids.
One dance, six robots
Human
Unitree G1
Booster T1
Fourier N1
Stanford Toddy
EngineAI PM01
PAL Talos
The tool's output for each robot, which we treat as the true answer.
Two dancers, same routine, on the Unitree G1
Dancer 1
Dancer 2
Difference kept
True answer
from the public retargeting tool
reference
Trained on true pairs
Trained on random pairs
paired at random within the same routine
Trained without pairs
never sees a pair
How much of the difference between performers each model keeps, compared with the true answer, over all 26 routines on the Unitree G1 (LAFAN1 study, denser setting). The two dancers shown are one of those routines. They are not in step, so part of the difference you see is timing.
Human motion: LAFAN1 by Ubisoft La Forge, CC BY-NC-ND 4.0. Robot motion made with GMR, a public retargeting tool (MIT License); robot models as distributed with GMR, credited to their original sources in its repository.
What do the methods do with other animals?
Each video fixes one source animal, one target animal and one action. Three source clips play on top, and each row below is one method's three outputs.
A row whose three outputs move alike has lost what made the source clips different.
How can I check my own method?
SIF needs only NumPy. Give it your source clips and your outputs, grouped by source animal, target animal and action.
pip install git+https://github.com/LiZhYun/NeurIPS2026-Cross-Skeleton-Retargeting
from sif import score_groups
# one group per source animal, target animal and action (format: docs/sif.md)
groups = [{"sources": [src_1, src_2, src_3], "outputs": [out_1, out_2, out_3], "block": "Horse"}]
result = score_groups(groups)
print(result.sif, result.ci)
How do I cite this?
@inproceedings{li2026retargeting,
title = {Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models},
author = {Li, Zhiyuan and Yang, Wenyan and Marttinen, Pekka and Pajarinen, Joni},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}