NeurIPS 2026

Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

Zhiyuan Li1, Wenyan Yang1, Pekka Marttinen2, Joni Pajarinen1

Aalto University, Finland

1 Department of Electrical Engineering and Automation 2 Department of Computer Science

The usual score can't tell these apart. SIF can.

In short

Retargeting carries a motion from one body to another. A model can show the right action on the new body without carrying over anything from the clip it was given. We show that when data are sparse and the bodies differ, the standard ways of training cannot tell these two cases apart, and neither can the usual action-level score. Our measure, SIF, can. On animal motion, nine of the fourteen methods we tested score at or near the level of a method that ignores its source.

Read the abstract

Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, conditional-mean degeneration: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.

What does retargeting have to get right?

It has to carry over how this particular clip moves, not only which action it is. Here each method gets the same bird attack clip and must turn it into a king cobra attack.

Bird → King cobra, attack

One bird clip on top. Below it, the king cobra motion each method made from that clip. The outputs differ a lot, though every method got the same clip.

Why can't training tell the difference?

The animal data shows which bodies do which actions. It never shows which target motion belongs to which source clip. The two usual ways of training leave that choice open.

Training without pairs

Each dot is one motion. Training sees only how each body's motions are spread out, never which goes with which.

Now turn the target body's dots. Their shape does not change.

The training loss does not change either.

But every source dot now lands on a different target dot.

Many different maps fit the training data equally well.

gauge non-identifiability

Training with pairs matched only by action

One source clip. Several target clips of the same action would all be valid.

Each training step pairs it with one of them at random.

The model's answer is pulled toward the middle.

It settles on the average, the same answer for every source clip.

In a test where the correct answer is known, going from 2 to 50 training pairs per group does not help. The model keeps about 2% of the variation it should.

conditional-mean degeneration

How do we measure it?

Take a few source clips of the same action and look at their outputs on one target body. Source clips that differ more should give outputs that differ more.

Three source clips: same body, same action.

A line joins every pair. The thicker the line, the more the two clips differ.

A faithful model: the output lines match the source lines.

A model that ignores the source: its output lines no longer follow the source lines.

SIF asks whether outputs differ from each other the way their sources do.

We scored fourteen methods on 1,891 combinations of source animal, target animal and action. Nine score at or near the floor: the range of scores a method gets when it ignores its source.

source-blind floor

SIF on 1,891 animal test cases

  • Above the floor
  • Near the floor
  • At the floor

Raw scores with 95% intervals. Methods that cannot handle a case are scored on the rest (at least 1,413 cases). The paper also reports scores with clip length controlled.

The methods above the floor keep only part of the source. The two highest in this raw view barely change their outputs when the source changes, and they fall to or below the next two when clip length is controlled.

What losing the source looks like

Bird → King cobra, attack

Three different bird clips go in on the left. Each row shows one method's king cobra motions on the right. The top two rows sit at the floor: their outputs move alike whatever the source. The bottom two rise above it but keep only part of the source. The fourth keeps the order of the sources, but its outputs barely differ.

Why aren't action-level scores enough?

An action-level score checks that the target does the right kind of thing. It cannot check that the motion came from the source.

Action-level score on animals left out of training

held-out AUC

Lookup of stored target clips ANCHOR

0.681

Random clip, same action group

0.678

answers 177 of 200 queries

Random clip, same exact action

0.923

answers only 57 of 200 queries

A random clip with the right label scores as well.

What happens when the true answer is known?

For human motion moved to humanoid robots, a public retargeting tool gives a robot version of every human clip, which we treat as the true answer. Trained on those true pairs, a model keeps what makes each clip different. Trained without pairs, or with pairs matched only by the routine, it loses most of it.

These pairs come from a retargeting tool, not animators, and the robots are all humanoids.

One dance, six robots

Human

Unitree G1

Booster T1

Fourier N1

Stanford Toddy

EngineAI PM01

PAL Talos

The tool's output for each robot, which we treat as the true answer.

Two dancers, same routine, on the Unitree G1

Dancer 1

Dancer 2

Difference kept

True answer

from the public retargeting tool

reference

Trained on true pairs

99%

Trained on random pairs

paired at random within the same routine

3.5%

Trained without pairs

never sees a pair

1.2%

How much of the difference between performers each model keeps, compared with the true answer, over all 26 routines on the Unitree G1 (LAFAN1 study, denser setting). The two dancers shown are one of those routines. They are not in step, so part of the difference you see is timing.

Human motion: LAFAN1 by Ubisoft La Forge, CC BY-NC-ND 4.0. Robot motion made with GMR, a public retargeting tool (MIT License); robot models as distributed with GMR, credited to their original sources in its repository.

What do the methods do with other animals?

Each video fixes one source animal, one target animal and one action. Three source clips play on top, and each row below is one method's three outputs.

A row whose three outputs move alike has lost what made the source clips different.

How can I check my own method?

SIF needs only NumPy. Give it your source clips and your outputs, grouped by source animal, target animal and action.

pip install git+https://github.com/LiZhYun/NeurIPS2026-Cross-Skeleton-Retargeting
from sif import score_groups

# one group per source animal, target animal and action (format: docs/sif.md)
groups = [{"sources": [src_1, src_2, src_3], "outputs": [out_1, out_2, out_3], "block": "Horse"}]
result = score_groups(groups)
print(result.sif, result.ci)

View the code on GitHub

How do I cite this?

@inproceedings{li2026retargeting,
  title     = {Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models},
  author    = {Li, Zhiyuan and Yang, Wenyan and Marttinen, Pekka and Pajarinen, Joni},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}