Newsletter

UniMate Can Animate Any Rig From a Text Prompt. Read the Data Licence First.

UniMate is a Princeton-led model that takes a text prompt plus any skeleton and produces animation for it, without retraining per rig or running test time optimisation. It covers bipeds, quadrupeds, birds, fish, insects, snakes plus articulated objects. The paper landed at SIGGRAPH Asia 2026. Code is out under MIT. Four pretrained checkpoints sit on Hugging Face at 26.1 GB total.

Then you reach the data section. UniML3D was built from Mixamo, Objaverse-XL plus Truebones ZOO, so the motion those MIT checkpoints learned from carries Adobe’s terms of use, thousands of per object licences, then a commercial licence you normally pay for. The authors wrote every bit of this down clearly. Almost nobody repeating the demo clip mentions it.

Best for technical artists and gamedev folks who want to test text driven animation on unusual rigs today. Not ideal for anyone who needs a clean commercial licence chain before shipping.


Somebody in your feed posted a dragon walking.

Then a crab. Then a fish, a snake, a desk lamp with an arm. All of them moving plausibly, all from a sentence typed into a box, all from one model that was never trained on any of those specific rigs.

It is a great clip. It deserves the attention it got.

What the clip does not tell you is that the motion library underneath it includes a commercial product you would normally buy a licence for. The people who built UniMate said so in writing, on the first page of their own README, weeks before the clip went around.


Why the Caveat Keeps Falling Off

That gap between what a project documents plus what the internet repeats is becoming the most reliable story in open source AI.

We have written it about Mistral’s open voice model, where the benchmark claim travelled and the conditions did not. We wrote it about Alibaba quietly closing Qwen, where a licence change reached almost nobody who had built on the open version.

UniMate is the cleanest example yet, because the research team did nothing wrong. They published the caveat. It simply did not survive the trip to your timeline.

There is a mechanical reason for that. A demo clip is self contained. You watch it, you understand it, you repost it in four seconds. A licence note requires opening a README, scrolling past the installation instructions, then caring about a sentence that names three asset libraries you have possibly never used. One of those travels. The other does not.

So the pattern repeats regardless of how honest the original authors were. Which means the useful work is no longer finding dishonest vendors. It is reading the thing everybody else scrolled past.


What UniMate Actually Is

Here is the verified spec, pulled from the arXiv paper, the GitHub README plus the Hugging Face model card.

ItemDetail
PaperarXiv:2609.05415v1, submitted 4 September 2026
VenueSIGGRAPH Asia 2026
AuthorsLinzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz
AffiliationsPrinceton University, UC Berkeley, MIT, Nanyang Technological University
ArchitectureTADiT, a topology aware diffusion transformer using flow matching
Code licenceMIT
Code released6 September 2026
Dataset released30 August 2026, raw assets plus processing pipeline
Checkpoints4 variants on Hugging Face at Linzhan/UniMate, 26.1 GB total, MIT
Largest checkpoint74.1M parameters
Training dataUniML3D, 13,006 text paired motion sequences, roughly 20 hours
Text prompts3,584 unique captions
Data sourcesMixamo, Objaverse-XL, Truebones ZOO
DemoInteractive 3D demo live since 1 August 2026
Downloads87 in the last month at time of writing

Every figure there came from a primary page rather than a summary. The star count is deliberately absent, because two sources disagreed about it by a wide margin. No claim in this piece needs one.


How It Animates Rigs It Has Never Seen

This paper exists because skeletons are not interchangeable. A dog has a different joint count, hierarchy plus proportion set than a human. A crab shares almost nothing with either.

Most text to motion work handles that by training one model per skeleton family, or by retargeting motion from a rig it knows onto a rig it does not. Both approaches break down the moment you hand them something unusual.

UniMate takes a different route. It treats the skeleton as a graph, then makes the model read that graph as part of its input.

The three pieces that make topology a first class input

Graph aware attention bias is the first. The model learns biases from pairwise joint relations plus geodesic distances along the skeleton. Attention between two joints therefore reflects how far apart they sit in the actual kinematic tree rather than their arbitrary index order.

Spec-RoPE is the second, which is the clever one. Standard rotary position embedding assumes a sequence. A skeleton is a tree. These authors generalise RoPE to arbitrary kinematic trees using the graph Laplacian spectrum, buying translation invariance plus permutation equivariance. In plain terms, renumbering the joints does not change the output. Nor does the model care where in the hierarchy a limb happens to be listed.

A global topological conditioner is the third. It modulates every transformer block through AdaLN-Zero, adapting layer statistics to whatever skeleton arrived. Attention is factorised into separate joint and temporal branches, so the model reasons about body structure then about time in different passes.

Taken together those let one set of weights produce motion for a rig it never saw during training, at inference time, with no optimisation loop. That claim is the whole paper. It is also a real contribution to graphics rather than a repackaged chatbot trick.

What it was measured against

Three baselines appear in the comparison. AnyTop is a prior topology agnostic method limited to a small animal corpus. How to Move Your Dragon is a cross topology retargeting approach. AnimaX is a skeleton free animation baseline.

Those choices tell you where the field sat before this. One method generalised across topology but only within a narrow set of animals. Another moved motion between rigs rather than generating it fresh. The third sidestepped skeletons altogether. UniMate reports improvements over all three on motion quality, generalisation to unseen skeletons, runtime efficiency plus zero shot cross topology transfer.

Hold that list lightly for now. We come back to why in a moment.


Where the Weights Actually Live

This part took three checks to pin down, so it is worth stating precisely.

The GitHub README lists pretrained checkpoints under TODO. Read only that, then you would conclude the model cannot be run. Plenty of people appear to have concluded exactly that.

Those checkpoints are public regardless. They sit on Hugging Face under Linzhan/UniMate, four variants totalling 26.1 GB, released MIT, with working inference instructions on the model card. The recommended one is unimate_uniml3d_f60_v2 at 74.1M parameters. There is also a Mixamo only humanoid checkpoint at 47.8M parameters for anyone who wants the narrower case.

The GitHub README does not link to any of them. Hugging Face does not appear in the repo’s release notes either. So the project is fully runnable today while its own front page implies otherwise, which is a documentation gap rather than a scandal. If you went looking once and gave up, go back.

One number is quietly instructive. Those checkpoints saw 87 downloads in the last month. Compare that against the attention the demo clip has pulled across X plus GitHub. A lot of people are saving this. Very few are running it. We saw the same split when we counted which Claude Code repos people actually run versus which ones they star.


The Licence Stack Nobody Is Repeating

Now the part that decides whether you can use this for money.

Code is MIT. Checkpoints are MIT. Both layers are clean, plus the team deserves credit for not inventing a bespoke research licence with a non commercial clause, which is the usual move.

UniML3D is where it gets complicated. The README says so directly, naming all three sources along with the terms each one keeps.

Three sources, three different problems

Mixamo assets stay governed by Adobe’s Mixamo terms of use. Mixamo is free to use with an Adobe account, plus its terms permit use of the characters and animations inside projects. They also constrain redistribution of the assets themselves, which is a separate question from using them in a game.

Objaverse-XL assets are governed by the licence attached to each original object. Not one licence. Thousands of them, gathered from across the web, ranging from permissive Creative Commons through to terms that forbid commercial use entirely. Working out which objects contributed to which motions is not something you can do from outside the project.

Truebones ZOO motions are governed by Truebones’ commercial licence. Truebones sells motion capture libraries. That is the business. The ZOO pack is a paid product, which is also the reason UniMate can animate a scorpion convincingly.

Why that combination is awkward

So the honest summary runs like this. Those weights are MIT, while the motion knowledge inside them was distilled from a paid commercial library plus an unmapped pile of per object terms.

Whether model weights trained on licensed data inherit that licence is an open legal question in 2026 rather than a settled one. We walked through the same unresolved territory when GitHub changed its Copilot training opt out. Nothing since has made the answer clearer.

For a hobby project this is interesting trivia. For a studio putting animation into a shipped product, your legal team will want to read the Truebones licence before you commit a pipeline to these checkpoints. That conversation costs an afternoon now. It costs considerably more after you have built tooling around the output.

How to check this on any model you download

UniMate is not unusual here, which is the reason this section exists at all. Almost every interesting 3D generation model of the last two years trained on some mixture of scraped web assets plus purchased commercial libraries, because that is where usable rigged motion lives. Text models get scraped books. Image models get scraped art. Motion models get asset store products, since nobody else bothered to capture a scorpion walking.

So build the habit. Open the model card, scroll to the training data section, then look for three specific things.

First, whether the data sources are named at all. Plenty of model cards say nothing, which tells you the authors either did not check or would rather you did not. UniMate names all three sources, putting it ahead of most.

Second, whether any named source is a paid product. A commercial asset library in the training mix is the clearest signal that somebody should read terms before money gets involved.

Third, whether the licence on the weights is stated separately from the licence on the data. Projects that collapse those into one line are hiding a question rather than answering it. UniMate separates them cleanly.

That check takes two minutes. It has saved teams considerably more than two minutes.


What 13,006 Sequences Buys You

UniML3D holds 13,006 text paired motion sequences, roughly 20 hours of motion, with 3,584 unique prompts covering locomotion, combat, idle, mechanical articulation plus object manipulation.

Twenty hours is not large. Human motion datasets that drove the last few years of text to motion research run into the tens of hours for humans alone. UniMate covers seven broad skeleton families plus thousands of distinct rigs on the same budget.

That is partly deliberate. This architecture is meant to generalise across topology rather than memorise per rig, so broad coverage matters more than depth in any single family. The authors also apply online skeletal augmentation during training, stretching the effective variety well past the raw count.

What that means for your expectations

Still, it sets a ceiling. A model trained on twenty hours of motion across every body plan it supports will produce plausible locomotion plus recognisable combat beats. It will not produce the nuanced performance work an animator delivers. The paper never claims it will.

If your mental model is a text box that replaces your animation department, recalibrate. If your mental model is a fast first pass on a creature rig that would otherwise take a day to block out, that matches what the thing does.

There is a version of this tool that becomes properly useful tomorrow, which looks like a studio feeding it their own motion library. Nothing stops you fine tuning on proprietary capture, at which point the licence question gets simpler plus the output gets closer to your house style. Nobody has published that workflow yet.


What You Cannot Reproduce

Here is where the TODO list matters for a different reason.

Released: code, raw dataset, processing pipeline, interactive demo. Published separately: checkpoints.

Still outstanding: the exact training configuration, the data manifest used to produce those checkpoints, the evaluation scripts plus the demo prompts.

Read that carefully. You can run the model. Rebuilding it is off the table. Re-running the paper’s own comparisons is also off the table, which is why those three baselines from earlier deserve a second look. Without the evaluation scripts, nobody outside the team can independently check the reported improvements against AnyTop, How to Move Your Dragon or AnimaX.

A sharper detail sits buried in the README. Captions in the released dataset were re processed for release. The public dataset is therefore not identical to whatever trained the published checkpoints, while the manifest that would tell you the difference remains on the TODO list. Even with unlimited GPU budget you could not currently reconstruct that training run.

None of this is unusual for a graphics paper three weeks after arXiv. Research code lands in stages. SIGGRAPH deadlines are brutal. It does mean the benchmark table should be read as author reported until somebody reproduces it, which is the standard we applied to Hermes Agent’s open source claims plus to every model launch where a lab ran its own evals.


The Authors Said It First

Put on record that the limitations section of this paper is unusually honest.

Three constraints appear plainly. Supervision is bounded by 4D animation data, so the model only knows what existing animation libraries contain. Conditioning is restricted to skeletons plus text, meaning no image input, no reference video, no style exemplar. There is no native channel for fine grained spatial motion control.

That last one is the practical ceiling. Text gets you a behaviour. It does not get you a precise spatial instruction, while precise spatial instruction is most of what animation direction consists of. You cannot tell it to put the left foot on that specific rock.

A research team that writes its own ceiling into the paper has done its job. Every failure here is downstream, inside a transmission chain that keeps the dragon clip then drops everything around it.


What You Should Actually Do

If you want to try it this week, the path is short. Clone the GitHub repo for the code. Pull unimate_uniml3d_f60_v2 from Hugging Face rather than waiting for the GitHub checkpoint release. Start with the interactive demo to see whether output quality clears your bar before you commit to a 26 GB download. Python 3.10 with setuptools pinned below 81 is the stated environment.

If you are evaluating it for commercial work, do the licence read before the technical evaluation. Pull up the Truebones ZOO terms first, because that is the restrictive one, then decide whether your organisation is comfortable with weights distilled from a paid motion library.

If you are a solo creator or a small team, the realistic use is blocking plus iteration rather than final animation. Generate a first pass on an unusual rig, then hand it to a human to clean up. That workflow saves real hours while sidestepping the quality ceiling the authors described.

If you are running an agent stack and wondering where this fits, it does not yet. There is no MCP server, no plugin, no skill. Somebody will wrap it, the same way everything in our OpenClaw deep dive eventually got wrapped. Nobody has.

And if you only came for the clip, that is fine too. Just do not budget a sprint around a tool whose evaluation scripts have not shipped.

One last practical note. Grab the smaller Mixamo humanoid checkpoint first if your characters are human shaped, since 47.8M parameters downloads faster plus tells you quickly whether the output quality suits your project at all.


The Part Worth Keeping

UniMate is good work. Princeton graphics faculty put their names on it. The architecture solves a problem that had resisted a clean solution. Code is MIT, weights are public, plus the team documented its own limits without being asked.

Every problem in this article exists one layer above the research.

Start with a README listing checkpoints as TODO while those same checkpoints sit downloadable on another platform. Then there is a benchmark table nobody outside the team can re-run. Underneath all of it sits a motion library carrying a commercial licence, stated clearly by the authors then dropped by everyone who passed the demo along.

Eighty seven downloads against a viral clip is the figure to sit with. That gap between what gets shared plus what gets run keeps widening, while the stuff falling into it is almost always the stuff that decides whether you can use the thing at all.


Charts and Blocks

Release timeline

Timeline
What shipped when
1 August 2026
Interactive 3D demo goes live on the project page
30 August 2026
Raw UniML3D dataset plus the data processing pipeline released
4 September 2026
Paper posted to arXiv as 2609.05415v1, accepted to SIGGRAPH Asia 2026
6 September 2026
Training plus inference code released under MIT
Still outstanding
Exact training configuration, data manifest, evaluation scripts plus demo prompts. Checkpoints are listed here too, yet they are already on Hugging Face

Dates taken from the project README release notes plus the arXiv submission record, read 1 October 2026.

What the licences actually say

Licence stack
MIT on top, three different sets of terms underneath
Layer
Terms
What that means for shipping
Code
MIT
Clean. Use it, modify it, ship it.
Checkpoints
MIT
Clean on their face. The open question is what they learned from.
Mixamo motions
Adobe ToU
Free with an account. Redistributing the assets is a separate matter from using them.
Objaverse-XL objects
Per object
Thousands of separate licences. Some permissive, some non commercial. Unmappable from outside.
Truebones ZOO
Commercial
A paid motion capture product. Read this one before a studio commits.

Licence attribution quoted from the project README plus the Hugging Face model card. Whether trained weights inherit source data terms is unsettled law rather than a resolved question.