FaceForge
Core
Solve, retarget between face vocabularies, and bake to a sequence and a montage.
Open source · Fab
Plugins / FaceForge
5 plugins in this set.
Runs on local, on cpu · self-hosted, or a rented gpu · self-hosted, alongside audio2face · compute · rented by the hour.
What it does
Solve a performing face straight from a sound wave — locally, on CPU, no account
Or solve on NVIDIA Audio2Face-3D, on your own card or a GPU rented by the hour
Emotion as an input, not only a reading: play the line angrier than it was recorded
Read the emotion out of the delivery instead, frame by frame, and drive the face with it
Retarget curves between face vocabularies, so the solve fits the rig you have
Check coverage before you bake: know the character can play every curve
Bake to a sequence and a montage, wired for the dialogue that plays it
Seen, not described
What it is
The solve takes the audio a voice pass just produced and returns control curves. The engine’s own solver runs locally, on CPU, with no account and no per-use cost, and is fast enough to be interactive.
The second solver is NVIDIA’s Audio2Face-3D, in a container the plugin starts and talks to over HTTP — nothing NVIDIA is linked into the editor. Its face models are ungated and ship inside the image, so lip-sync still needs no account; it runs on your own card, or on a GPU rented by the hour when you have none.
What that second solver adds is emotion. Ten named emotions that mix, taken as an *input* — so a line can be angrier than it was recorded, because the player just failed the quest and the recording could not have known. Add Audio2Emotion and it reads the emotion out of the delivery instead, frame by frame; give it both and the performance drives while the emotion you named leans it further.
The part that is easy to get wrong is the vocabulary. A solver emits curves in one naming scheme and a rig reads another — and a curve nothing recognises animates nothing and warns nobody. FaceForge retargets between vocabularies and reports coverage, because a pipeline that says success while the face never moves is worse than one that fails.
FaceForge knows nothing about any particular solver — providers register themselves — and nothing about any gameplay framework.
Measured on our own project
1 : 4placeholder
solve time to audio length
$0placeholder
per solve — local, on CPU
What is in the set
Install what you need. A toolset can be deleted and the capability beside it behaves identically; a provider can be deleted and the core still loads, with one fewer option.
Core
Solve, retarget between face vocabularies, and bake to a sequence and a montage.
Open source · Fab
Adapter
Registers the engine's own audio-driven solver. A sound wave in, 81 face-board control curves out, with no MetaHuman asset required anywhere. Remove it and FaceForge still loads, with one fewer solver.
Open source · Fab
Provider
Registers NVIDIA's Audio2Face-3D as a second solver, in a container this plugin starts and talks to over HTTP — no CUDA, no TensorRT and no SDK inside the editor. Its models are ungated and ship inside the image, so lip-sync needs no account at all.
Free on Fab
Agent toolset
The runner, its container and a rented GPU as agent tools. Solving is deliberately absent — a bank pointed at Audio2Face goes through FaceForge like any other.
Open source · Fab
Agent toolset
Author face banks, check whether a character can actually play the curves, solve, and bake.
Open source · Fab
Chosen per asset, not per project
No stage hard-depends on a particular model, and every stage records what produced its output. A better model arrives as a config change, not a rewrite — today’s is the worst one this will ever run on.
Local, on CPU
The engine's own solver, fast enough to be interactive — a second of solve per four seconds of audio.
You need
Nothing. No account, no per-use cost, and no MetaHuman licence for the solve itself.
Self-hosted, or a rented GPU
Open-sourced models that need no engine to solve, so the same runner serves an editor, a batch or a browser. Emotion is an input here rather than only a reading: a line can be angrier than it was recorded, because the player just failed the quest.
You need
An NVIDIA card and Docker. No account at all — the face models are ungated and ship inside the container image.
Self-hosted, alongside Audio2Face
Reads the emotion out of the delivery frame by frame and drives the face with it, so a line performed angry looks angry with nobody labelling it. Measured on real dialogue: it moves the brows and the mouth corners while leaving the lip-sync alone.
You need
A free Hugging Face token and one licence click. Optional: without it every face still solves, with the emotion you specify.
Compute · rented by the hour
For a machine with no NVIDIA card, or a whole-game batch. A rented pod pulls the same container image a desktop builds, so there is one artefact to trust. A face solve is a fraction of a second, so renting is never a way to make one line faster.
You need
A RunPod key with write access. Shared with every set that rents hardware, so it is entered once.
Asked first
If the head deforms, yes. A face is the one thing in this family you cannot drop onto a rigid head: it needs a MetaHuman, or a head an artist rigged. The solve itself is character-agnostic — only the control names are MetaHuman-shaped.
That is the problem FaceForge exists to solve. Solvers emit one vocabulary, rigs read another, and a curve nothing recognises animates nothing and warns nobody. FaceForge retargets between vocabularies and reports coverage, so silence cannot pass as success.
Not for the solve. The engine solver takes audio in and emits face-board control curves with no MetaHuman asset involved anywhere. The MetaHuman adapter exists for when your character is one.
A gesture, yes. Lipsync, no — a face solved for one reading puts the wrong words in the mouth of another. Re-solve for the take you keep; at a second of solve per four of audio, that costs seconds.
Next
Opens a message to bojan@blackcode.ch. One email when something ships. Not a newsletter.