Skip to content
Automation Forge

Plugins / FaceForge

5 plugins in this set.

Runs on local, on cpu · self-hosted, or a rented gpu · self-hosted, alongside audio2face · compute · rented by the hour.

What it can talk to →

The face that goes with the voice.

Spoken audio becomes a performing face — solved locally on CPU for nothing or on NVIDIA Audio2Face-3D, with emotion you set or emotion read out of the delivery, retargeted into the vocabulary your rig actually reads, and baked ready to play.

What it does

01

Solve a performing face straight from a sound wave — locally, on CPU, no account

02

Or solve on NVIDIA Audio2Face-3D, on your own card or a GPU rented by the hour

03

Emotion as an input, not only a reading: play the line angrier than it was recorded

04

Read the emotion out of the delivery instead, frame by frame, and drive the face with it

05

Retarget curves between face vocabularies, so the solve fits the rig you have

06

Check coverage before you bake: know the character can play every curve

07

Bake to a sequence and a montage, wired for the dialogue that plays it

Seen, not described

PlaceholderVideo: FaceForge, in the editor, start to finishAssets A2 and A3 — one video per set, plus stills or GIFs per headline capability.

What it is

The solve takes the audio a voice pass just produced and returns control curves. The engine’s own solver runs locally, on CPU, with no account and no per-use cost, and is fast enough to be interactive.

The second solver is NVIDIA’s Audio2Face-3D, in a container the plugin starts and talks to over HTTP — nothing NVIDIA is linked into the editor. Its face models are ungated and ship inside the image, so lip-sync still needs no account; it runs on your own card, or on a GPU rented by the hour when you have none.

What that second solver adds is emotion. Ten named emotions that mix, taken as an *input* — so a line can be angrier than it was recorded, because the player just failed the quest and the recording could not have known. Add Audio2Emotion and it reads the emotion out of the delivery instead, frame by frame; give it both and the performance drives while the emotion you named leans it further.

The part that is easy to get wrong is the vocabulary. A solver emits curves in one naming scheme and a rig reads another — and a curve nothing recognises animates nothing and warns nobody. FaceForge retargets between vocabularies and reports coverage, because a pipeline that says success while the face never moves is worse than one that fails.

FaceForge knows nothing about any particular solver — providers register themselves — and nothing about any gameplay framework.

Measured on our own project

1 : 4placeholder

solve time to audio length

$0placeholder

per solve — local, on CPU

What is in the set

5 plugins, and what each one is for.

Install what you need. A toolset can be deleted and the capability beside it behaves identically; a provider can be deleted and the core still loads, with one fewer option.

FaceForge

Core

Solve, retarget between face vocabularies, and bake to a sequence and a montage.

Open source · Fab

GitHub

FaceForge MetaHuman

Adapter

Registers the engine's own audio-driven solver. A sound wave in, 81 face-board control curves out, with no MetaHuman asset required anywhere. Remove it and FaceForge still loads, with one fewer solver.

Open source · Fab

GitHub

FaceForge ACE

Provider

Registers NVIDIA's Audio2Face-3D as a second solver, in a container this plugin starts and talks to over HTTP — no CUDA, no TensorRT and no SDK inside the editor. Its models are ungated and ship inside the image, so lip-sync needs no account at all.

Free on Fab

FaceForge ACE Toolset

Agent toolset

The runner, its container and a rented GPU as agent tools. Solving is deliberately absent — a bank pointed at Audio2Face goes through FaceForge like any other.

Open source · Fab

GitHub

FaceForge Toolset

Agent toolset

Author face banks, check whether a character can actually play the curves, solve, and bake.

Open source · Fab

GitHub

Chosen per asset, not per project

What it can talk to.

No stage hard-depends on a particular model, and every stage records what produced its output. A better model arrives as a config change, not a rewrite — today’s is the worst one this will ever run on.

Unreal Engine audio-to-face

Local, on CPU

The engine's own solver, fast enough to be interactive — a second of solve per four seconds of audio.

You need

Nothing. No account, no per-use cost, and no MetaHuman licence for the solve itself.

NVIDIA Audio2Face-3D

Self-hosted, or a rented GPU

Open-sourced models that need no engine to solve, so the same runner serves an editor, a batch or a browser. Emotion is an input here rather than only a reading: a line can be angrier than it was recorded, because the player just failed the quest.

You need

An NVIDIA card and Docker. No account at all — the face models are ungated and ship inside the container image.

NVIDIA Audio2Emotion

Self-hosted, alongside Audio2Face

Reads the emotion out of the delivery frame by frame and drives the face with it, so a line performed angry looks angry with nobody labelling it. Measured on real dialogue: it moves the brows and the mouth corners while leaving the lip-sync alone.

You need

A free Hugging Face token and one licence click. Optional: without it every face still solves, with the emotion you specify.

RunPod

Compute · rented by the hour

For a machine with no NVIDIA card, or a whole-game batch. A rented pod pulls the same container image a desktop builds, so there is one artefact to trust. A face solve is a fraction of a second, so renting is never a way to make one line faster.

You need

A RunPod key with write access. Shared with every set that rents hardware, so it is entered once.

Asked first

What studios ask about FaceForge.

Will it animate the character I already have?

If the head deforms, yes. A face is the one thing in this family you cannot drop onto a rigid head: it needs a MetaHuman, or a head an artist rigged. The solve itself is character-agnostic — only the control names are MetaHuman-shaped.

What if my rig reads different curve names?

That is the problem FaceForge exists to solve. Solvers emit one vocabulary, rigs read another, and a curve nothing recognises animates nothing and warns nobody. FaceForge retargets between vocabularies and reports coverage, so silence cannot pass as success.

Do I need a MetaHuman licence?

Not for the solve. The engine solver takes audio in and emits face-board control curves with no MetaHuman asset involved anywhere. The MetaHuman adapter exists for when your character is one.

Can I reuse a face across takes?

A gesture, yes. Lipsync, no — a face solved for one reading puts the wrong words in the mouth of another. Re-solve for the take you keep; at a second of solve per four of audio, that costs seconds.