Real-Time Lip Sync
An Unreal Engine client that animates a MetaHuman's face in real time on an LLM-generated voice.
Unreal 5.5C++MetaHumanLiveLink

The problem
The “Meta-Serious Game” project already has a REST middleware that lets the player talk to NPCs through a large language model (LLM). But while the NPC answers, the web client only shows a still image: the text is believable, the scene is not.
The goal of my thesis: an Unreal Engine client where a MetaHuman actually speaks the answer, around one research question: what latency and quality budgets does the exchange need to feel both responsive and believable?
The pipeline
- The middleware returns the LLM answer and the synthesized audio.
- The client analyzes the audio with Rhubarb Lip Sync, which extracts a sequence of visemes (mouth shapes).
- Each viseme is converted into ARKit blendshape curves, interpolated to smooth the transitions.
- A custom
ILiveLinkSourcewritten in C++ pushes these curves into the MetaHuman’s facial rig (RigLogic), frame by frame. - Simple animations complete it: random eye blinks and an idle body animation (retargeted from Mixamo).
Evaluation
- Technical: every step of the pipeline is timed and automatically logged to a CSV file, over 57 runs.
- Perceptual: a two-round, video-based pilot study (15 responses in total) to judge how believable the animation is.
Results
- Around 2 seconds of end-to-end latency, split almost evenly between the network round trip to the middleware and the Rhubarb analysis.
- On the client side, audio analysis alone accounts for 99.7% of the time: that’s the bottleneck to tackle first.
- On quality, the mere presence of viseme-based lip sync and a few idle animations is enough to make the exchange believable, regardless of how finely it is tuned.


