Skip to main content
The European High Performance Computing Joint Undertaking (EuroHPC JU)

Diffusion based human avatar video generation

500 000 Awarded Resources (in node hours)
MareNostrum5 ACC System Partition
February 2026 - 12 months Allocation Period

This project aims to develop a unified, diffusion-based video generation framework capable of synthesising high-fidelity human avatars with generalisable fine motor control and physical interaction capabilities. While current generative models excel at static portraiture, they fail to address the "dynamic interaction gap"—specifically, the ability for digital humans to handle objects, gesture naturally with hands, and maintain temporal coherence during complex motion.

This research targets three foundational computer vision challenges:

Fine-Grained Articulation: Solving the topological complexity of hand-object interaction (e.g., grasping tools, intricate signing) and self-occlusion to prevent visual artifacts.

Multi-Modal Synchronisation: Establishing a robust alignment between audio, text, and motion inputs to drive not just lip-sync, but full-body emotive expression.

Real-Time Efficiency: Distilling massive diffusion priors into lightweight student models suitable for live, low-latency applications.

By solving these fundamental problems, the project establishes a base layer technology with broad applicability across sectors. The resulting system will power next-generation use cases ranging from AI Sign Language Interpreters for accessibility and immersive online teaching assistants, to interactive virtual performers for the entertainment industry.

Principal Investigator, Company and Country

Miloš Lokajíček, Valka AI, Czechia