Skip to main content
The European High Performance Computing Joint Undertaking (EuroHPC JU)

EuroDistil-Compact: A Sovereign, Multilingual 3B Dense Model via Multi-Teacher Knowledge Distillation

80,000 Awarded Resources (in node hours)
MareNostrum5 ACC System Partition
August 2026 - February 2027 Allocation Period

EuroDistil-Compact trains a compact, sovereign, multilingual 3B dense language model for five major European languages (English, German, French, Italian, Spanish) by knowledge distillation rather than from scratch. Capable compact models exist, but none combine a European sovereign provenance, broad coverage of the major European languages, and a fully open, reproducible recipe.

The project distils a dense decoder-only student from strong open foundation models that share a single tokenizer with the student: NVIDIA’s Nemotron-3 family and the European sovereign Soofi line. Teacher top-k logits are pre-computed once over a ~4T-token corpus (processed in 1T-token chunks) and reused throughout, using our PyTorch-native Modalities framework [4]. Beyond standard logit distillation the team studies three questions: how teacher capability affects distillation quality; a teacher-interpolation curriculum that blends two teachers over a configurable window to smooth distribution shift; and a stage-free curriculum that removes the hard pretraining/SFT/RL boundaries by continuously interpolating the data mixture. All models, data accounting and code are released openly.

The team request 90,000 MareNostrum 5 (BSC) node-hours (≈360,000 GPU-hours) for a 6-month allocation.

Principal Investigator, Company and Country

Max Lübbering, Fraunhofer IAIS, Germany