Skip to main content
The European High Performance Computing Joint Undertaking (EuroHPC JU)

LAIONsynthscale: a factory for scalable synthetic data generation to compose canonical datasets for training open multi-modal foundation models with strong reasoning and

4 560 000 Awarded Resources (in node hours)
MareNostrum5 ACC System Partition
March 2026 - 12 months Allocation Period

Recent advances in creating reasoning and agentic models demonstrate that dataset composition — generation, selection and mixing of training data — is as critical as model architecture and compute scale for achieving strong reasoning, generalisation and robustness of foundation models. Yet no standardised, openly validated canonical datasets exist that show how to systematically combine real and synthetic data across reasoning, agentic, and multi-modal pre- and post-training settings to serve as reproducible common ground for both research and industry. 

Building on our previous successful work

OpenThoughts [1] (open reasoning datasets with >1000 ablation experiments, producing models matching state-of-the-art on major benchmarks), 

OpenThoughts-Agent [2] (open end-to-end pipelines for agentic trajectory data composition, SFT+RL training and evaluation), 

Terminal-Bench [3] (reproducible evaluation for reasoning and agentic models), 

MixtureVitae [4] (permissive web-scale pre-training dataset with instruction data) and 

Self-Flow [5] (Self-supervised flow matching for scalable multi-modal synthesis) 

this project proposes to investigate and to create canonical datasets at unprecedented scale, validated through reference model training and standardised evaluation. 

In WP1 (Scalable Synthetic Data Generation, ~2.9M H100 GPU-h), the project researchers 

(1) scale reasoning data from 1.2M to ~100M traces (~1T tokens, "OpenThoughts XXL"), systematically comparing teacher models and studying data scale effects on downstream reasoning capabilities; 

(2) generate ~100M agentic trajectories (~1T tokens) combining SFT traces from teacher agents with RL tasks featuring verifiable rewards for models with tool-use and multi-step task completion; 

(3) produce ~47B tokens of multi-modal captions (~300M samples) providing high-quality language descriptions for images, audio, and video using vision-language (Qwen3-VL) and omni-modal (Qwen3-Omni) models.

 In WP2 (Dataset Validation via Reference Training, ~1.7M H100 GPU-h), the researchers validate the generated datasets by post-training reasoning, agentic, and multi-modal models at 1.7B–32B parameter scales using LlamaFactory, SkyRL, and vision-language/audio-language instruction tuning (Qwen3-VL, Qwen3-Omni), evaluating on established benchmarks including AIME, GPQA Diamond, LiveCodeBench, Terminal-Bench 2.0, SWE-Bench Verified, and multi-modal benchmarks. 

This creates an iteratively improvable, validated canonical dataset base that can standardise experimental and developmental workflows on open foundation models and datasets across research and industry. The research team includes members of LAION, FZJ (JSC), Black Forest Labs, Prior Labs, Ellamind, ELLIS Institute Tuebingen, Tuebingen AI Center, BSC, CINECA and collaborators from large EU projects OpenEuroLLM, ELLIOT and MINERVA, with demonstrated track of record and experience of handling foundation model and dataset research at large scales, using routinely > 256 nodes for research and development on various supercomputers. 

Data generation throughput is measured on MN5 ACC, confirming feasibility of the resource plan, and further parts are already tested and measured on various supercomputers (Leonardo, JUWELS Booster, JUPITER). All datasets, generation pipelines, trained models, and evaluation results will be openly released, directly contributing to European AI sovereignty, while providing validated, shareable, strong and further scalable data foundations for the broader open ML/AI collaborative research and development ecosystem.

Principal Investigator, Company and Country

Robin Rombach, Black Forest Labs, Germany