This project proposes to train and validate a novel language model architecture based on energy-based neural networks that provides native explainability — a critical requirement under the EU AI Act for high-risk AI systems. Unlike standard transformer architectures, our approach produces built-in confidence scores and adaptive compute allocation as natural byproducts of inference, without requiring post-hoc approximation methods.
This project will:
(1) establish scaling laws for this architecture class from 59M to 7B parameters,
(2) benchmark against equivalent transformer baselines at each scale, and
(3) validate the architecture for enterprise deployment in regulated sectors.
Preliminary results on 13M-parameter models demonstrate competitive performance with standard transformers while providing native confidence metrics at zero additional inference cost. EuroHPC JU resources will enable scaling these results to 1B-7B parameters.
Principal Investigator, Company and Country
Marius Dima, Qriton Technologies, Romania