Two-dimensional gas chromatography coupled with time-of-flight mass spectrometry (GC×GC–ToF-MS) captures rich chemical information from complex mixtures, but produces high-dimensional measurements whose scale and structure make automated analysis challenging.
Building on the team's previous work on sex classification and identity verification from raw human-scent measurements, this project will investigate the feasibility of employing large language models for general-purpose GC×GC–ToF-MS analysis. In contrast to conventional approaches that require expert-driven compound identification and task-specific preprocessing, the proposed framework will learn directly from raw or minimally processed chromatographic data. Each GC×GC–ToF-MS measurement will be transformed into a compact sequence of chromatographic and spectral tokens preserving the two retention-time dimensions, mass-to-charge information, and detector intensity.
A scalable encoder will connect these tokens to an instruction-tuned language model, enabling the joint interpretation of measurements, sample metadata, analytical tasks, and reference spectral knowledge. The model will be trained using self-supervised, contrastive, and task-specific objectives and evaluated on sex classification, identity verification, sample retrieval, measurement comparison, and the generation of evidence-grounded analytical reports. EuroHPC resources are essential for distributed training and systematic experimentation on multi-terabyte datasets. The team's existing identity-verification dataset alone comprises 2,528 raw samples from 252 individuals and approximately 7.5 TB of data.
The requested resources will enable scaling studies across model sizes, chromatographic representations, tokenization strategies, spatial-alignment methods, and numerical precision. Particular attention will be given to uncertainty calibration, reproducibility, computational efficiency, and prevention of unsupported model outputs.The project will deliver a reusable software pipeline, pretrained model components, evaluation benchmarks, and practical recommendations for applying scientific foundation models to chromatography–mass spectrometry. By replacing isolated task-specific models with a shared representation and natural-language analytical interface, the project aims to reduce expert-driven processing and accelerate GC×GC–ToF-MS research across forensic, environmental, biomedical, and industrial applications.
Principal Investigator, Company and Country
Radim Spetlik, Czech Technical University in Prague, Czechia