Efficient and scalable insertion of large genetic payloads remains a major bottleneck in mammalian genome engineering. Current approaches, including viral vectors, DNA transposons, integrases and recombinases, provide powerful capabilities but remain limited by delivery complexity, episomal background, payload constraints, insertion profiles or cellular-context dependence.
This project will develop a workflow to discover and engineer RNA-driven retroelement-derived gene writers. Building on their preliminary LINE-1 work, where genome mining, protein language models and reinforcement learning enabled the discovery and design of hyperactive ORF2p variants, the researchers will scale this strategy towards a broader non-LTR retrotransposon space, including LINE-1, CR1, BovB/RTE and Jockey.
The project will combine own curated retroelement datasets, ProGen2 fine-tuning, alternative biological language models, rDPO/GRPO-based reinforcement learning, large-scale synthetic sequence generation, structural prediction and multi-objective in silico ranking. The requested EuroHPC allocation will enable the GPU-intensive stages required to train, align, generate and prioritise millions of long multi-domain retroelement proteins. Expected outputs include family-specific generative models, synthetic retroelement protein libraries, reusable computational workflows and a ranked shortlist of candidates for experimental validation as RNA-based gene-writing tools.
Principal Investigator, Company and Country
Gerard Mingarro Broch, Universitat Pompeu Fabra, Spain