The deployment of state-of-the-art automatic speech recognition on edge devices is fundamentally limited by the immense energy consumption of dense, continuous neural networks. This project builds upon the successful development of the Spiking Ternary GRU, a strictly MAC-free neuromorphic architecture that achieved single-digit Word Error Rates at a fraction of the energy cost of continuous models. The project aims to scale this discrete paradigm to more advanced sequence modelling frameworks and massive, highly diverse datasets. The primary objective is to adapt the highly parallelised xLSTM architecture into a fully discrete, ternary-spiking format. By designing novel deferred-scaling operators, the team will eliminate power-hungry floating-point multiply-accumulate operations from the xLSTM gating mechanisms, preserving its parallel efficiency while harvesting the extreme energy savings of event-driven computation. Furthermore, they will rigorously evaluate these advanced MAC-free networks on the massive Mozilla Common Voice English dataset combined with LibriSpeech. A critical focus of this continuation project is to introduce comprehensive noise robustness testing protocols, determining how effectively highly quantised, bio-inspired architectures can segregate speech from chaotic background interference compared to standard continuous models. This research aims to deliver a scalable, edge-ready blueprint for sustainable, robust artificial intelligence.
Principal Investigator, Company and Country
Themos Stafylakis, Athens University of Economics and Business, Greece