Research

I am a researcher and engineer specializing in speech technology and deep learning, with more than 15 years of experience developing speech systems from research to production. Throughout my career, I have designed novel deep learning models, developed production software, built machine learning infrastructure, and led multidisciplinary research and engineering teams. I enjoy working at the intersection of research and engineering, transforming new ideas into robust technologies used in real-world products.

My research focuses on learning meaningful representations from speech and developing end-to-end deep learning approaches for speech understanding. Over the years, I have worked on automatic speech recognition, speech emotion recognition, non-verbal vocalisation detection, voice activity detection, and, more recently, self-supervised learning and acoustic-to-articulatory inversion. My work has resulted in more than fifteen peer-reviewed publications, over 1,400 citations, and two journal best paper awards.

Since 2017, I have been with Speech Graphics, where I currently serve as Principal Research Scientist. My work combines research, software engineering, and technical leadership to develop production-ready speech technologies for audio-driven facial animation, a technology widely adopted across the video game industry. Beyond designing deep learning models, I contribute to production C++ software, machine learning infrastructure, and technical strategy, helping bridge the gap between research, engineering, and product development.

I received my PhD in Electrical Engineering from École Polytechnique Fédérale de Lausanne (EPFL) in 2016. My doctoral research was carried out at the Idiap Research Institute under the supervision of Ronan Collobert, Mathew Magimai-Doss, and Hervé Bourlard, where I explored end-to-end deep learning approaches for speech recognition.

Awards

2023 Eurasip Best Paper Award for Speech Communication Journal

“End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition”, Palaz D., Magimai-Doss M., and Collobert R., Volume 108, April, 2019

award-specom-23

ISCA Award for Best Paper Published in Speech Communication (2018-2022)

“End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition”, Palaz D., Magimai-Doss M., and Collobert R., Volume 108, April, 2019

award-is23

Publications

E. Stanley, E. DeMattos, A. Klementiev, P. Ozimek, G. Clarke, M. Berger, and D. Palaz, Emotion Label Encoding Using Word Embeddings for Speech Emotion Recognition. Proceedings of Interspeech 2023, pp. 2418-2422. [pdf]

J. Parry, E. DeMattos, A. Klementiev, A. Ind, D. Morse-Kopp, G. Clarke, and D. Palaz. Speech Emotion Recognition in the Wild using Multi-task and Adversarial Learning. Proceedings of Interspeech 2022, pp. 1158-1162. [pdf]

S. Condron, G. Clarke, A. Klementiev, D. Morse-Kopp, J. Parry, D. Palaz, Non-Verbal Vocalisation and Laughter Detection Using Sequence-to-Sequence Models and Multi-Label Training, Proceedings of Interspeech 2021, pp. 2506-2510. [pdf]

J. Parry, D. Palaz, G. Clarke, P. Lecomte, R. Mead, M. Berger, G. Hofer, Analysis of Deep Learning Architectures for Cross-Corpus Speech Emotion Recognition, Proceedings of Interspeech 2019, pp. 1656-1660. [pdf]

D. Palaz, M. Magimai-Doss, R. Collobert, End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition, Speech Communication 108, pp. 15-32, 2019. [pdf]

D. Palaz, Towards end-to-end speech recognition, Ph.D. Thesis, EPFL,  2016. [pdf]

D. Palaz, G. Synnaeve, R. Collobert, Jointly Learning to Locate and Classify Words Using Convolutional Network, Proceedings of Interspeech 2016, pp. 2741-2745. [pdf]

D. Palaz, M. Magimai-Doss, R. Collobert, Convolutional neural networks-based continuous speech recognition using raw speech signal, Proceedings of ICASSP 2015. [pdf]

D. Palaz, M. Magimai-Doss, R. Collobert, Analysis of CNN-based Speech Recognition System using Raw Speech as Input, Proceedings of Interspeech 2015, pp. 11-15. [pdf]

D. Palaz, M. Magimai-Doss, R. Collobert, Learning linearly separable features for speech recognition using convolutional neural networks, ICLR 2015 workshop. [pdf]

D. Palaz, M. Magimai-Doss, R. Collobert, Joint phoneme segmentation inference and classification using CRFs, Proceedings of IEEE Global Conference on Signal and Information Processing (GlobalSIP) 2014. [pdf]

D. Palaz, R. Collobert, M. Magimai-Doss, End-to-end phoneme sequence recognition using convolutional neural networks, NIPS 2013 Deep Learning workshop. [pdf]

D. Palaz, R. Collobert, M. Magimai Doss, Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks, Proceedings of Interspeech 2013, pp. 1766-1770. [pdf]

D. Palaz, I. Tošić, P. Frossard, Sparse stereo image coding with learned dictionaries, IEEE International Conference on Image Processing 2011, pp. 133-136.

M. Borgeaud, D. Palaz, P. Deleglise, Monitoring of Land Cover Charge Using SAR and Optical Data from the ESA Rolling Archives, Proceedings of ESA Living Planet Symposium, 2010.