Skip to content

Jonathan Le Roux

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

115

Venues

10

Active years

2002–2025

Best venue rank

A*

Where they publish

Papers

115 indexed papers, newest first.

YearVenueTitleAuthors
2025ASRURobot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM.Chiori Hori, Yoshiki Masuyama, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux
2025ICASSPNo Class Left Behind: A Closer Look at Class Balancing for Audio Tagging.Janek Ebbers, Franois G. Germain, Kevin Wilkinghoff, Gordon Wichern, Jonathan Le Roux
2025ICASSPO-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization.Elio Gruttadauria, Mathieu Fontaine, Jonathan Le Roux, Slim Essid
2025ICASSPInteractive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos.Chiori Hori, Motonari Kambara, Komei Sugiura, Kei Ota, Sameer Khurana, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux
2025ICASSPRetrieval-Augmented Neural Field for HRTF Upsampling and Personalization.Yoshiki Masuyama, Gordon Wichern, Franois G. Germain, Christopher Ick, Jonathan Le Roux
2025ICASSPLeveraging Audio-Only Data for Text-Queried Target Sound Extraction.Kohei Saijo, Janek Ebbers, Franois G. Germain, Sameer Khurana, Gordon Wichern, Jonathan Le Roux
2025ICASSPTask-Aware Unified Source Separation.Kohei Saijo, Janek Ebbers, Franois G. Germain, Gordon Wichern, Jonathan Le Roux
2025ICASSPKeeping the Balance: Anomaly Score Calculation for Domain Generalization.Kevin Wilkinghoff, Haici Yang, Janek Ebbers, Franois G. Germain, Gordon Wichern, Jonathan Le Roux
2025InterspeechHASRD: Hierarchical Acoustic and Semantic Representation Disentanglement.Amir Hussein, Sameer Khurana, Gordon Wichern, Franois G. Germain, Jonathan Le Roux
2025InterspeechDirection-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses.Christopher Ick, Gordon Wichern, Yoshiki Masuyama, Franois G. Germain, Jonathan Le Roux
2025InterspeechFactorized RVQ-GAN For Disentangled Speech Tokenization.Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, Franois G. Germain, Gordon Wichern, Jonathan Le Roux
2025InterspeechInvestigating continuous autoregressive generative speech enhancement.Haici Yang, Gordon Wichern, Ryo Aihara, Yoshiki Masuyama, Sameer Khurana, Franois G. Germain, Jonathan Le Roux
2024CVPRRILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation.Zeyuan Yang, Jiageng Lin, Peihao Chen, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan
2024ICASSPSpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis.Teysir Baoueb, Haocheng Liu, Mathieu Fontaine, Jonathan Le Roux, Gal Richard
2024ICASSPGeneration or Replication: Auscultating Audio Latent Diffusion Models.Dimitrios Bralios, Gordon Wichern, Franois G. Germain, Zexu Pan, Sameer Khurana, Chiori Hori, Jonathan Le Roux
2024ICASSPWI-FI based Indoor Monitoring Enhanced by Multimodal Fusion.Chiori Hori, Pu Wang, Mahbub Rahman, Cristian J. Vaca-Rubio, Sameer Khurana, Anoop Cherian, Jonathan Le Roux
2024ICASSPWhy Does Music Source Separation Benefit from Cacophony?Chang-Bin Jeon, Gordon Wichern, Franois G. Germain, Jonathan Le Roux
2024ICASSPGLA-GRAD: A Griffin-Lim Extended Waveform Generation Diffusion Model.Haocheng Liu, Teysir Baoueb, Mathieu Fontaine, Jonathan Le Roux, Gal Richard
2024ICASSPNIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization.Yoshiki Masuyama, Gordon Wichern, Franois G. Germain, Zexu Pan, Sameer Khurana, Chiori Hori, Jonathan Le Roux
2024ICASSPNeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention Detection.Zexu Pan, Gordon Wichern, Franois G. Germain, Sameer Khurana, Jonathan Le Roux
2024ICASSPLate Audio-Visual Fusion for in-the-Wild Speaker Diarization.Zexu Pan, Gordon Wichern, Franois G. Germain, Aswin Shanmugam Subramanian, Jonathan Le Roux
2024ICASSPImproving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation.Shih-Lun Wu, Xuankai Chang, Gordon Wichern, Jee-Weon Jung, Franois G. Germain, Jonathan Le Roux, Shinji Watanabe
2024InterspeechSpeech dereverberation constrained on room impulse response characteristics.Louis Bahrman, Mathieu Fontaine, Jonathan Le Roux, Gal Richard
2024InterspeechSound Event Bounding Boxes.Janek Ebbers, Franois G. Germain, Gordon Wichern, Jonathan Le Roux
2024InterspeechZeroST: Zero-Shot Speech Translation.Sameer Khurana, Chiori Hori, Antoine Laurent, Gordon Wichern, Jonathan Le Roux
2024InterspeechPARIS: Pseudo-AutoRegressIve Siamese Training for Online Speech Separation.Zexu Pan, Gordon Wichern, Franois G. Germain, Kohei Saijo, Jonathan Le Roux
2024InterspeechEnhanced Reverberation as Supervision for Unsupervised Speech Separation.Kohei Saijo, Gordon Wichern, Franois G. Germain, Zexu Pan, Jonathan Le Roux
2024IROSDisentangled Acoustic Fields For Multimodal Physical Scene Understanding.Jie Yin, Andrew Luo, Yilun Du, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan
2023ASRUScenario-Aware Audio-Visual TF-Gridnet for Target Speech Extraction.Zexu Pan, Gordon Wichern, Yoshiki Masuyama, Franois G. Germain, Sameer Khurana, Chiori Hori, Jonathan Le Roux
2023ICASSPReverberation as Supervision For Speech Separation.Rohith Aralikatti, Christoph Bddeker, Gordon Wichern, Aswin Shanmugam Subramanian, Jonathan Le Roux
2023ICASSPLatent Iterative Refinement for Modular Source Separation.Dimitrios Bralios, Efthymios Tzinis, Gordon Wichern, Paris Smaragdis, Jonathan Le Roux
2023ICASSPPaᗧ-HuBERT: Self-Supervised Music Source Separation Via Primitive Auditory Clustering And Hidden-Unit Bert.Ke Chen, Gordon Wichern, Franois G. Germain, Jonathan Le Roux
2023ICASSPHyperbolic Audio Source Separation.Darius Petermann, Gordon Wichern, Aswin Shanmugam Subramanian, Jonathan Le Roux
2023ICASSPOptimal Condition Training for Target Source Separation.Efthymios Tzinis, Gordon Wichern, Paris Smaragdis, Jonathan Le Roux
2023ICASSPCold Diffusion for Speech Enhancement.Hao Yen, Franois G. Germain, Gordon Wichern, Jonathan Le Roux
2023InterspeechStyle-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos.Chiori Hori, Puyuan Peng, David Harwath, Xinyu Liu, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux
2022AAAI(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering.Anoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le Roux
2022ICASSPExtended Graph Temporal Classification for Multi-Speaker End-to-End ASR.Xuankai Chang, Niko Moritz, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux
2022ICASSPAdvancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy.Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, Takaaki Hori
2022ICASSPSequence Transduction with Graph-Based Supervision.Niko Moritz, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux
2022ICASSPThe Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks.Darius Petermann, Gordon Wichern, Zhong-Qiu Wang, Jonathan Le Roux
2022ICASSPAudio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning.Ankit P. Shah, Shijie Geng, Peng Gao, Anoop Cherian, Takaaki Hori, Tim K. Marks, Jonathan Le Roux, Chiori Hori
2022ICASSPLocate This, Not that: Class-Conditioned Sound Event DOA Estimation.Olga Slizovskaia, Gordon Wichern, Zhong-Qiu Wang, Jonathan Le Roux
2022InterspeechLow-Latency Online Streaming VideoQA Using Audio-Visual Transformers.Chiori Hori, Takaaki Hori, Jonathan Le Roux
2022InterspeechHeterogeneous Target Speech Separation.Efthymios Tzinis, Gordon Wichern, Aswin Shanmugam Subramanian, Paris Smaragdis, Jonathan Le Roux
2021AAAIDynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers.Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, Anoop Cherian
2021ICASSPTranscription Is All You Need: Learning To Separate Musical Mixtures With Score As Supervision.Yun-Ning Hung, Gordon Wichern, Jonathan Le Roux
2021ICASSPUnsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training.Sameer Khurana, Niko Moritz, Takaaki Hori, Jonathan Le Roux
2021ICASSPCapturing Multi-Resolution Context by Dilated Self-Attention.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2021ICASSPSemi-Supervised Speech Recognition Via Graph-Based Temporal Classification.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2021ICCVVisual Scene Graphs for Audio Source Separation.Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian
2021InterspeechMomentum Pseudo-Labeling for Semi-Supervised Speech Recognition.Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, Takaaki Hori
2021InterspeechOptimizing Latency for Online Video Captioning Using Audio-Visual Transformers.Chiori Hori, Takaaki Hori, Jonathan Le Roux
2021InterspeechAdvanced Long-Context End-to-End Speech Recognition Using Context-Expanded Transformers.Takaaki Hori, Niko Moritz, Chiori Hori, Jonathan Le Roux
2021InterspeechDual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2020ICASSPEnd-To-End Multi-Speaker Speech Recognition With Transformer.Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux, Shinji Watanabe
2020ICASSPWHAMR!: Noisy and Reverberant Single-Channel Speech Separation.Matthew Maciejewski, Gordon Wichern, Emmett McQuinn, Jonathan Le Roux
2020ICASSPStreaming Automatic Speech Recognition with the Transformer Model.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2020ICASSPLearning to Separate Sounds from Weakly Labeled Scenes.Fatemeh Pishdadian, Gordon Wichern, Jonathan Le Roux
2020ICASSPUnsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR.Leda Sari, Niko Moritz, Takaaki Hori, Jonathan Le Roux
2020InterspeechTransformer-Based Long-Context End-to-End Speech Recognition.Takaaki Hori, Niko Moritz, Chiori Hori, Jonathan Le Roux
2020InterspeechDetecting Audio Attacks on ASR Systems with Dropout Uncertainty.Tejas Jayashankar, Jonathan Le Roux, Pierre Moulin
2020InterspeechAll-in-One Transformer: Unifying Speech Recognition, Audio Tagging, and Event Detection.Niko Moritz, Gordon Wichern, Takaaki Hori, Jonathan Le Roux
2019ASRUMIMO-Speech: End-to-End Multi-Channel Multi-Speaker Speech Recognition.Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux, Shinji Watanabe
2019ASRUStreaming End-to-End Speech Recognition with Joint CTC-Attention Based Models.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2019ICASSPTeacher-student Deep Clustering for Low-delay Single Channel Speech Separation.Ryo Aihara, Toshiyuki Hanazawa, Yohei Okato, Gordon Wichern, Jonathan Le Roux
2019ICASSPCycle-consistency Training for End-to-end Speech Recognition.Takaaki Hori, Ramn Fernandez Astudillo, Tomoki Hayashi, Yu Zhang, Shinji Watanabe, Jonathan Le Roux
2019ICASSPTriggered Attention for End-to-end Speech Recognition.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2019ICASSPSDR - Half-baked or Well Done?Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, John R. Hershey
2019ICASSPThe Phasebook: Building Complex Masks via Discrete Representations for Source Separation.Jonathan Le Roux, Gordon Wichern, Shinji Watanabe, Andy M. Sarroff, John R. Hershey
2019ICASSPBootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo Mixtures.Prem Seetharaman, Gordon Wichern, Jonathan Le Roux, Bryan Pardo
2019ICASSPClass-conditional Embeddings for Music Source Separation.Prem Seetharaman, Gordon Wichern, Shrikant Venkataramani, Jonathan Le Roux
2019InterspeechUnidirectional Neural Network Architectures for End-to-End Automatic Speech Recognition.Niko Moritz, Takaaki Hori, Jonathan Le Roux
2019InterspeechVectorized Beam Search for CTC-Attention-Based Speech Recognition.Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Niko Moritz, Jonathan Le Roux
2019InterspeechEnd-to-End Multilingual Multi-Speaker Speech Recognition.Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2019InterspeechWHAM!: Extending Speech Separation to Noisy Environments.Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, Jonathan Le Roux
2018ACLA Purely End-to-End System for Multi-speaker Speech Recognition.Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2018ICASSPAn End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech.Hiroshi Seki, Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey
2018ICASSPEnd-to-End Multi-Speaker Speech Recognition.Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, John R. Hershey
2018ICASSPMulti-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation.Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey
2018ICASSPAlternative Objective Functions for Deep Clustering.Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey
2018InterspeechEnd-to-End Speech Separation with Unfolded Iterative Phase Reconstruction.Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, John R. Hershey
2017ICASSPBLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection.Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda
2017ICASSPDeep clustering and conventional networks for music separation: Stronger together.Yi Luo, Zhuo Chen, John R. Hershey, Jonathan Le Roux, Nima Mesgarani
2017ICASSPStudent-teacher network learning with enhanced features.Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey
2017InterspeechCoupled Initialization of Multi-Channel Non-Negative Matrix Factorization Based on Spatial and Spectral Information.Yuuki Tachioka, Tomohiro Narita, Iori Miura, Takanobu Uramoto, Natsuki Monta, Shingo Uenohara, Ken'ichi Furuya, Shinji Watanabe, Jonathan Le Roux
2016ICASSPDeep clustering: Discriminative embeddings for segmentation and separation.John R. Hershey, Zhuo Chen, Jonathan Le Roux, Shinji Watanabe
2016ICASSPDeep unfolding for multichannel source separation.Scott Wisdom, John R. Hershey, Jonathan Le Roux, Shinji Watanabe
2016InterspeechImproved MVDR Beamforming Using Single-Channel Mask Prediction Networks.Hakan Erdogan, John R. Hershey, Shinji Watanabe, Michael I. Mandel, Jonathan Le Roux
2016InterspeechSingle-Channel Multi-Speaker Separation Using Deep Clustering.Yusuf Ziya Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, John R. Hershey
2015ASRUThe MERL/SRI system for the 3RD CHiME challenge using beamforming, robust feature extraction, and advanced speech recognition.Takaaki Hori, Zhuo Chen, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe
2015ICASSPPhase-sensitive and recognition-boosted speech separation using deep recurrent neural networks.Hakan Erdogan, John R. Hershey, Shinji Watanabe, Jonathan Le Roux
2015ICASSPDeep NMF for speech separation.Jonathan Le Roux, John R. Hershey, Felix Weninger
2015ICASSPMicbots: Collecting large realistic datasets for speech and audio research using mobile robots.Jonathan Le Roux, Emmanuel Vincent, John R. Hershey, Daniel P. W. Ellis
2014ICASSPNon-negative source-filter dynamical system for speech enhancement.Umut Simsekli, Jonathan Le Roux, John R. Hershey
2014ICASSPBlack box optimization for automatic speech recognition.Shinji Watanabe, Jonathan Le Roux
2014InterspeechSequential maximum mutual information linear discriminant analysis for speech recognition.Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2014InterspeechDiscriminative NMF and its application to single-channel source separation.Felix Weninger, Jonathan Le Roux, John R. Hershey, Shinji Watanabe
2013ASRUA generalized discriminative training framework for system combination.Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2013ASRUThe second 'CHiME' speech separation and recognition challenge: An overview of challenge systems and outcomes.Emmanuel Vincent, Jon Barker, Shinji Watanabe, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
2013ICASSPNon-negative dynamical system with application to speech and audio.Cdric Fvotte, Jonathan Le Roux, John R. Hershey
2013ICASSPSource localization in reverberant environments using sparse optimization.Jonathan Le Roux, Petros T. Boufounos, Kang Kang, John R. Hershey
2013ICASSPThe second 'chime' speech separation and recognition challenge: Datasets, tasks and baselines.Emmanuel Vincent, Jon Barker, Shinji Watanabe, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
2013IJCNLPStatistical Dialogue Management using Intention Dependency Graph.Koichiro Yoshino, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2012ICASSPIndirect model-based speech enhancement.Jonathan Le Roux, John R. Hershey
2011ICASSPInfinite-state spectrum model for music signal analysis.Masahiro Nakano, Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama
2010InterspeechA statistical model of speech F0 contours.Hirokazu Kameoka, Jonathan Le Roux, Yasunori Ohishi
2008ICASSPModulation analysis of speech through orthogonal FIR filterbank optimization.Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama, Alain de Cheveign
2008InterspeechComputational auditory induction by missing-data non-negative matrix factorization.Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveign, Shigeki Sagayama
2008InterspeechExplicit consistency constraints for STFT spectrograms and their application to phase reconstruction.Jonathan Le Roux, Nobutaka Ono, Shigeki Sagayama
2007ICASSPMEG Signal Denoising Based on Time-Shift PCA.Alain de Cheveign, Jonathan Le Roux, Jonathan Z. Simon
2007ICASSPHarmonic-Temporal Clustering of Speech for Single and Multiple F0 Contour Estimation in Noisy Environments.Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveign, Shigeki Sagayama
2006InterspeechSpeech analyzer using a joint estimation model of spectral envelope and fine structure.Hirokazu Kameoka, Jonathan Le Roux, Nobutaka Ono, Shigeki Sagayama
2005InterspeechOptimization methods for discriminative training.Jonathan Le Roux, Erik McDermott
2002WSCGFast Algorithms of Plant Computation Based on Substructure Instances.Hongping Yan, Jean Franois Barczi, Philippe de Reffye, Bao-Gang Hu, Marc Jaeger, Jonathan Le Roux