Hagen Soltau
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
58
Venues
7
Active years
1998–2025
Best venue rank
A*
Where they publish
Papers
58 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | CVPR | Learning Visual Composition through Improved Semantic Guidance. | Austin Stone, Hagen Soltau, Robert Geirhos, Xi Yi, Ye Xia, Bingyi Cao, Kaifeng Chen, Abhijit Ogale, Jonathon Shlens |
| 2024 | ICASSP | Retrieval Augmented End-to-End Spoken Dialog Models. | Mingqiu Wang, Izhak Shafran, Hagen Soltau, Wei Han, Yuan Cao, Dian Yu, Laurent El Shafey |
| 2023 | ASRU | Detecting Speech Abnormalities With a Perceiver-Based Sequence Classifier that Leverages a Universal Speech Model. | Hagen Soltau, Izhak Shafran, Alex Ottenwess, Joseph R. Duffy, Rene L. Utianski, Leland R. Barnard, John L. Stricker, Daniela A. Wiepert, David T. Jones, Hugo Botha |
| 2023 | ASRU | SLM: Bridge the Thin Gap Between Speech and Text Foundation Models. | Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu, Chung-Cheng Chiu, Yuan Cao, Nanxin Chen, Yu Zhang, Hagen Soltau, Paul K. Rubenstein, Lukas Zilka, Dian Yu, Golan Pundak, Nikhil Siddhartha, Johan Schalkwyk, Yonghui Wu |
| 2023 | EMNLP | AnyTOD: A Programmable Task-Oriented Dialog System. | Jeffrey Zhao, Yuan Cao, Raghav Gupta, Harrison Lee, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran, Yonghui Wu |
| 2023 | Interspeech | Speech Aware Dialog System Technology Challenge (DSTC11). | Hagen Soltau, Izhak Shafran, Mingqiu Wang, Abhinav Rastogi, Jeffrey Zhao, Ye Jia, Wei Han, Yuan Cao, Aramys Miranda |
| 2022 | EMNLP | Knowledge-grounded Dialog State Tracking. | Dian Yu, Mingqiu Wang, Yuan Cao, Laurent El Shafey, Izhak Shafran, Hagen Soltau |
| 2022 | Interspeech | RNN Transducers for Named Entity Recognition with constraints on alignment for understanding medical conversations. | Hagen Soltau, Izhak Shafran, Mingqiu Wang, Laurent El Shafey |
| 2022 | NAACL | Unsupervised Slot Schema Induction for Task-oriented Dialog. | Dian Yu, Mingqiu Wang, Yuan Cao, Izhak Shafran, Laurent El Shafey, Hagen Soltau |
| 2021 | ASRU | Word-Level Confidence Estimation for RNN Transducers. | Mingqiu Wang, Hagen Soltau, Laurent El Shafey, Izhak Shafran |
| 2021 | Interspeech | Understanding Medical Conversations: Rich Transcription, Confidence Scores & Information Extraction. | Hagen Soltau, Mingqiu Wang, Izhak Shafran, Laurent El Shafey |
| 2020 | LREC | The Medical Scribe: Corpus Development and Model Performance Analyses. | Izhak Shafran, Nan Du, Linh Tran, Amanda Perry, Lauren Keyes, Mark Knichel, Ashley Domin, Lei Huang, Yuhui Chen, Gang Li, Mingqiu Wang, Laurent El Shafey, Hagen Soltau, Justin S. Paul |
| 2019 | ASRU | Monotonic Recurrent Neural Network Transducer and Decoding Strategies. | Anshuman Tripathi, Han Lu, Hasim Sak, Hagen Soltau |
| 2019 | Interspeech | Joint Speech Recognition and Speaker Diarization via Sequence Transduction. | Laurent El Shafey, Hagen Soltau, Izhak Shafran |
| 2017 | ASRU | Reducing the computational complexity for whole word models. | Hagen Soltau, Hank Liao, Hasim Sak |
| 2017 | Interspeech | Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition. | Hagen Soltau, Hank Liao, Hasim Sak |
| 2014 | ICASSP | Out-of-vocabulary word detection in a speech-to-speech translation system. | Hong-Kwang Kuo, Ellen Eide Kislal, Lidia Mangu, Hagen Soltau, Toms Beran |
| 2014 | ICASSP | Efficient spoken term detection using confusion networks. | Lidia Mangu, Brian Kingsbury, Hagen Soltau, Hong-Kwang Kuo, Michael Picheny |
| 2014 | ICASSP | Progress in dynamic network decoding. | David Nolden, Hagen Soltau, Hermann Ney |
| 2014 | ICASSP | A comparison of two optimization techniques for sequence discriminative training of deep neural networks. | George Saon, Hagen Soltau |
| 2014 | ICASSP | Joint training of convolutional and non-convolutional neural networks. | Hagen Soltau, George Saon, Tara N. Sainath |
| 2014 | ICASSP | Analyzing convolutional neural networks for speech activity detection in mismatched acoustic conditions. | Samuel Thomas, Sriram Ganapathy, George Saon, Hagen Soltau |
| 2014 | Interspeech | Removing redundancy from lattices. | David Nolden, Hagen Soltau, Daniel Povey, Pegah Ghahremani, Lidia Mangu, Hermann Ney |
| 2014 | Interspeech | Unfolded recurrent neural networks for speech recognition. | George Saon, Hagen Soltau, Ahmad Emami, Michael Picheny |
| 2013 | ASRU | The IBM keyword search system for the DARPA RATS program. | Lidia Mangu, Hagen Soltau, Hong-Kwang Kuo, George Saon |
| 2013 | ASRU | Improvements to Deep Convolutional Neural Networks for LVCSR. | Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Toms Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
| 2013 | ASRU | Speaker adaptation of neural network acoustic models using i-vectors. | George Saon, Hagen Soltau, David Nahamoo, Michael Picheny |
| 2013 | ICASSP | Exploiting diversity for spoken term detection. | Lidia Mangu, Hagen Soltau, Hong-Kwang Kuo, Brian Kingsbury, George Saon |
| 2013 | ICASSP | Morpheme-based feature-rich language models using Deep Neural Networks for LVCSR of Egyptian Arabic. | Amr El-Desoky Mousa, Hong-Kwang Jeff Kuo, Lidia Mangu, Hagen Soltau |
| 2013 | Interspeech | The IBM speech activity detection system for the DARPA RATS program. | George Saon, Samuel Thomas, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury |
| 2013 | Interspeech | Neural network acoustic models for the DARPA RATS program. | Hagen Soltau, Hong-Kwang Kuo, Lidia Mangu, George Saon, Toms Beran |
| 2012 | Interspeech | Scalable Minimum Bayes Risk Training of Deep Neural Network Acoustic Models Using Distributed Hessian-free Optimization. | Brian Kingsbury, Tara N. Sainath, Hagen Soltau |
| 2011 | ASRU | The IBM 2011 GALE Arabic speech transcription system. | Lidia Mangu, Hong-Kwang Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy |
| 2011 | ASRU | From Modern Standard Arabic to Levantine ASR: Leveraging GALE for dialects. | Hagen Soltau, Lidia Mangu, Fadi Biadsy |
| 2011 | ICASSP | The IBM 2009 GALE Arabic speech transcription system. | Brian Kingsbury, Hagen Soltau, George Saon, Stephen M. Chu, Hong-Kwang Kuo, Lidia Mangu, Suman V. Ravuri, Nelson Morgan, Adam Janin |
| 2010 | ICASSP | A comparative study on system combination schemes for LVCSR. | Chengyuan Ma, Hong-Kwang Jeff Kuo, Hagen Soltau, Xiaodong Cui, Upendra V. Chaudhari, Lidia Mangu, Chin-Hui Lee |
| 2010 | ICASSP | The IBM 2008 GALE Arabic speech transcription system. | George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Kuo, Lidia Mangu, Daniel Povey |
| 2010 | Interspeech | Decoding with shrinkage-based language models. | Ahmad Emami, Stanley F. Chen, Abraham Ittycheriah, Hagen Soltau, Bing Zhao |
| 2010 | Interspeech | Boosting systems for LVCSR. | George Saon, Hagen Soltau |
| 2009 | ASRU | Dynamic network decoding revisited. | Hagen Soltau, George Saon |
| 2009 | ICASSP | Large margin semi-tied covariance transforms for discriminative training. | George Saon, Daniel Povey, Hagen Soltau |
| 2008 | Interspeech | Fast speaker adaptive training for speech recognition. | Daniel Povey, Hong-Kwang Jeff Kuo, Hagen Soltau |
| 2007 | ICASSP | The IBM 2006 Gale Arabic ASR System. | Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Geoffrey Zweig |
| 2005 | ICASSP | fMPE: Discriminatively Trained Features for Speech Recognition. | Daniel Povey, Brian Kingsbury, Lidia Mangu, George Saon, Hagen Soltau, Geoffrey Zweig |
| 2005 | ICASSP | The IBM 2004 Conversational Telephony System for Rich Transcription. | Hagen Soltau, Brian Kingsbury, Lidia Mangu, Daniel Povey, George Saon, Geoffrey Zweig |
| 2004 | ICASSP | The 2003 ISL rich transcription system for conversational telephony speech. | Hagen Soltau, Hua Yu, Florian Metze, Christian Fgen, Qin Jin, Szu-Chen Stan Jou |
| 2002 | ICASSP | Efficient language model lookahead through polymorphic linguistic context assignment. | Hagen Soltau, Florian Metze, Christian Fgen, Alex Waibel |
| 2002 | Interspeech | Compensating for hyperarticulation by modeling articulatory properties. | Hagen Soltau, Florian Metze, Alex Waibel |
| 2001 | ICASSP | Speaker compensation with sine-log all-pass transforms. | John W. McDonough, Florian Metze, Hagen Soltau, Alex Waibel |
| 2001 | ICASSP | The ISL evaluation system for Verbmobil-II. | Hagen Soltau, Thomas Schaaf, Florian Metze, Alex Waibel |
| 2001 | ICASSP | Advances in automatic meeting record creation and access. | Alex Waibel, Michael Bett, Florian Metze, Klaus Ries, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu, Klaus Zechner |
| 2001 | Interspeech | Speech recognition over netmeeting connections. | Florian Metze, John W. McDonough, Hagen Soltau |
| 2001 | NAACL | Advances in meeting recognition. | Alex Waibel, Hua Yu, Tanja Schultz, Yue Pan, Michael Bett, Martin Westphal, Hagen Soltau, Thomas Schaaf, Florian Metze |
| 2000 | ICASSP | Confidence measure based language identification. | Florian Metze, Thomas Kemp, Thomas Schaaf, Tanja Schultz, Hagen Soltau |
| 2000 | ICASSP | Specialized acoustic models for hyperarticulated speech. | Hagen Soltau, Alex Waibel |
| 2000 | Interspeech | Phone dependent modeling of hyperarticulated effects#. | Hagen Soltau, Alex Waibel |
| 1998 | ICASSP | Recognition of music types. | Hagen Soltau, Tanja Schultz, Martin Westphal, Alex Waibel |
| 1998 | Interspeech | On the influence of hyperarticulated speech on recognition performance. | Hagen Soltau, Alex Waibel |