Hasim Sak
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
41
Venues
5
Active years
2007–2024
Best venue rank
A*
Where they publish
Papers
41 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2024 | ICASSP | Monte Carlo Self-Training for Speech Recognition. | Anshuman Tripathi, Soheil Khorram, Han Lu, Jaeyoung Kim, Qian Zhang, Hasim Sak |
| 2023 | ICASSP | Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition. | Soheil Khorram, Anshuman Tripathi, Jaeyoung Kim, Han Lu, Qian Zhang, Rohit Prabhavalkar, Hasim Sak |
| 2022 | ICASSP | Contrastive Siamese Network for Semi-Supervised Speech Recognition. | Soheil Khorram, Jaeyoung Kim, Anshuman Tripathi, Han Lu, Qian Zhang, Hasim Sak |
| 2022 | ICASSP | Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection. | Wei Xia, Han Lu, Quan Wang, Anshuman Tripathi, Yiling Huang, Ignacio Lpez-Moreno, Hasim Sak |
| 2021 | Interspeech | Reducing Streaming ASR Model Delay with Self Alignment. | Jaeyoung Kim, Han Lu, Anshuman Tripathi, Qian Zhang, Hasim Sak |
| 2020 | ICASSP | End-To-End Multi-Talker Overlapping Speech Recognition. | Anshuman Tripathi, Han Lu, Hasim Sak |
| 2020 | ICASSP | Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss. | Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, Shankar Kumar |
| 2020 | Interspeech | Multilingual Speech Recognition with Self-Attention Structured Parameterization. | Yun Zhu, Parisa Haghani, Anshuman Tripathi, Bhuvana Ramabhadran, Brian Farris, Hainan Xu, Han Lu, Hasim Sak, Isabel Leal, Neeraj Gaur, Pedro J. Moreno, Qian Zhang |
| 2019 | ASRU | A Density Ratio Approach to Language Model Fusion in End-to-End Automatic Speech Recognition. | Erik McDermott, Hasim Sak, Ehsan Variani |
| 2019 | ASRU | Monotonic Recurrent Neural Network Transducer and Decoding Strategies. | Anshuman Tripathi, Han Lu, Hasim Sak, Hagen Soltau |
| 2019 | Interspeech | Large-Scale Visual Speech Recognition. | Brendan Shillingford, Yannis M. Assael, Matthew W. Hoffman, Thomas Paine, Can Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew W. Senior, Nando de Freitas |
| 2018 | Interspeech | Speech Recognition for Medical Conversations. | Chung-Cheng Chiu, Anshuman Tripathi, Katherine Chou, Chris Co, Navdeep Jaitly, Diana Jaunzeikare, Anjuli Kannan, Patrick Nguyen, Hasim Sak, Ananth Sankar, Justin Tansuwan, Nathan Wan, Yonghui Wu, Xuedong Zhang |
| 2017 | ASRU | Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer. | Kanishka Rao, Hasim Sak, Rohit Prabhavalkar |
| 2017 | ASRU | Reducing the computational complexity for whole word models. | Hagen Soltau, Hank Liao, Hasim Sak |
| 2017 | ICASSP | Multi-accent speech recognition with hierarchical grapheme based models. | Kanishka Rao, Hasim Sak |
| 2017 | Interspeech | Acoustic Modeling for Google Home. | Bo Li, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
| 2017 | Interspeech | Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping. | Hasim Sak, Matt Shannon, Kanishka Rao, Franoise Beaufays |
| 2017 | Interspeech | Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition. | Hagen Soltau, Hank Liao, Hasim Sak |
| 2016 | ICASSP | Personalized speech recognition on mobile devices. | Ian McGraw, Rohit Prabhavalkar, Raziel Alvarez, Montse Gonzalez Arenas, Kanishka Rao, David Rybach, Ouais Alsharif, Hasim Sak, Alexander Gruenstein, Franoise Beaufays, Carolina Parada |
| 2016 | ICASSP | Flat start training of CD-CTC-SMBR LSTM RNN acoustic models. | Kanishka Rao, Andrew W. Senior, Hasim Sak |
| 2015 | ASRU | Acoustic modelling with CD-CTC-SMBR LSTM RNNS. | Andrew W. Senior, Hasim Sak, Felix de Chaumont Quitry, Tara N. Sainath, Kanishka Rao |
| 2015 | ICASSP | Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks. | Kanishka Rao, Fuchun Peng, Hasim Sak, Franoise Beaufays |
| 2015 | ICASSP | Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks. | Tara N. Sainath, Oriol Vinyals, Andrew W. Senior, Hasim Sak |
| 2015 | ICASSP | Learning acoustic frame labeling for speech recognition with recurrent neural networks. | Hasim Sak, Andrew W. Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Franoise Beaufays, Johan Schalkwyk |
| 2015 | ICASSP | Context dependent phone models for LSTM RNN acoustic modelling. | Andrew W. Senior, Hasim Sak, Izhak Shafran |
| 2015 | ICASSP | Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis. | Heiga Zen, Hasim Sak |
| 2015 | Interspeech | Fast and accurate recurrent neural network acoustic models for speech recognition. | Hasim Sak, Andrew W. Senior, Kanishka Rao, Franoise Beaufays |
| 2014 | Interspeech | Automatic language identification using long short-term memory recurrent neural networks. | Javier Gonzalez-Dominguez, Ignacio Lpez-Moreno, Hasim Sak, Joaquin Gonzalez-Rodriguez, Pedro J. Moreno |
| 2014 | Interspeech | Long short-term memory recurrent neural network architectures for large scale acoustic modeling. | Hasim Sak, Andrew W. Senior, Franoise Beaufays |
| 2014 | Interspeech | Sequence discriminative distributed training of long short-term memory recurrent neural networks. | Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao |
| 2013 | ASRU | Mixture of mixture n-gram language models. | Hasim Sak, Cyril Allauzen, Kaisuke Nakajima, Franoise Beaufays |
| 2013 | ICASSP | Language model verbalization for automatic speech recognition. | Hasim Sak, Franoise Beaufays, Kaisuke Nakajima, Cyril Allauzen |
| 2013 | Interspeech | Written-domain language modeling for automatic speech recognition. | Hasim Sak, Yun-Hsuan Sung, Franoise Beaufays, Cyril Allauzen |
| 2012 | ICASSP | Semi-supervised discriminative language modeling for Turkish ASR. | Arda elebi, Hasim Sak, Erin Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Daniel M. Bikel, Chris Callison-Burch, Yuan Cao, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
| 2011 | ASRU | Discriminative reranking of ASR hypotheses with morpholexical and N-best-list features. | Hasim Sak, Murat Saraclar, Tunga Gungor |
| 2010 | ICASSP | Morphology-based and sub-word language modeling for Turkish speech recognition. | Hasim Sak, Murat Saraclar, Tunga Gngr |
| 2010 | Interspeech | On-the-fly lattice rescoring for real-time automatic speech recognition. | Hasim Sak, Murat Saraclar, Tunga Gngr |
| 2009 | ACL | A Stochastic Finite-State Morphological Parser for Turkish. | Hasim Sak, Tunga Gngr, Murat Saraclar |
| 2009 | ASRU | Integrating morphology into automatic speech recognition. | Hasim Sak, Murat Saraclar, Tunga Gngr |
| 2007 | CICLING | Morphological Disambiguation of Turkish Text with Perceptron Algorithm. | Hasim Sak, Tunga Gngr, Murat Saraclar |
| 2007 | Interspeech | Language modeling for automatic turkish broadcast news transcription. | Ebru Arisoy, Hasim Sak, Murat Saraclar |