Florian Metze
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
163
Venues
17
Active years
2000–2026
Best venue rank
A*
Where they publish
Papers
163 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | EACL | Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning. | Minseok Kim, Jingxiang Chen, Seong-Gyun Leem, Yin Huang, Rashi Rungta, Zhicheng Ouyang, Haibin Wu, Surya Teja Appini, Ankur Bansal, Yang Bai, Yue Liu, Florian Metze, Ahmed Aly, Anuj Kumar, Ariya Rastrow, Zhaojiang Lin |
| 2025 | ASRU | Long-Form Fuzzy Speech-to-Text Alignment for 1000+ Languages. | Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun, Florian Metze |
| 2025 | ASRU | MMW: Side Talk Rejection Multi-Microphone Whisper On Smart Glasses. | Yang Liu, Li Wan, Yiteng Huang, Yong Xu, Yangyang Shi, Saurabh Adya, Ming Sun, Florian Metze |
| 2025 | Interspeech | Directional Speech Recognition with Full-Duplex Capability. | Ju Lin, Yiteng Huang, Ming Sun, Frank Seide, Florian Metze |
| 2025 | Interspeech | MASV: Speaker Verification with Global and Local Context Mamba. | Yang Liu, Li Wan, Yiteng Huang, Ming Sun, Xinhao Mei, Xubo Liu, Yangyang Shi, Florian Metze |
| 2025 | Interspeech | Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition. | Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Zhaojun Yang, Peng Su, Prashant Rawat, Sangeeta Srivastava, Ming Sun, Florian Metze |
| 2024 | ICASSP | Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation. | Jackson Michaels, Juncheng B. Li, Laura Yao, Lijun Yu, Zach Wood-Doughty, Florian Metze |
| 2023 | EACL | CTC Alignments Improve Autoregressive Translation. | Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W. Black, Shinji Watanabe |
| 2022 | ACL | Zero-shot Learning for Grapheme to Phoneme Conversion with Language Ensemble. | Xinjian Li, Florian Metze, David R. Mortensen, Shinji Watanabe, Alan W. Black |
| 2022 | CVPR | Self-supervised object detection from audio-visual correspondence. | Triantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi, Florian Metze |
| 2022 | EMNLP | Token-level Sequence Labeling for Spoken Language Understanding using Compositional End-to-End Models. | Siddhant Arora, Siddharth Dalmia, Brian Yan, Florian Metze, Alan W. Black, Shinji Watanabe |
| 2022 | EMNLP | On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization. | Shruti Palaskar, Akshita Bhagia, Yonatan Bisk, Florian Metze, Alan W. Black, Ana Marasovic |
| 2022 | EMNLP | Normalized Contrastive Learning for Text-Video Retrieval. | Yookoon Park, Mahmoud Azab, Seungwhan Moon, Bo Xiong, Florian Metze, Gourab Kundu, Ahmed Kirmani |
| 2022 | ICASSP | On Adversarial Robustness Of Large-Scale Audio Visual Learning. | Juncheng B. Li, Shuhui Qu, Xinjian Li, Bernie Po-Yao Huang, Florian Metze |
| 2022 | ICASSP | End-to-End Speech Summarization Using Restricted Self-Attention. | Roshan Sharma, Shruti Palaskar, Alan W. Black, Florian Metze |
| 2022 | Interspeech | AudioTagging Done Right: 2nd comparison of deep learning methods for environmental sound classification. | Juncheng Li, Shuhui Qu, Po-Yao Huang, Florian Metze |
| 2022 | Interspeech | ASR2K: Speech Recognition for Around 2000 Languages without Audio. | Xinjian Li, Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe |
| 2022 | LREC | Phone Inventories and Recognition for Every Language. | Xinjian Li, Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe |
| 2021 | ACL | VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding. | Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, Luke Zettlemoyer |
| 2021 | CVPR | How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign Language. | Amanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, Xavier Gir-i-Nieto |
| 2021 | EACL | NoiseQA: Challenge Set Evaluation for User-Centric Question Answering. | Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard H. Hovy, Alan W. Black |
| 2021 | EMNLP | VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. | Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, Christoph Feichtenhofer |
| 2021 | ICASSP | Phone Distribution Estimation for Low Resource Languages. | Xinjian Li, Juncheng Li, Jiali Yao, Alan W. Black, Florian Metze |
| 2021 | ICASSP | Multilingual Phonetic Dataset for Low Resource Speech Recognition. | Xinjian Li, David R. Mortensen, Florian Metze, Alan W. Black |
| 2021 | ICASSP | Audio-Visual Event Recognition Through the Lens of Adversary. | Juncheng B. Li, Kaixin Ma, Shuhui Qu, Po-Yao Huang, Florian Metze |
| 2021 | ICCV | Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning. | Mandela Patrick, Po-Yao Huang, Ishan Misra, Florian Metze, Andrea Vedaldi, Yuki M. Asano, Joo F. Henriques |
| 2021 | ICLR | Support-set bottlenecks for video-text representation learning. | Mandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze, Alexander G. Hauptmann, Joo F. Henriques, Andrea Vedaldi |
| 2021 | Interspeech | Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language Understanding. | Siddhant Arora, Alissa Ostapenko, Vijay Viswanathan, Siddharth Dalmia, Florian Metze, Shinji Watanabe, Alan W. Black |
| 2021 | Interspeech | Hierarchical Phone Recognition with Compositional Phonetics. | Xinjian Li, Juncheng Li, Florian Metze, Alan W. Black |
| 2021 | Interspeech | Multimodal Speech Summarization Through Semantic Concept Learning. | Shruti Palaskar, Ruslan Salakhutdinov, Alan W. Black, Florian Metze |
| 2021 | Interspeech | Differentiable Allophone Graphs for Language-Universal Speech Recognition. | Brian Yan, Siddharth Dalmia, David R. Mortensen, Florian Metze, Shinji Watanabe |
| 2021 | NAACL | Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks. | Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze, Shinji Watanabe |
| 2021 | NAACL | Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models. | Poyao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alex Hauptmann |
| 2020 | AAAI | Towards Zero-Shot Learning for Automatic Phonemic Transcription. | Xinjian Li, Siddharth Dalmia, David R. Mortensen, Juncheng Li, Alan W. Black, Florian Metze |
| 2020 | EMNLP | On Long-Tailed Phenomena in Neural Machine Translation. | Vikas Raunak, Siddharth Dalmia, Vivek Gupta, Florian Metze |
| 2020 | EMNLP | Fine-Grained Grounding for Multimodal Speech Recognition. | Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott |
| 2020 | ICASSP | Universal Phone Recognition with a Multilingual Allophone System. | Xinjian Li, Siddharth Dalmia, Juncheng Li, Matthew Lee, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze |
| 2020 | ICASSP | ASR Error Correction and Domain Adaptation Using Machine Translation. | Anirudh Mani, Shruti Palaskar, Nimshi Venkat Meripo, Sandeep Konam, Florian Metze |
| 2020 | ICASSP | Looking Enhances Listening: Recovering Missing Speech Using Images. | Tejas Srinivasan, Ramon Sanabria, Florian Metze |
| 2020 | Interspeech | Contextual RNN-T for Open Domain ASR. | Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig, Florian Metze, Yatharth Saraf |
| 2020 | Interspeech | Towards Context-Aware End-to-End Code-Switching Speech Recognition. | Zimeng Qiu, Yiyuan Li, Xinjian Li, Florian Metze, William M. Campbell |
| 2020 | LREC | AlloVera: A Multilingual Allophone Database. | David R. Mortensen, Xinjian Li, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W. Black, Florian Metze, Graham Neubig |
| 2019 | ACL | Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion. | Suyoun Kim, Siddharth Dalmia, Florian Metze |
| 2019 | ACL | Multimodal Abstractive Summarization for How2 Videos. | Shruti Palaskar, Jindrich Libovick, Spandana Gella, Florian Metze |
| 2019 | ICASSP | A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling. | Yun Wang, Juncheng Li, Florian Metze |
| 2019 | ICASSP | Connectionist Temporal Localization for Sound Event Detection with Sequential Labeling. | Yun Wang, Florian Metze |
| 2019 | ICASSP | Multimodal Grounding for Sequence-to-sequence Speech Recognition. | Ozan Caglayan, Ramon Sanabria, Shruti Palaskar, Loc Barrault, Florian Metze |
| 2019 | ICASSP | Phoneme Level Language Models for Sequence Based Low Resource ASR. | Siddharth Dalmia, Xinjian Li, Alan W. Black, Florian Metze |
| 2019 | ICASSP | Learning from Multiview Correlations in Open-domain Videos. | Nils Holzenberger, Shruti Palaskar, Pranava Madhyastha, Florian Metze, Raman Arora |
| 2019 | ICASSP | Learned in Speech Recognition: Contextual Acoustic Word Embeddings. | Shruti Palaskar, Vikas Raunak, Florian Metze |
| 2019 | INLG | On Leveraging the Visual Modality for Neural Machine Translation. | Vikas Raunak, Sang Keun Choe, Quanyang Lu, Yi Xu, Florian Metze |
| 2019 | Interspeech | Cross-Attention End-to-End ASR for Two-Party Conversations. | Suyoun Kim, Siddharth Dalmia, Florian Metze |
| 2019 | Interspeech | Multilingual Speech Recognition with Corpus Relatedness Sampling. | Xinjian Li, Siddharth Dalmia, Alan W. Black, Florian Metze |
| 2019 | Interspeech | SANTLR: Speech Annotation Toolkit for Low Resource Languages. | Xinjian Li, Zhong Zhou, Siddharth Dalmia, Alan W. Black, Florian Metze |
| 2019 | Interspeech | Survey Talk: Multimodal Processing of Speech and Language. | Florian Metze |
| 2019 | NAACL | Acoustic-to-Word Models with Conversational Context Information. | Suyoun Kim, Florian Metze |
| 2018 | ICASSP | Sequence-Based Multi-Lingual Low Resource Speech Recognition. | Siddharth Dalmia, Ramon Sanabria, Florian Metze, Alan W. Black |
| 2018 | ICASSP | A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging. | Juncheng Li, Yun Wang, Joseph Szurley, Florian Metze, Samarjit Das |
| 2018 | ICASSP | End-to-end Multimodal Speech Recognition. | Shruti Palaskar, Ramon Sanabria, Florian Metze |
| 2018 | ICASSP | Enhancement and Analysis of Conversational Speech: JSALT 2017. | Neville Ryant, Elika Bergelson, Kenneth Church, Alejandrina Cristi, Jun Du, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Jn Profant, Lei Sun, Yu Tsao, Zhou Yu |
| 2018 | ICASSP | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop. | Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stker, Pierre Godard, Markus Mller, Lucas Ondel, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang, Emmanuel Dupoux |
| 2018 | Interspeech | The ACLEW DiViMe: An Easy-to-use Diarization Tool. | Adrien Le Franc, Eric Riebling, Julien Karadayi, Yun Wang, Camila Scaff, Florian Metze, Alejandrina Cristi |
| 2018 | Interspeech | Multiple Instance Deep Learning for Weakly Supervised Small-Footprint Audio Event Detection. | Shao-Yen Tseng, Juncheng Li, Yun Wang, Florian Metze, Joseph Szurley, Samarjit Das |
| 2018 | Interspeech | Comparing the Max and Noisy-Or Pooling Functions in Multiple Instance Learning for Weakly Supervised Sequence Learning Tasks. | Yun Wang, Juncheng Li, Florian Metze |
| 2018 | Interspeech | Subword and Crossword Units for CTC Acoustic Models. | Thomas Zenkel, Ramon Sanabria, Florian Metze, Alex Waibel |
| 2018 | LREC | Annotating High-Level Structures of Short Stories and Personal Anecdotes. | Boyang Li, Beth Cardier, Tong Wang, Florian Metze |
| 2017 | ICASSP | Visual features for context-aware speech recognition. | Abhinav Gupta, Yajie Miao, Leonardo Neves, Florian Metze |
| 2017 | ICASSP | A comparison of Deep Learning methods for environmental sound detection. | Juncheng Li, Wei Dai, Florian Metze, Shuhui Qu, Samarjit Das |
| 2017 | ICASSP | A first attempt at polyphonic sound event detection using connectionist temporal classification. | Yun Wang, Florian Metze |
| 2017 | Interspeech | A Transfer Learning Based Feature Extractor for Polyphonic Sound Event Detection Using Connectionist Temporal Classification. | Yun Wang, Florian Metze |
| 2017 | Interspeech | Comparison of Decoding Strategies for CTC Acoustic Models. | Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stker, Alex Waibel |
| 2016 | ICASSP | An empirical exploration of CTC acoustic models. | Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, Alexander Waibel |
| 2016 | ICASSP | Audio-based multimedia event detection using deep recurrent neural networks. | Yun Wang, Leonardo Neves, Florian Metze |
| 2016 | Interspeech | Experiences with Shared Resources for Research and Education in Speech and Language Processing. | Rebecca Bates, Eric Fosler-Lussier, Florian Metze, Martha A. Larson, Gina-Anne Levow, Emily Mower Provost |
| 2016 | Interspeech | Manipulating Word Lattices to Incorporate Human Corrections. | Yashesh Gaur, Florian Metze, Jeffrey P. Bigham |
| 2016 | Interspeech | Virtual Machines and Containers as a Platform for Experimentation. | Florian Metze, Eric Riebling, Anne S. Warlaumont, Elika Bergelson |
| 2016 | Interspeech | Open-Domain Audio-Visual Speech Recognition: A Deep Learning Approach. | Yajie Miao, Florian Metze |
| 2015 | ASRU | EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding. | Yajie Miao, Mohammad Gowayyed, Florian Metze |
| 2015 | ICASSP | QUESST2014: Evaluating Query-by-Example Speech Search in a zero-resource setting with real-life queries. | Xavier Anguera, Luis Javier Rodrguez-Fuentes, Andi Buzo, Florian Metze, Igor Szke, Mikel Peagarikano |
| 2015 | ICASSP | Semi-supervised training in low-resource ASR and KWS. | Florian Metze, Ankur Gandhe, Yajie Miao, Zaid Sheikh, Yun Wang, Di Xu, Hao Zhang, Jungsuk Kim, Ian R. Lane, Wonkyum Lee, Sebastian Stker, Markus Mller |
| 2015 | ICASSP | Regularizing DNN acoustic models with Gaussian stochastic neurons. | Hao Zhang, Yajie Miao, Florian Metze |
| 2015 | Interspeech | Using keyword spotting to help humans correct captioning faster. | Yashesh Gaur, Florian Metze, Yajie Miao, Jeffrey P. Bigham |
| 2015 | Interspeech | The speech recognition virtual kitchen turns one. | Florian Metze, Eric Riebling, Eric Fosler-Lussier, Andrew R. Plummer, Rebecca Bates |
| 2015 | Interspeech | Distance-aware DNNs for robust speech recognition. | Yajie Miao, Florian Metze |
| 2015 | Interspeech | On speaker adaptation of long short-term memory recurrent neural networks. | Yajie Miao, Florian Metze |
| 2014 | ACL | Semantics for Large-Scale Multimedia: New Challenges for NLP. | Florian Metze, Koichi Shinoda |
| 2014 | EACL | Augmenting Translation Models with Simulated Acoustic Confusions for Improved Spoken Language Translation. | Yulia Tsvetkov, Florian Metze, Chris Dyer |
| 2014 | ICASSP | Optimization of Neural Network Language Models for keyword search. | Ankur Gandhe, Florian Metze, Alex Waibel, Ian R. Lane |
| 2014 | ICASSP | Exploring audio semantic concepts for event-based video retrieval. | Yipei Wang, Shourabh Rawat, Florian Metze |
| 2014 | ICASSP | Semi-automatic audio semantic concept discovery for multimedia retrieval. | Yipei Wang, Shourabh Rawat, Florian Metze |
| 2014 | Interspeech | Query-by-example spoken term detection on multilingual unconstrained speech. | Xavier Anguera, Luis Javier Rodrguez-Fuentes, Igor Szke, Andi Buzo, Florian Metze, Mikel Peagarikano |
| 2014 | Interspeech | Neural network language models for low resource languages. | Ankur Gandhe, Florian Metze, Ian R. Lane |
| 2014 | Interspeech | Improving language-universal feature extraction with deep maxout and convolutional neural networks. | Yajie Miao, Florian Metze |
| 2014 | Interspeech | Distributed learning of multilingual DNN feature extractors using GPUs. | Yajie Miao, Hao Zhang, Florian Metze |
| 2014 | Interspeech | Towards speaker adaptive training of deep neural network acoustic models. | Yajie Miao, Hao Zhang, Florian Metze |
| 2014 | Interspeech | The speech recognition virtual kitchen: launch party. | Andrew R. Plummer, Eric Riebling, Anuj Kumar, Florian Metze, Eric Fosler-Lussier, Rebecca Bates |
| 2014 | Interspeech | An in-depth comparison of keyword specific thresholding and sum-to-one score normalization. | Yun Wang, Florian Metze |
| 2014 | Interspeech | Word-based probabilistic phonetic retrieval for low-resource spoken term detection. | Di Xu, Florian Metze |
| 2013 | ASRU | Using web text to improve keyword spotting in speech. | Ankur Gandhe, Long Qin, Florian Metze, Alexander I. Rudnicky, Ian R. Lane, Matthias Eck |
| 2013 | ASRU | DNN acoustic modeling with modular multi-lingual feature extraction networks. | Jonas Gehring, Quoc Bao Nguyen, Florian Metze, Alex Waibel |
| 2013 | ASRU | Models of tone for tonal and non-tonal languages. | Florian Metze, Zaid Sheikh, Alex Waibel, Jonas Gehring, Kevin Kilgour, Quoc Bao Nguyen, Van Huy Nguyen |
| 2013 | ASRU | Deep maxout networks for low-resource speech recognition. | Yajie Miao, Florian Metze, Shourabh Rawat |
| 2013 | ASRU | Neighbour selection and adaptation for rapid speaker-dependent ASR. | Udhyakumar Nallasamy, Mark C. Fuhs, Monika Woszczyna, Florian Metze, Tanja Schultz |
| 2013 | ICASSP | Extracting deep bottleneck features using stacked auto-encoders. | Jonas Gehring, Yajie Miao, Florian Metze, Alex Waibel |
| 2013 | ICASSP | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition. | Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson, Sanjeev Khudanpur, Kenneth Church, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin Bennett, Benjamin Brschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas |
| 2013 | ICASSP | The spoken web search task at MediaEval 2012. | Florian Metze, Xavier Anguera, Etienne Barnard, Marelie H. Davel, Guillaume Gravier |
| 2013 | ICASSP | Subspace mixture model for low-resource speech recognition in cross-lingual settings. | Yajie Miao, Florian Metze, Alex Waibel |
| 2013 | ICASSP | Learning discriminative basis coefficients for eigenspace MLLR unsupervised adaptation. | Yajie Miao, Florian Metze, Alex Waibel |
| 2013 | ICASSP | Identification and modeling of word fragments in spontaneous speech. | Yulia Tsvetkov, Zaid Sheikh, Florian Metze |
| 2013 | IJCNLP | Prosody-Based Unsupervised Speech Summarization with Two-Layer Mutually Reinforced Random Walk. | Sujay Kumar Jauhar, Yun-Nung Chen, Florian Metze |
| 2013 | Interspeech | Multi-layer mutually reinforced random walk with hidden parameters for improved multi-party meeting summarization. | Yun-Nung Chen, Florian Metze |
| 2013 | Interspeech | Formalizing expert knowledge for developing accurate speech recognizers. | Anuj Kumar, Florian Metze, Wenyi Wang, Matthew Kam |
| 2013 | Interspeech | The speech recognition virtual kitchen. | Florian Metze, Eric Fosler-Lussier, Rebecca Bates |
| 2013 | Interspeech | Improving low-resource CD-DNN-HMM using dropout and multilingual DNN training. | Yajie Miao, Florian Metze |
| 2013 | Interspeech | Robust audio-codebooks for large-scale event detection in consumer videos. | Shourabh Rawat, Peter F. Schulam, Susanne Burger, Duo Ding, Yipei Wang, Florian Metze |
| 2012 | ICASSP | Articulatory features for expressive speech synthesis. | Alan W. Black, H. Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry, Tim Polzehl, Kishore Prahallad, Stefan Steidl, Callie Vaughn |
| 2012 | ICASSP | The Spoken Web Search Task at MediaEval 2011. | Florian Metze, Nitendra Rajput, Xavier Anguera, Marelie H. Davel, Guillaume Gravier, Charl Johannes van Heerden, Gautam Varma Mantena, Armando Muscariello, Kishore Prahallad, Igor Szke, Javier Tejedor |
| 2012 | INLG | Generating Natural Language Summaries for Multimedia. | Duo Ding, Florian Metze, Shourabh Rawat, Peter Franz Schulam, Susanne Burger |
| 2012 | Interspeech | Integrating Intra-Speaker Topic Modeling and Temporal-Based Inter-Speaker Topic Modeling in Random Walk for Improved Multi-Party Meeting Summarization. | Yun-Nung Chen, Florian Metze |
| 2012 | Interspeech | Event-based Video Retrieval Using Audio. | Qin Jin, Peter Franz Schulam, Shourabh Rawat, Susanne Burger, Duo Ding, Florian Metze |
| 2012 | Interspeech | The Speech Recognition Virtual Kitchen: An Initial Prototype. | Florian Metze, Eric Fosler-Lussier |
| 2012 | Interspeech | Enhanced Polyphone Decision Tree Adaptation for Accented Speech Recognition. | Udhyakumar Nallasamy, Florian Metze, Tanja Schultz |
| 2012 | Interspeech | On Speaker-Independent Personality Perception and Prediction from Speech. | Tim Polzehl, Katrin Schoenenberg, Sebastian Mller, Florian Metze, Gelareh Mohammadi, Alessandro Vinciarelli |
| 2012 | Interspeech | Initialization Schemes for Multilayer Perceptron Training and their Impact on ASR Performance using Multilingual Data. | Ngoc Thang Vu, Wojtek Breiter, Florian Metze, Tanja Schultz |
| 2012 | NAACL | Intra-Speaker Topic Modeling for Improved Multi-Party Meeting Summarization with Integrated Random Walk. | Yun-Nung Chen, Florian Metze |
| 2011 | HCI | A Review of Personality in Voice-Based Man Machine Interaction. | Florian Metze, Alan W. Black, Tim Polzehl |
| 2011 | Interspeech | Analysis of Dialectal Influence in Pan-Arabic ASR. | Udhyakumar Nallasamy, Michael Garbus, Florian Metze, Qin Jin, Thomas Schaaf, Tanja Schultz |
| 2011 | Interspeech | Modeling Speaker Personality Using Voice. | Tim Polzehl, Sebastian Mller, Florian Metze |
| 2010 | ICASSP | Late fusion of individual engines for improved recognition of negative emotion in speech - learning vs. democratic vote. | Bjrn W. Schuller, Florian Metze, Stefan Steidl, Anton Batliner, Florian Eyben, Tim Polzehl |
| 2010 | Interspeech | Improvements to generalized discriminative feature transformation for speech recognition. | Roger Hsiao, Florian Metze, Tanja Schultz |
| 2010 | Interspeech | Emotion recognition using imperfect speech recognition. | Florian Metze, Anton Batliner, Florian Eyben, Tim Polzehl, Bjrn W. Schuller, Stefan Steidl |
| 2010 | Interspeech | The 2010 CMU GALE speech-to-text system. | Florian Metze, Roger Hsiao, Qin Jin, Udhyakumar Nallasamy, Tanja Schultz |
| 2010 | Interspeech | Analysis of gender normalization using MLP and VTLN features. | Thomas Schaaf, Florian Metze |
| 2009 | HCI | Reliable Evaluation of Multimodal Dialogue Systems. | Florian Metze, Ina Wechsung, Stefan Schaffer, Julia Seebode, Sebastian Mller |
| 2009 | HCI | Usability Evaluation of Multimodal Interfaces: Is the Whole the Sum of Its Parts? | Ina Wechsung, Klaus-Peter Engelbrecht, Stefan Schaffer, Julia Seebode, Florian Metze, Sebastian Mller |
| 2009 | ICASSP | Detecting real life anger. | Felix Burkhardt, Tim Polzehl, Joachim Stegmann, Florian Metze, Richard Huber |
| 2009 | Interspeech | Emotion classification in children's speech using fusion of acoustic and linguistic features. | Tim Polzehl, Shiva Sundaram, Hamed Ketabdar, Michael Wagner, Florian Metze |
| 2009 | Interspeech | Influence of training on direct and indirect measures for the evaluation of multimodal systems. | Julia Seebode, Stefan Schaffer, Ina Wechsung, Florian Metze |
| 2009 | Interspeech | Predicting the quality of multimodal systems based on judgments of single modalities. | Ina Wechsung, Klaus-Peter Engelbrecht, Anja B. Naumann, Stefan Schaffer, Julia Seebode, Florian Metze, Sebastian Mller |
| 2008 | ICPR | Detecting trends in social bookmarking systems using a probabilistic generative model and smoothing. | Robert Wetzker, Till Plumbaum, Alexander Korth, Christian Bauckhage, Tansu Alpcan, Florian Metze |
| 2008 | Interspeech | User perception of multi-modal interfaces for mobile applications. | Florian Metze, Roman Englert, Udo Bub, Ingmar Kliche, Thomas Scheerbarth |
| 2007 | ICASSP | Spotting using Durational Entropy. | Jitendra Ajmera, Florian Metze |
| 2007 | ICASSP | Comparison of Four Approaches to Age and Gender Recognition for Telephone Applications. | Florian Metze, Jitendra Ajmera, Roman Englert, Udo Bub, Felix Burkhardt, Joachim Stegmann, Christian A. Mller, Richard Huber, Bernt Andrassy, Josef G. Bauer, Bernhard Littel |
| 2007 | NAACL | On using Articulatory Features for Discriminative Speaker Adaptation. | Florian Metze |
| 2007 | SMC | An intelligent knowledge sharing system for web communities. | Christian Bauckhage, Tansu Alpcan, Sachin Agarwal, Florian Metze, Robert Wetzker, Milena Ilic, Sahin Albayrak |
| 2006 | Interspeech | Articulatory features for "meeting" speech recognition. | Florian Metze |
| 2005 | ICASSP | Automatically Transcribing Meetings using Distant Microphones. | Florian Metze, Christian Fgen, Yue Pan, Alex Waibel |
| 2004 | ICASSP | The 2003 ISL rich transcription system for conversational telephony speech. | Hagen Soltau, Hua Yu, Florian Metze, Christian Fgen, Qin Jin, Szu-Chen Stan Jou |
| 2004 | Interspeech | Issues in meeting transcription - the ISL meeting transcription system. | Tanja Schultz, Qin Jin, Kornel Laskowski, Yue Pan, Florian Metze, Christian Fgen |
| 2003 | ICASSP | Multilingual articulatory features. | Sebastian Stker, Tanja Schultz, Florian Metze, Alex Waibel |
| 2003 | Interspeech | The NESPOLE! voIP multilingual corpora in tourism and medical domains. | Nadia Mana, Susanne Burger, Roldano Cattoni, Laurent Besacier, Victoria MacLaren, John W. McDonough, Florian Metze |
| 2003 | Interspeech | Integrating multilingual articulatory features into speech recognition. | Sebastian Stker, Florian Metze, Tanja Schultz, Alex Waibel |
| 2002 | ACL | A Multi-Perspective Evaluation of the NESPOLE! Speech-to-Speech Translation System. | Alon Lavie, Florian Metze, Roldano Cattoni, Erica Costantini |
| 2002 | ICASSP | Efficient language model lookahead through polymorphic linguistic context assignment. | Hagen Soltau, Florian Metze, Christian Fgen, Alex Waibel |
| 2002 | Interspeech | A flexible stream architecture for ASR using articulatory features. | Florian Metze, Alex Waibel |
| 2002 | Interspeech | Compensating for hyperarticulation by modeling articulatory properties. | Hagen Soltau, Florian Metze, Alex Waibel |
| 2001 | ICASSP | Speaker compensation with sine-log all-pass transforms. | John W. McDonough, Florian Metze, Hagen Soltau, Alex Waibel |
| 2001 | ICASSP | The ISL evaluation system for Verbmobil-II. | Hagen Soltau, Thomas Schaaf, Florian Metze, Alex Waibel |
| 2001 | ICASSP | Advances in automatic meeting record creation and access. | Alex Waibel, Michael Bett, Florian Metze, Klaus Ries, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu, Klaus Zechner |
| 2001 | Interspeech | The nespole! voIP dialogue database. | Susanne Burger, Laurent Besacier, Paolo Coletti, Florian Metze, Cline Morel |
| 2001 | Interspeech | Speech recognition over netmeeting connections. | Florian Metze, John W. McDonough, Hagen Soltau |
| 2001 | NAACL | Advances in meeting recognition. | Alex Waibel, Hua Yu, Tanja Schultz, Yue Pan, Michael Bett, Martin Westphal, Hagen Soltau, Thomas Schaaf, Florian Metze |
| 2000 | ICASSP | Confidence measure based language identification. | Florian Metze, Thomas Kemp, Thomas Schaaf, Tanja Schultz, Hagen Soltau |