Skip to content

Hisashi Kawai

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

121

Venues

9

Active years

1986–2025

Best venue rank

A*

Where they publish

Papers

121 indexed papers, newest first.

YearVenueTitleAuthors
2025ASRULayer-wise Analysis for Quality of Multilingual Synthesized Speech.Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
2025ASRUVoice Factor Control Using FIR-Based Fast Neural Vocoder for Speech Generation Applications.Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai
2025ICASSPMora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels.Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai
2025InterspeechCross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR.Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai
2025InterspeechGST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens.Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai
2024ICASSPHierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASR.Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai
2024ICASSPFIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter.Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai
2024ICASSPConvnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion.Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
2024InterspeechInvestigating ASR Error Correction with Large Language Model and Multilingual 1-best Hypotheses.Sheng Li, Chen Chen, Kwok Chin Yuen, Chenhui Chu, Eng Siong Chng, Hisashi Kawai
2024InterspeechMobile PresenTra: NICT fast neural text-to-speech system on smartphones with incremental inference of MS-FC-HiFi-GAN for law-latency synthesis.Takuma Okamoto, Yamato Ohtani, Hisashi Kawai
2024InterspeechChallenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder.Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai
2023ASRUCross-Modal Alignment With Optimal Transport For CTC-Based ASR.Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai
2023ASRUWaveNeXt: ConvNeXt-Based Fast Neural Vocoder Without ISTFT layer.Takuma Okamoto, Haruki Yamashita, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
2023ASRUGenerative Linguistic Representation for Spoken Language Identification.Peng Shen, Xuguang Lu, Hisashi Kawai
2023InterspeechE2E-S2S-VC: End-To-End Sequence-To-Sequence Voice Conversion.Takuma Okamoto, Tomoki Toda, Hisashi Kawai
2023SMCHomeostatic System Design Based on Understanding the Living Environmental Determinants of Falls.Mikiko Oono, Ayano Nomura, Koji Kitamura, Yoshifumi Nishida, Shunsaburo Nakahara, Hisashi Kawai
2022InterspeechTransducer-based language embedding for spoken language identification.Peng Shen, Xugang Lu, Hisashi Kawai
2021ASRUMulti-Stream HiFi-GAN with Data-Driven Waveform Decomposition.Takuma Okamoto, Tomoki Toda, Hisashi Kawai
2021ICASSPUnsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification.Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai
2021ICASSPHigh-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC.Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2021ICASSPNoise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders.Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2021InterspeechNoise Robust Acoustic Modeling for Single-Channel Speech Recognition Based on a Stream-Wise Transformer Architecture.Masakiyo Fujimoto, Hisashi Kawai
2020ICASSPTransformer-Based Text-to-Speech with Weighted Forced Attention.Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2020InterspeechInvestigation of NICT Submission for Short-Duration Speaker Verification Challenge 2020.Peng Shen, Xugang Lu, Hisashi Kawai
2020InterspeechQuasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech Generation.Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda
2019ASRUTacotron-Based Acoustic Model Using Phoneme Alignment for Practical Neural Text-to-Speech Systems.Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2019CoRLMultimodal Attention Branch Network for Perspective-Free Sentence Generation.Aly Magassouba, Komei Sugiura, Hisashi Kawai
2019ICASSPInvestigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features.Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2019ICASSPInteractive Learning of Teacher-student Model for Short Utterance Spoken Language Identification.Peng Shen, Xugang Lu, Sheng Li, Hisashi Kawai
2019ICASSPInvestigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models.Ryoichi Takashima, Sheng Li, Hisashi Kawai
2019InterspeechEnd-to-End Articulatory Attribute Modeling for Low-Resource Multilingual Speech Recognition.Sheng Li, Chenchen Ding, Xugang Lu, Peng Shen, Tatsuya Kawahara, Hisashi Kawai
2019InterspeechOne-Pass Single-Channel Noisy Speech Recognition Using a Combination of Noisy and Enhanced Features.Masakiyo Fujimoto, Hisashi Kawai
2019InterspeechIncorporating Symbolic Sequential Modeling for Speech Enhancement.Chien-Feng Liao, Yu Tsao, Xugang Lu, Hisashi Kawai
2019InterspeechInvestigating Radical-Based End-to-End Speech Recognition Systems for Chinese Dialects and Japanese.Sheng Li, Xugang Lu, Chenchen Ding, Peng Shen, Tatsuya Kawahara, Hisashi Kawai
2019InterspeechImproving Transformer-Based Speech Recognition Systems with Compressed Structure and Speech Attributes Augmentation.Sheng Li, Raj Dabre, Xugang Lu, Peng Shen, Tatsuya Kawahara, Hisashi Kawai
2019InterspeechClass-Wise Centroid Distance Metric Learning for Acoustic Event Detection.Xugang Lu, Peng Shen, Sheng Li, Yu Tsao, Hisashi Kawai
2019InterspeechDuration Modeling with Global Phoneme-Duration Vectors.Jinfu Ni, Yoshinori Shiga, Hisashi Kawai
2019InterspeechReal-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders.Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2018ICASSPComparative Evaluations of Various Factored Deep Convolutional Rnn Architectures for Noise Robust Speech Recognition.Masakiyo Fujimoto, Hisashi Kawai
2018ICASSPAn Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features.Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2018ICASSPAn Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation.Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2018ICASSPAn Investigation of a Knowledge Distillation Method for CTC Acoustic Models.Ryoichi Takashima, Sheng Li, Hisashi Kawai
2018ICASSPCTC Loss Function with a Unit-Level Ambiguity Penalty.Ryoichi Takashima, Sheng Li, Hisashi Kawai
2018InterspeechImproving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks.Sheng Li, Xugang Lu, Ryoichi Takashima, Peng Shen, Tatsuya Kawahara, Hisashi Kawai
2018InterspeechTemporal Attentive Pooling for Acoustic Event Detection.Xugang Lu, Peng Shen, Sheng Li, Yu Tsao, Hisashi Kawai
2018InterspeechMultilingual Grapheme-to-Phoneme Conversion with Global Character Vectors.Jinfu Ni, Yoshinori Shiga, Hisashi Kawai
2018InterspeechFeature Representation of Short Utterances Based on Knowledge Distillation for Spoken Language Identification.Peng Shen, Xugang Lu, Sheng Li, Hisashi Kawai
2017ASRUIncremental training and constructing the very deep convolutional residual network acoustic models.Sheng Li, Xugang Lu, Peng Shen, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai
2017ASRUSubband wavenet with overlapped single-sideband filterbanks.Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2017ASRUGrounded language understanding for manipulation instructions using GAN-based classification.Komei Sugiura, Hisashi Kawai
2017ICASSPMinimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework.Naoyuki Kanda, Xugang Lu, Hisashi Kawai
2017InterspeechGlobal Syllable Vectors for Building TTS Front-End with Deep Learning.Jinfu Ni, Yoshinori Shiga, Hisashi Kawai
2017InterspeechConditional Generative Adversarial Nets Classifier for Spoken Language Identification.Peng Shen, Xugang Lu, Sheng Li, Hisashi Kawai
2016ICASSPBottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizer.Tsubasa Ochiai, Shigeki Matsuda, Hideyuki Watanabe, Xugang Lu, Hisashi Kawai, Shigeru Katagiri
2016ICASSPLocal fisher discriminant analysis for spoken language identification.Peng Shen, Xugang Lu, Lemao Liu, Hisashi Kawai
2016InterspeechInvestigation of Semi-Supervised Acoustic Model Training Based on the Committee of Heterogeneous Neural Networks.Naoyuki Kanda, Shoji Harada, Xugang Lu, Hisashi Kawai
2016InterspeechMaximum a posteriori Based Decoding for CTC Acoustic Models.Naoyuki Kanda, Xugang Lu, Hisashi Kawai
2016InterspeechPair-Wise Distance Metric Learning of Neural Network Model for Spoken Language Identification.Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai
2016InterspeechUsing Zero-Frequency Resonator to Extract Multilingual Intonation Structure.Jinfu Ni, Yoshinori Shiga, Hisashi Kawai
2016InterspeechModel Integration for HMM- and DNN-Based Speech Synthesis Using Product-of-Experts Framework.Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
2016InterspeechFXiaoyun Wang, Xugang Lu, Hisashi Kawai, Seiichi Yamamoto
2015ASRUTraining data pseudo-shuffling and direct decoding framework for recurrent neural network based acoustic modeling.Naoyuki Kanda, Mitsuyoshi Tachimori, Xugang Lu, Hisashi Kawai
2015InterspeechSparse representation with temporal max-smoothing for acoustic event detection.Xugang Lu, Peng Shen, Yu Tsao, Chiori Hori, Hisashi Kawai
2015InterspeechHMM based myanmar text to speech system.Ye Kyaw Thu, Win Pa Pa, Jinfu Ni, Yoshinori Shiga, Andrew M. Finch, Chiori Hori, Hisashi Kawai, Eiichiro Sumita
2014ICRANon-monologue HMM-based speech synthesis for service robots: A cloud robotics approach.Komei Sugiura, Yoshinori Shiga, Hisashi Kawai, Teruhisa Misu, Chiori Hori
2013MDMMultilingual Speech-to-Speech Translation System: VoiceTra.Shigeki Matsuda, Xinhui Hu, Yoshinori Shiga, Hideki Kashioka, Chiori Hori, Keiji Yasuda, Hideo Okuma, Masao Uchiyama, Eiichiro Sumita, Hisashi Kawai, Satoshi Nakamura
2012InterspeechAn Evaluation of Parameter Generation Methods with Rich Context Models in HMM-Based Speech Synthesis.Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai, Sakriani Sakti, Satoshi Nakamura
2011ICASSPUnsupervised determination of efficient Korean LVCSR units using a Bayesian Dirichlet process model.Sakriani Sakti, Andrew M. Finch, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2011ICASSPIncreasing discriminative capability on MAP-based mapping function estimation for acoustic model adaptation.Yu Tsao, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2011ICASSPA sampling-based environment population projection approach for rapid acoustic model adaptation.Yu Tsao, Shigeki Matsuda, Shinsuke Sakai, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2011IJCNLPImproving Related Entity Finding via Incorporating Homepages and Recognizing Fine-grained Entities.Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka
2011IJCNLPAnswering Complex Questions via Exploiting Social Q&A Collection.Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka
2011InterspeechSpeaker-Adaptive Speech Synthesis Based on Eigenvoice Conversion and Language-Dependent Prosodic Conversion in Speech-to-Speech Translation.Nobuhiko Hattori, Tomoki Toda, Hisashi Kawai, Hiroshi Saruwatari, Kiyohiro Shikano
2011InterspeechAdaptive Regularization Framework for Robust Voice Activity Detection.Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2011InterspeechUser Study of Spoken Decision Support System.Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hisashi Kawai, Satoshi Nakamura
2011InterspeechIncorporating Regional Information to Enhance MAP-Based Stochastic Feature Compensation for Robust Speech Recognition.Yu Tsao, Paul R. Dixon, Chiori Hori, Hisashi Kawai
2011InterspeechEstimation of Perceptual Spaces for Speaker Identities Based on the Cross-Lingual Discrimination Task.Minoru Tsuzaki, Keiichi Tokuda, Hisashi Kawai, Jinfu Ni
2011SIGdialToward Construction of Spoken Dialogue System that Evokes Users' Spontaneous Backchannels.Teruhisa Misu, Etsuo Mizukami, Yoshinori Shiga, Shinichi Kawamoto, Hisashi Kawai, Satoshi Nakamura
2010InterspeechBrazilian portuguese acoustic model training based on data borrowing from other language.Kazuhiko Abe, Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2010InterspeechCluster-based language model for spoken document retrieval using NMF-based document clustering.Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2010InterspeechConstruction and evaluations of an annotated Chinese conversational corpus in travel domain for the language model of speech recognition.Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2010InterspeechVoice activity detection in a reguarized reproducing kernel hilbert space.Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2010InterspeechAn unsupervised approach to creating web audio contents-based HMM voices.Jinfu Ni, Hisashi Kawai
2010InterspeechUtilizing a noisy-channel approach for Korean LVCSR.Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura
2010InterspeechImproved training of excitation for HMM-based parametric speech synthesis.Yoshinori Shiga, Tomoki Toda, Shinsuke Sakai, Hisashi Kawai
2010SIGdialModeling Spoken Decision Making Dialogue and Optimization of its Dialogue Strategy.Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura
2009InterspeechA close look into the probabilistic concatenation model for corpus-based speech synthesis.Shinsuke Sakai, Ranniery Maia, Hisashi Kawai, Satoshi Nakamura
2008ICASSPUnit database pruning based on the cost degradation criterion for concatenative speech synthesis.Nobuyuki Nishizawa, Hisashi Kawai
2007InterspeechA preselection method based on cost degradation from the optimal sequence for concatenative speech synthesis.Nobuyuki Nishizawa, Hisashi Kawai
2006ICASSPConstructing a Phonetic-Rich Speech Corpus While Controlling Time-Dependent Voice Quality Variability for English Speech Synthesis.Jinfu Ni, Toshio Hirai, Hisashi Kawai
2006ICASSPA Short-Latency Unit Selection Method with Redundant Search for Concatenative Speech Synthesis.Nobuyuki Nishizawa, Hisashi Kawai
2006InterspeechQuick individual fitting methods of simplified hearing compensation for elderly people.Kengo Fujita, Tsuneo Kato, Hisashi Kawai
2006InterspeechA text-prompted distributed speaker verification system implemented on a cellular phone and a mobile terminal.Tsuneo Kato, Hisashi Kawai
2005InterspeechSNR-dependent background noise compensation of PESQ values for cellular phone speech.Kengo Fujita, Tsuneo Kato, Hideaki Yamada, Hisashi Kawai
2005InterspeechAnalysis of major factors of naturalness degradation in concatenative synthesis.Toshio Hirai, Hisashi Kawai, Minoru Tsuzaki, Nobuyuki Nishizawa
2005InterspeechEstimation of intonation variation with constrained tone transformations.Jinfu Ni, Hisashi Kawai, Keikichi Hirose
2005InterspeechImprovement of rejection performance of keyword spotting using anti-keywords derived from large vocabulary considering acoustical similarity to keywords.Makoto Yamada, Tsuneo Kato, Masaki Naito, Hisashi Kawai
2004ICASSPAn evaluation of automatic phone segmentation for concatenative speech synthesis.Hisashi Kawai, Tomoki Toda
2004ICASSPScaling of waveform segments along the time axis for concatenative speech synthesis.Nobuyuki Nishizawa, Hisashi Kawai
2004ICASSPOptimizing sub-cost functions for segment selection based on perceptual evaluations in concatenative speech synthesis.Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki
2004ICASSPMinimum segmentation error based discriminative training for speech synthesis application.Yi-Jian Wu, Hisashi Kawai, Jinfu Ni, Ren-Hua Wang
2004InterspeechFormulating contextual tonal variations in Mandarin.Jinfu Ni, Hisashi Kawai, Keikichi Hirose
2004InterspeechUsing a depth-restricted search to reduce delays in unit selection.Nobuyuki Nishizawa, Hisashi Kawai
2004InterspeechA study on automatic detection of Japanese vowel devoicing for speech synthesis.Yi-Jian Wu, Hisashi Kawai, Jinfu Ni, Ren-Hua Wang
2003ICASSPTone feature extraction through parametric modeling and analysis-by-synthesis-based pattern matching.Jinfu Ni, Hisashi Kawai
2003ICASSPSegment selection considering local degradation of naturalness in concatenative speech synthesis.Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano
2003InterspeechTone pattern discrimination combining parametric modeling and maximum likelihood estimation.Jinfu Ni, Hisashi Kawai
2003InterspeechOptimizing integrated cost function for segment selection in concatenative speech synthesis based on perceptual evaluations.Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki
2002ICASSPUnit selection algorithm for Japanese speech synthesis based on both phoneme unit and diphone unit.Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano
2002InterspeechAcoustic measures vs. phonetic features as predictors of audible discontinuity in concatenative speech synthesis.Hisashi Kawai, Minoru Tsuzaki
2002InterspeechPerceptual evaluation of naturalness due to substitution of Chinese syllable for concatenative speech synthesis.Jinlin Lu, Hisashi Kawai
2002InterspeechDesign of a Mandarin sentence set for corpus-based speech synthesis by use of a multi-tier algorithm taking account of the varied prosodic and spectral characteristics.Jinfu Ni, Hisashi Kawai
2002InterspeechFeature extraction for unit selection in concatenative speech synthesis: comparison between AIM, LPC, and MFCC.Minoru Tsuzaki, Hisashi Kawai
2000InterspeechA design method of speech corpus for text-to-speech synthesis taking account of prosody.Hisashi Kawai, Seiichi Yamamoto, Norio Higuchi, Tohru Shimizu
1998InterspeechRecognition of connected digit speech in Japanese collected over the telephone network.Hisashi Kawai, Norio Higuchi
1994ICASSPDevelopment of a text-to-speech system for Japanese based on waveform splicing.Hisashi Kawai, Norio Higuchi, Tohru Shimizu, Seiichi Yamamoto
1990ICASSPA system for synthesizing Japanese speech from orthographic text.Hiroya Fujisaki, Keikichi Hirose, Hisashi Kawai, Yasuharu Asano
1990InterspeechImprovement of the synthetic speech quality of the formant-type speech synthesizer and its subjective evaluation.Norio Higuchi, Hisashi Kawai, Tohru Shimizu, Seiichi Yamamoto
1990InterspeechThe linguistic processing module for Japanese text-to-speech system.Tohru Shimizu, Norio Higuchi, Hisashi Kawai, Seiichi Yamamoto
1988ICASSPRealization of linguistic information in the voice fundamental frequency contour of the spoken Japanese.Hiroya Fujisaki, Hisashi Kawai
1986ICASSPGeneration of prosodic symbols for rule-synthesis of connected speech of Japanese.Keikichi Hirose, Hiroya Fujisaki, Hisashi Kawai