Recognition of spectrally shaped speech in speech-modulated noise: Effects of age, spectral shape, speech level, and vocoding [PDF]
This study examined the recognition of spectrally shaped syllables and sentences in speech-modulated noise by younger and older adults. The effect of spectral shaping and speech level on temporal amplitude modulation cues was explored through speech ...
Daniel Fogerty +2 more
doaj +2 more sources
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models [PDF]
In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis.
Yinghao Aaron Li +4 more
semanticscholar +1 more source
Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information [PDF]
Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task.
Heqing Zou +4 more
semanticscholar +1 more source
Changes of Voice Production in Artificial Acoustic Environments
The physical production of speech level dynamic range is directly affected by the physiological features of the speaker such as vocal tract size and lung capacity; however, the regulation of these production systems is affected by the perception of the ...
Tomás Sierra-Polanco +4 more
doaj +1 more source
Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech [PDF]
Automatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and
Farhad Javanmardi +4 more
semanticscholar +1 more source
PHONEME-LEVEL BERT FOR ENHANCED PROSODY OF TEXT-TO-SPEECH WITH GRAPHEME PREDICTIONS [PDF]
Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-
Yinghao Aaron Li +3 more
semanticscholar +1 more source
Speech level is one of the essential Sundanese language elements. As Indonesian mixed within Sundanese language use, the usage of speech level is gradually degrading.
R. B. Primandhika +2 more
semanticscholar +1 more source
This research investigates the relationship between phonology and morphology in influencing the changing Balinese speech level, namely ‘singgih’ (high) and ‘sor’ (low). The analysis focuses on utilizing the internal modification of a formal morphological
I Gusti Ayu Sundari Okasunu +2 more
doaj +1 more source
SAMU-XLSR: Semantically-Aligned Multimodal Utterance-Level Cross-Lingual Speech Representation [PDF]
We propose the ($\tt SAMU\text{-}XLSR$): Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework. Unlike previous works on speech representation learning, which learns multilingual contextual speech ...
Sameer Khurana +2 more
semanticscholar +1 more source
Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization [PDF]
Code-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and ...
Hexin Liu +5 more
semanticscholar +1 more source

