Results 1 to 10 of about 28,309,773 (304)

Recognition of spectrally shaped speech in speech-modulated noise: Effects of age, spectral shape, speech level, and vocoding [PDF]

open access: yesJASA Express Letters, 2023
This study examined the recognition of spectrally shaped syllables and sentences in speech-modulated noise by younger and older adults. The effect of spectral shaping and speech level on temporal amplitude modulation cues was explored through speech ...
Daniel Fogerty   +2 more
doaj   +2 more sources

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models [PDF]

open access: yesNeural Information Processing Systems, 2023
In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis.
Yinghao Aaron Li   +4 more
semanticscholar   +1 more source

Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information [PDF]

open access: yesIEEE International Conference on Acoustics, Speech, and Signal Processing, 2022
Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task.
Heqing Zou   +4 more
semanticscholar   +1 more source

Changes of Voice Production in Artificial Acoustic Environments

open access: yesFrontiers in Built Environment, 2021
The physical production of speech level dynamic range is directly affected by the physiological features of the speaker such as vocal tract size and lung capacity; however, the regulation of these production systems is affected by the perception of the ...
Tomás Sierra-Polanco   +4 more
doaj   +1 more source

Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech [PDF]

open access: yesIEEE International Conference on Acoustics, Speech, and Signal Processing, 2023
Automatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and
Farhad Javanmardi   +4 more
semanticscholar   +1 more source

PHONEME-LEVEL BERT FOR ENHANCED PROSODY OF TEXT-TO-SPEECH WITH GRAPHEME PREDICTIONS [PDF]

open access: yesIEEE International Conference on Acoustics, Speech, and Signal Processing, 2023
Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-
Yinghao Aaron Li   +3 more
semanticscholar   +1 more source

Experiment on a Transformer Model Indonesian-to-Sundanese Neural Machine Translation with Sundanese Speech Level Evaluation

open access: yesProceedings of the Thirteenth Conference on Applied Linguistics (CONAPLIN 2020), 2021
Speech level is one of the essential Sundanese language elements. As Indonesian mixed within Sundanese language use, the usage of speech level is gradually degrading.
R. B. Primandhika   +2 more
semanticscholar   +1 more source

The Changing of ‘Sor Singgih Basa’ in Balinese Root Based on the Internal Modification: Morpho-Phonology Study

open access: yesJournal of Language and Literature, 2023
This research investigates the relationship between phonology and morphology in influencing the changing Balinese speech level, namely ‘singgih’ (high) and ‘sor’ (low). The analysis focuses on utilizing the internal modification of a formal morphological
I Gusti Ayu Sundari Okasunu   +2 more
doaj   +1 more source

SAMU-XLSR: Semantically-Aligned Multimodal Utterance-Level Cross-Lingual Speech Representation [PDF]

open access: yesIEEE Journal on Selected Topics in Signal Processing, 2022
We propose the ($\tt SAMU\text{-}XLSR$): Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework. Unlike previous works on speech representation learning, which learns multilingual contextual speech ...
Sameer Khurana   +2 more
semanticscholar   +1 more source

Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization [PDF]

open access: yesIEEE International Conference on Acoustics, Speech, and Signal Processing, 2022
Code-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and ...
Hexin Liu   +5 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy