Results 261 to 270 of about 28,309,773 (304)

Impact of Amazonian dance on speech performance in people with Parkinson's disease. [PDF]

open access: yesiScience
Prates RA   +9 more
europepmc   +1 more source

Tutorial and Guidelines on Measurement of Sound Pressure Level in Voice and Speech.

open access: yesJournal of Speech, Language and Hearing Research, 2018
Purpose Sound pressure level (SPL) measurement of voice and speech is often considered a trivial matter, but the measured levels are often reported incorrectly or incompletely, making them difficult to compare among various ...
J. Švec, S. Granqvist
semanticscholar   +2 more sources
Some of the next articles are maybe not open access.

Related searches:

NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality

IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022
Text-to-speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality, how to define/judge that quality, and how to achieve it.
Xu Tan   +13 more
semanticscholar   +1 more source

Residual Neural Network precisely quantifies dysarthria severity-level based on short-duration speech segments

Neural Networks, 2021
Recently, we have witnessed Deep Learning methodologies gaining significant attention for severity-based classification of dysarthric speech. Detecting dysarthria, quantifying its severity, are of paramount importance in various real-life applications ...
Siddhant Gupta   +6 more
semanticscholar   +1 more source

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Annual Meeting of the Association for Computational Linguistics, 2023
We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss during pre-training.
Ziyang Ma   +6 more
semanticscholar   +1 more source

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Neural Information Processing Systems
Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction.
Chaoyou Fu   +15 more
semanticscholar   +1 more source

FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

arXiv.org
This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications.
Haohan Guo   +8 more
semanticscholar   +1 more source

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

International Conference on Machine Learning
Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users.
Xiong Wang   +8 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy