Results 41 to 50 of about 36,891 (314)
Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach
Recent advancements in end-to-end (E2E) modeling have notably improved automatic speech recognition (ASR) systems; however, far-field speech recognition (FSR) remains challenging due to signal degradation from factors such as low signal-to-noise ratio ...
Tong Niu, Yaqi Chen, Dan Qu, Hengbo Hu
doaj +1 more source
The objectives of the research that the researchers conducted were to find out how the ability of Maharah Istima' was for Class II students at Madrasah Diniyah Pondok Pesantren Sunan Giri Surabaya, to find out how Maharah Istima' was able to play through
Faruq Abdul Muid
doaj +1 more source
Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models
A multimodal human–robot interaction framework integrates human pose estimation (HPE) and a large language model (LLM) for gesture‐ and voice‐based robot control. Speech‐to‐text (STT) enables voice command interpretation, while a safety‐aware arbitration mechanism prioritizes gesture input for rapid intervention.
Nasiru Aboki +2 more
wiley +1 more source
To the Editor: There is an interesting game called “telephone” or “whispers,” in which a message is passed on, in a whisper, down a line of people, and then the last person speaks the message out loud. The final version of the message is usually radically changed from the original.
openaire +2 more sources
LLM‐Integrated Human–Robot Interaction System for Microrobots
This paper proposes an LLM‐based control framework for guiding microrobots using human natural language. This framework can convert the natural human speech into safe and executable command sets for reliable navigation in complex environments. The experimental results show high accuracy and robustness in task performance, demonstrating the potential of
Bairong Zhu, Amar Salehi, Tingting Yu
wiley +1 more source
Exploration of Whisper fine-tuning strategies for low-resource ASR
Limited data availability remains a significant challenge for Whisper’s low-resource speech recognition performance, falling short of practical application requirements.
Yunpeng Liu, Xukui Yang, Dan Qu
semanticscholar +1 more source
Green and scalable synthesis of CsPbBr3 perovskite quantum dots using bio‐sourced lecithin and hexane enables single‐batch production over 10 grams. Suppressed Auger recombination and electron‐phonon coupling yields 95% PLQY, enhanced stability, and low ASE thresholds of 230.4 (ns) and 50.1 (fs) µJ cm−2 in neat films. ABSTRACT Halide perovskite quantum
Yongfeng Liu +10 more
wiley +1 more source
Symmetry‐Imposed Selection Rules for Excitations of Nontrivial Plasmonic Topologies
A unified group‐theory selection rule governs the excitation of vectorial nearfield topologies across three plasmonic spin states. Derived from first principles, it predicts spin–orbit vortex splitting and multidimensional nested vortices, confirmed by phase‐resolved in situ measurements.
Jie Yang +14 more
wiley +1 more source
BackgroundHypernasality, a hallmark of velopharyngeal insufficiency (VPI), is a speech disorder with significant psychosocial and functional implications.
Myranda Uselton Shirk +14 more
doaj +1 more source
Patrick Modiano‟s Voice: from La Place de l’Etoile to Dora Bruder [PDF]
The way an author individualizes his writing is expressed through voice, a feature of writing that is often overlooked, generally not analyzed. The phenomenon of voice is not easy to grasp and when we think of Patrick Modiano, an adjective comes to ...
Ruth AMAR
doaj

