Results 241 to 250 of about 107,975 (302)
Some of the next articles are maybe not open access.

Multimodaltrace: Deepfake Detection using Audiovisual Representation Learning

2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2023
By employing generative deep learning techniques, Deepfakes are created with the intent to create mistrust in society, manipulate public opinion and political decisions, and for other malicious purposes such as blackmail, scamming, and even cyberstalking.
Muhammad Anas Raza, K. Malik
semanticscholar   +1 more source

Machine learning based approaches for clinical and non-clinical depression recognition and depression relapse prediction using audiovisual and EEG modalities: A comprehensive review

Comput. Biol. Medicine, 2023
Mental disorders are rapidly increasing each year and have become a major challenge affecting the social and financial well-being of individuals. There is a need for phenotypic characterization of psychiatric disorders with biomarkers to provide a rich ...
Sana Yasin   +3 more
semanticscholar   +1 more source

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

Neural Information Processing Systems, 2020
Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines.
Di Hu   +7 more
semanticscholar   +1 more source

Differential audiovisual information processing in emotion recognition: An eye-tracking study.

Emotion, 2022
Recent research has suggested that dynamic emotion recognition involves strong audiovisual association; that is, facial or vocal information alone automatically induces perceptual processes in the other modality.
Yueyuan Zheng, J. Hsiao
semanticscholar   +1 more source

Synthesis in the Audiovisual

Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, 2016
The S T R A T I C audiovisual project is based on the phenomenon that occurs when filming a pulsating light -- lines appear on the screen. The thickness, color and movement of these lines are directly related to the frequency of the sound. In other words, the sound generates the visuals in real-time. The visuals are examined by the use of shutter speed
Vygandas 'Vegas' Simbelis   +1 more
openaire   +1 more source

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

arXiv.org
Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual
Xinlong Chen   +11 more
semanticscholar   +1 more source

Aligning Audiovisual Features for Audiovisual Speech Recognition

2018 IEEE International Conference on Multimedia and Expo (ICME), 2018
Visual information can improve the performance of automatic speech recognition (ASR), especially in the presence of background noise or different speech modes. A key problem is how to fuse the acoustic and visual features leveraging their complementary information and overcoming the alignment differences between modalities.
Fei Tao 0003, Carlos Busso
openaire   +1 more source

Transformer-Based Spiking Neural Networks for Multimodal Audiovisual Classification

IEEE Transactions on Cognitive and Developmental Systems
The spiking neural networks (SNNs), as brain-inspired neural networks, have received noteworthy attention due to their advantages of low power consumption, high parallelism, and high fault tolerance.
Lingyue Guo   +6 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy