Learning Semantic Features for Dense Video Captioning
Sujin Lee, Incheol Kim
openaire +1 more source
Refer-ASV: Referring Multi-Object Tracking in Autonomous Surface Vehicle Navigation Scenes. [PDF]
Xue B +5 more
europepmc +1 more source
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding. [PDF]
Tang F +20 more
europepmc +1 more source
Will multimodal large language models ever achieve deep understanding of the world? [PDF]
Farkaš I, Vavrečka M, Wermter S.
europepmc +1 more source
Enhancing interior design and space planning via human-machine intelligent interaction for artistic cognition. [PDF]
Jiang J.
europepmc +1 more source
Interpol review of detection of AI-generated image and video deepfakes, 2022-2025. [PDF]
van Lierop S +3 more
europepmc +1 more source
Confidence-Guided Fusion for Self-Supervised Monocular Depth Estimation in Endoscopy. [PDF]
Li S, Wang H, Hu Z, Chu T, Li Y, Zhao L.
europepmc +1 more source
EchoNet++: A multilingual soccer match audio commentary dataset. [PDF]
Majeed F, Nazir M, Agus M, Schneider J.
europepmc +1 more source
Social robot navigation: a review and benchmarking of learning-based methods. [PDF]
Alyassi R +3 more
europepmc +1 more source
Enhancing Surveillance Systems: Integration of Object, Behavior, and Space Information in Captions for Advanced Risk Assessment. [PDF]
Jeon M, Ko J, Cheoi K.
europepmc +1 more source

