Multimodal generative AI for interpreting 3D medical images and videos. [PDF]
Lee JO +4 more
europepmc +1 more source
Human-like scene graph generation and evaluation. [PDF]
Milewski V, Moens MF, Trusca MM.
europepmc +1 more source
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment. [PDF]
Kulkarni Y, Fazli P.
europepmc +1 more source
Industrial Object Counting from Traditional Machine Vision to Open-World Foundation Models: A Systematic Review. [PDF]
Wang W +5 more
europepmc +1 more source
Training strategies for semi-supervised remote sensing image captioning. [PDF]
Cheng Q +5 more
europepmc +1 more source
Adaptive graph signal processing for robust multimodal fusion with dynamic semantic alignment. [PDF]
Karthikeya KV +4 more
europepmc +1 more source
Contrastive language image pretraining for a cardiac magnetic resonance image embedding with zero-shot capabilities. [PDF]
Nakashima M +10 more
europepmc +1 more source
Lightweight detection network based on bidirectional weighted feature fusion with small target enhancement for USVs. [PDF]
Li Y, Lian D, Du J, Bu C, Gao D.
europepmc +1 more source
Multi-Spectral Band Analysis for Satellite-to-Aerial Image Registration: A Comparative Study of Deep Learning and Traditional Feature-Matching Methods. [PDF]
Han D, Song JH, Lee SG.
europepmc +1 more source
Insights into Object Semantics: Leveraging Transformer Networks for Advanced Image Captioning. [PDF]
Abdal Hafeth D, Kollias S.
europepmc +1 more source

