SynTrackThinking improves multimodal multi-object tracking for autonomous driving through frequency-aware fusion and temporal contrastive learning. [PDF]
Su Z, Chang W.
europepmc +1 more source
A comparative study of vision-language models for food ingredient recognition and nutrient estimation. [PDF]
Wang S, Sheng G, Yan H, Min W, Jiang S.
europepmc +1 more source
CIRS: A Multi-Agent Machine Learning Framework for Real-Time Accident Detection and Emergency Response. [PDF]
Ayesha S, Aslam A, Zaheer MH, Khan MB.
europepmc +1 more source
Multimodal Large Language Models in Medical Imaging: Current State and Future Directions. [PDF]
Nam Y +14 more
europepmc +1 more source
The Structural Similarity Can Identify the Presence of Noise in Video Data from Unmanned Vehicles. [PDF]
Orazaev A +3 more
europepmc +1 more source
Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users. [PDF]
Zeraati FZ +4 more
europepmc +1 more source
Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. [PDF]
Wu C +5 more
europepmc +1 more source
Scaling up biomedical vision-language models: Fine-tuning, instruction tuning, and multi-modal learning. [PDF]
Peng C +5 more
europepmc +1 more source
Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models. [PDF]
Khanal B +9 more
europepmc +1 more source
An Egocentric Life-Saving Interventional Procedure Dataset of Actions, Medical Questions, Maneuvers and Tools. [PDF]
Zhuo Y +16 more
europepmc +1 more source

