TUMAČENJA ENGLESKOG PREZENT PERFEKTA: GLAGOLSKO VREME, FAZA, GLAGOLSKI VID
2022Because of the complexity of its meanings and uses, different linguists within their theoretical approaches discussed English present perfect as tense, phase or verbal aspect, thus giving priority to one of the components of its usage over the others.
openaire +1 more source
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
Computer Vision and Pattern RecognitionExisting video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability.
Weixi Feng +5 more
semanticscholar +1 more source
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
Neural Information Processing SystemsMultimodal large language models (MLLMs) demonstrate remarkable capabilities in handling complex multimodal tasks and are increasingly adopted in video understanding applications.
Qi Li, Runpeng Yu, Xinchao Wang
semanticscholar +1 more source
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs
arXiv.orgVideo Large Language Models (VLLMs) excel in video understanding, but their excessive visual tokens pose a significant computational challenge for real-world applications.
Junpeng Ma +6 more
semanticscholar +1 more source
RealCam-Vid: High-resolution Video Dataset with Dynamic Scenes and Metric-scale Camera Movements
arXiv.orgRecent advances in camera-controllable video generation have been constrained by the reliance on static-scene datasets with relative-scale camera annotations, such as RealEstate10K.
Guangcong Zheng +3 more
semanticscholar +1 more source
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
arXiv.orgWe present xGen-MM-Vid (BLIP-3-Video): a multimodal language model for videos, particularly designed to efficiently capture temporal information over multiple frames.
Michael S. Ryoo +9 more
semanticscholar +1 more source
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
arXiv.orgWe introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for
Junyoung Seo +11 more
semanticscholar +1 more source
Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
arXiv.orgThe rapid progress of image-to-video (I2V) generation models has introduced significant risks by enabling deceptive or malicious video synthesis from a single image.
Rohit Chowdhury +3 more
semanticscholar +1 more source
FWAF-VID: A Flapping-Wing Aggressive Flight Benchmark Dataset for Visual-Inertial Localization
IEEE Robotics and Automation LettersAccurate state estimation of micro aerial vehicles (MAVs) in high-speed and dynamic environments poses a significant challenge for visual-inertial odometry (VIO) algorithms.
Jizhou Jiang +4 more
semanticscholar +1 more source
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
arXiv.orgRecent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge.
Haijier Chen +5 more
semanticscholar +1 more source

