Results 131 to 140 of about 1,468,270 (262)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
Accepted in EMNLP ...
MinJu Jeon +4 more
openaire +4 more sources
UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos
Abstract Extracting actionable insights from long‐duration urban videos is often labor‐intensive: analysts must manually sift through raw footage to pinpoint target events or uncover broader behavioral trends. In this work, we present UrbanClipAtlas, a visual analytics system for exploring long urban videos recorded at street intersections ...
Joel Perca +5 more
wiley +1 more source
Local feature‐based video captioning with multiple classifier and CARU‐attention
Video captioning aims to identify multiple objects and their behaviours in a video event and generate captions for the current scene. This task aims to generate a detailed description of the current video in real‐time using natural language, which ...
Sio‐Kei Im, Ka‐Hou Chan
doaj +1 more source
IntroductionAccess to non-speech information (NSI) in video content is essential to creating accessible and engaging video content, particularly for D/deaf and Hard-of-Hearing (DHH) audiences.
Lloyd May +8 more
doaj +1 more source
Single Line Drawing Generation via Semantics‐Driven Optimization
We present a method for automatically generating single‐line drawings in vector format, guided by a text prompt or an input image. Our approach leverages score distillation sampling to optimize the parameters of a uniform rational B‐spline (URBS) curve, ensuring that the drawing consists of a single continuous stroke by design.
Tanguy Magne +3 more
wiley +1 more source
Advances in 4D Representation: Geometry, Motion, and Interaction
We survey 4D representation through three key pillars — geometry, motion, and interaction — offering a selective, representation‐centric perspective to guide researchers in choosing and customizing the right 4D representation for their tasks. Abstract We present a survey on 4D generation and reconstruction, a fast‐evolving subfield of computer graphics
M. Zhao +7 more
wiley +1 more source
Bilingual video captioning model for enhanced video retrieval
Many video platforms rely on the descriptions that uploaders provide for video retrieval. However, this reliance may cause inaccuracies. Although deep learning-based video captioning can resolve this problem, it has some limitations: (1) traditional ...
Norah Alrebdi, Amal A. Al-Shargabi
doaj +1 more source
The Evidence Base of Fluoride Videos on TikTok
ABSTRACT Objective The fluoride controversy has intensified in recent years, partly due to new research suggesting negative health effects. Simultaneously, TikTok has become a prominent platform for disseminating oral health‐related content, but its accuracy remains uncertain.
Annel Lueth, Lance Brendan Young
wiley +1 more source
Survey of Dense Video Captioning: Techniques, Resources, and Future Perspectives
Dense Video Captioning (DVC) represents the cutting edge of advanced multimedia tasks, focusing on generating a series of temporally precise descriptions for events unfolding within a video.
Zhandong Liu, Ruixia Song
doaj +1 more source
Pre‐Task Explicit Instruction, Input Modality, and Working Memory in L2 Oral Self‐Repair
ABSTRACT Despite the central role of tasks in language education and ensuing research documenting how task‐related variables might affect language performance and learning, it remains unclear whether pre‐task explicit instruction, input modality, and working memory (WM) influence how learners monitor and repair grammatical structures in real‐time ...
Reza Yadollahpour +2 more
wiley +1 more source

