Results 121 to 130 of about 1,468,270 (262)
Abstract Lecture capture is ubiquitous in higher education. Lecture capture recordings are typically accompanied by automatically generated closed captions that are sometimes corrected by humans. Students self‐report that they benefit from captions, and particularly human‐corrected captions.
Peter J. Allen +4 more
wiley +1 more source
No evidence that same‐language subtitles improve children's reading fluency
Abstract High‐profile campaigns globally have argued that same‐language television subtitles may help children improve their reading. In this intervention study, we tested the causal hypothesis that exposure to subtitles improves children's reading fluency.
Anastasiya Lopukhina +4 more
wiley +1 more source
Integrated Bayesian-Bidirectional attention network for advanced contextual video captioning
Automatic video description integrates visual and audio analysis to generate written summaries or captions, crucial for enhancing accessibility and user engagement. However, ensuring accurate and meaningful natural language descriptions remains a primary
Supriya Kurlekar, Manasi R Dixit, Dr.
doaj +1 more source
Story2Board: A Training‐Free Approach for Expressive Visual Storytelling
Abstract We present Story2Board, a training‐free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing.
D. Dinkevich +4 more
wiley +1 more source
PBR‐Inspired Controllable Diffusion for Image Generation
Abstract Despite recent advances in text‐to‐image generation, controlling geometric layout and PBR material properties in synthesized scenes remains challenging. We present a pipeline that first produces a G‐buffer (albedo, normals, depth, roughness, shading, and metallic) from a text prompt and then renders a final image through a PBR‐inspired branch ...
Bowen Xue +3 more
wiley +1 more source
Approach of dense video captioning based on multimodal memory knowledge
Dense video captioning aims to localize events in an untrimmed video and generate a corresponding captions for each meaningful event. Existing methods mainly utilize the source video input to generate captions, and these methods are unable to capture the
FANG Haojie +3 more
doaj
Ask and focus more: Question-prompt uncertainty allocation for dual-controllable video captioning
Video captioning aims to generate natural language descriptions from video content via hierarchical architectures that capture key visual elements. Although entities, predicates, and syntactic structures are critical for coherent descriptions, existing ...
Chen, Yi +11 more
core +1 more source
Spatio-temporal attention models for grounded video captioning
Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and generate a description that relies ...
Zanfir, Mihai +5 more
core +1 more source
MultiCOIN: Multi‐Modal COntrollable INbetweening
Abstract Video inbetweening creates smooth transitions between two frames making it an indispensable tool for video editing and longform video synthesis. Existing methods struggle with large or complex motion and offer limited control over intermediate frames, often misaligning with user intent.
M. Tanveer +6 more
wiley +1 more source
Neural image and video captioning
In today’s digital age, the proliferation of visual content has underscored the critical importance of multimedia comprehension and interpretation. Video uses images and sound to convey information.
Lam, Ting En
core

