Results 41 to 50 of about 400 (142)
Two‐Photon Endomicroscopy for Image‐Guided Magnetic Micro‐Agents
Image‐guided manipulation of magnetic micro‐agents requires localized fluorescence feedback capable of resolving dynamic interactions in confined environments. This study introduces a two‐photon endomicroscopy platform for real‐time motion tracking, multicolor micro‐agent discrimination, and imaging through biological tissue.
Juan J. Huaroto +3 more
wiley +1 more source
Abstract Automating bridge inspections requires more than detecting individual damage instances. It demands systems capable of describing, contextualizing, and interpreting damage in an inspection‐relevant manner. Conventional computer vision approaches, such as object detection and segmentation, primarily address visual recognition tasks and are ...
Rona Firdes Çelik +2 more
wiley +1 more source
Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning [PDF]
Dense video captioning is a newly emerging task that aims at both localizing and describing all events in a video. We identify and tackle two challenges on this task, namely, (1) how to utilize both past and future contexts for accurate event proposal predictions, and (2) how to construct informative input to the decoder for generating natural event ...
Jingwen Wang 0003 +4 more
openaire +3 more sources
ABSTRACT High voltage equipment, as a core component of power systems, plays an indispensable role in ensuring the reliability of power supply through its safe and stable operation. Traditional visual defect detection for high voltage equipment, however, often relies on manual inspection and experience‐based judgement, which struggles to meet the ...
Zhenbing Zhao +7 more
wiley +1 more source
With the rapid advancement of intelligent coal mine construction, the volume of underground operational video data has surged dramatically. Current video processing and storage methods predominantly rely on single-scene video analysis and raw-format ...
Xiang FU +6 more
doaj +1 more source
Read the free Plain Language Summary for this article on the Journal blog. Abstract Locomotion consumes a large proportion of individual energy budgets and may impose energetic constraints on other fitness‐related traits particularly under variable environmental conditions.
Miki Jahn, Frank Seebacher
wiley +1 more source
Generating an image/video caption has always been a fundamental problem of Artificial Intelligence, which is usually performed using the potential of Deep Learning Methods, Computer Vision, Knowledge Graphs, and Natural Language Processing (NLP).
Mohammad Saif Wajid +3 more
doaj +1 more source
Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning
Multimodal video understanding (MVU) has emerged as a fast-growing research frontier, driven by major advances in video-language pre-training and large multimodal models over the past decade.
Rongyong Zhao +5 more
doaj +1 more source
MultiCOIN: Multi‐Modal COntrollable INbetweening
Abstract Video inbetweening creates smooth transitions between two frames making it an indispensable tool for video editing and longform video synthesis. Existing methods struggle with large or complex motion and offer limited control over intermediate frames, often misaligning with user intent.
M. Tanveer +6 more
wiley +1 more source
The increasing popularity of digital twins, alongside the rapid evolution of connectivity driven by the Internet of Things, highlights their potential to greatly aid in the development of smart cities.
Mohammad Saif Wajid +4 more
doaj +1 more source

