Musicians are more consistent: Gestural cross-modal mappings of pitch, loudness and tempo in real-time [PDF]
Cross-modal mappings of auditory stimuli reveal valuable insights into how humans make sense of sound and music. Whereas researchers have investigated cross-modal mappings of sound features varied in isolation within paradigms such as speeded ...
Mats B. Küssner +3 more
doaj +4 more sources
Symmetry-Driven Multimodal Adversarial Attacks: An Information-Theoretic Perspective on Cross-Modal Invariance and Robustness [PDF]
Multimodal models such as CLIP and ALBEF essentially maximize cross-modal mutual information to align heterogeneous modalities, utilizing semantic consistency as an implicit prior. However, this alignment mechanism creates a structural vulnerability: the
Jin Wei +3 more
doaj +2 more sources
Beyond “Taiji Diagram”: how multimodal metaphor deciphers Yin Xu and Yang Xu through Traditional Chinese Medicine science communication short videos [PDF]
IntroductionThis study investigates how multimodal metaphors convey two core Traditional Chinese Medicine (TCM) concepts— “Yin Xu” (Yin deficiency) and “Yang Xu” (Yang deficiency)—in short science communication videos on Douyin.
Shaoci Wang +4 more
doaj +2 more sources
Natural cross-modal mappings between visual and auditory features. [PDF]
The brain may combine information from different sense modalities to enhance the speed and accuracy of detection of objects and events, and the choice of appropriate responses. There is mounting evidence that perceptual experiences that appear to be modality-specific are also influenced by activity from other sensory modalities, even in the absence of ...
Evans KK, Treisman A.
europepmc +3 more sources
Learning visual to auditory sensory substitution reveals flexibility in image to sound mapping [PDF]
Visual-to-auditory sensory substitution devices (SSDs) translate images to sounds. One SSD, The vOICe, translates a pixel’s vertical position into pitch and horizontal position into time.
Asa Kucinkas +5 more
doaj +2 more sources
Hearing movement, seeing sound: multimodal predictive coding in pianist-dancer interaction [PDF]
Live piano accompaniment for dance poses a “zero-latency paradox”: performers achieve near-simultaneous audiovisual alignment despite sensory and integration delays that should make purely reactive control too slow. This review argues that pianist–dancer
Xinyu Cao, Xinlei Shi
doaj +2 more sources
What Sound Does That Taste? Cross-Modal Mappings across Gustation and Audition [PDF]
All people share implicit mappings across the senses, which give us preferences for certain sensory combinations over others (eg light colours are preferentially paired with higher-pitch sounds; Ward et al, 2006 Cortex42 264–280). Although previous work has tended to focus on the cross-modality of vision with other senses, here we present evidence of ...
Christine Cuskley +2 more
exaly +4 more sources
Weakly Supervised Fine-Grained Discrimination of Wheat Mold Using Local RGB–HSI Fusion [PDF]
Wheat is a major staple crop, and storage mold growth poses a severe threat to grain safety and quality stability. Natural mold development in stored wheat exhibits subtle, localized, and highly heterogeneous characteristics.
Le Xiao, Shengtong Wang, Lulu Niu
doaj +2 more sources
Cross-modal distillation for flood extent mapping
Abstract The increasing intensity and frequency of floods is one of the many consequences of our changing climate. In this work, we explore ML techniques that improve the flood detection module of an operational early flood warning system. Our method exploits an unlabeled dataset of paired multi-spectral and synthetic aperture radar (SAR) imagery to
Shubhika Garg +6 more
openaire +3 more sources
A Universal Model for Cross Modality Mapping by Relational Reasoning
With the aim of matching a pair of instances from two different modalities, cross modality mapping has attracted growing attention in the computer vision community. Existing methods usually formulate the mapping function as the similarity measure between the pair of instance features, which are embedded to a common space.
Zun Li 0001 +6 more
openaire +2 more sources

