English
Related papers

Related papers: Multi-scale Geometric Summaries for Similarity-bas…

200 papers

Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements can help improve the performance. In this work, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-30 Rohith Aralikatti , Sharad Roy , Abhinav Thanda , Dilip Kumar Margam , Pujitha Appan Kandala , Tanay Sharma , Shankar M Venkatesan

Image fusion technology is widely used to fuse the complementary information between multi-source remote sensing images. Inspired by the frontier of deep learning, this paper first proposes a heterogeneous-integrated framework based on a…

Image and Video Processing · Electrical Eng. & Systems 2024-05-15 Menghui Jiang , Huanfeng Shen , Jie Li , Liangpei Zhang

Hyperspectral video (HSV) offers valuable spatial, spectral, and temporal information simultaneously, making it highly suitable for handling challenges such as background clutter and visual similarity in object tracking. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Hanzheng Wang , Wei Li , Xiang-Gen Xia , Qian Du , Jing Tian

Multimodal change detection (MMCD) identifies changed areas in multimodal remote sensing (RS) data, demonstrating significant application value in land use monitoring, disaster assessment, and urban sustainable development. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Xuanguang Liu , Lei Ding , Yujie Li , Chenguang Dai , Zhenchao Zhang , Mengmeng Li , Ziyi Yang , Yifan Sun , Yongqi Sun , Hanyun Wang

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance changes caused by…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Abhinav Valada , Rohit Mohan , Wolfram Burgard

Effective fusion of multi-scale features is crucial for improving speaker verification performance. While most existing methods aggregate multi-scale features in a layer-wise manner via simple operations, such as summation or concatenation.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-04 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Jiajun Qi

We present an approach named JSFusion (Joint Sequence Fusion) that can measure semantic similarity between any pairs of multimodal sequence data (e.g. a video clip and a language sentence). Our multimodal matching network consists of two…

Computer Vision and Pattern Recognition · Computer Science 2018-08-09 Youngjae Yu , Jongseok Kim , Gunhee Kim

In a previous work \citep{luo2016sparse2d_spej}, the authors proposed an ensemble-based 4D seismic history matching (SHM) framework, which has some relatively new ingredients, in terms of the type of seismic data in choice, the way to…

Data Analysis, Statistics and Probability · Physics 2018-03-13 Xiaodong Luo , Tuhin Bhakta , Morten Jakobsen , Geir Nævdal

Event detection has been one of the most important research topics in social media analysis. Most of the traditional approaches detect events based on fixed temporal and spatial resolutions, while in reality events of different scales…

Social and Information Networks · Computer Science 2015-08-31 Xiaowen Dong , Dimitrios Mavroeidis , Francesco Calabrese , Pascal Frossard

Oceanic processes at fine scales are crucial yet difficult to observe accurately due to limitations in satellite and in-situ measurements. The Surface Water and Ocean Topography (SWOT) mission provides high-resolution Sea Surface Height…

Atmospheric and Oceanic Physics · Physics 2025-03-28 Eugenio Cutolo , Carlos Granero-Belinchon , Ptashanna Thiraux , Jinbo Wang , Ronan Fablet

In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly…

Sound · Computer Science 2025-05-07 Ya Li , Bin Zhou , Bo Hu

Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches…

Sound · Computer Science 2024-06-14 Xueyu Liu , Jie Lin , Chao Wang

Leveraging information across diverse modalities is known to enhance performance on multimodal segmentation tasks. However, effectively fusing information from different modalities remains challenging due to the unique characteristics of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Md Kaykobad Reza , Ashley Prater-Bennette , M. Salman Asif

This article deals with the fusion of flaw detections from multi-sensor nondestructive materials testing. Because each testing method makes use of different physical effects for defect localization, a multi-method approach is promising to…

Computational Engineering, Finance, and Science · Computer Science 2015-07-01 René Heideklang , Parisa Shokouhi

Multivariate time series (MTS) classification is widely applied in fields such as industry, healthcare, and finance, aiming to extract key features from complex time series data for accurate decision-making and prediction. However, existing…

Machine Learning · Computer Science 2025-06-19 Mingsen Du , Meng Chen , Yongjian Li , Cun Ji , Shoushui Wei

The extent to which different neural or artificial neural networks (models) rely on equivalent representations to support similar tasks remains a central question in neuroscience and machine learning. Prior work has typically compared…

Neurons and Cognition · Quantitative Biology 2026-04-06 Jialin Wu , Shreya Saha , Yiqing Bo , Meenakshi Khosla

This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consideration. Specifically, we propose to use generative learning…

Sound · Computer Science 2021-02-04 Yucheng Zhao , Dacheng Yin , Chong Luo , Zhiyuan Zhao , Chuanxin Tang , Wenjun Zeng , Zheng-Jun Zha

We introduce SMUTF (Schema Matching Using Generative Tags and Hybrid Features), a unique approach for large-scale tabular data schema matching (SM), which assumes that supervised learning does not affect performance in open-domain tasks,…

Computation and Language · Computer Science 2025-05-06 Yu Zhang , Mei Di , Haozheng Luo , Chenwei Xu , Richard Tzong-Han Tsai

Video object detection is a tough task due to the deteriorated quality of video sequences captured under complex environments. Currently, this area is dominated by a series of feature enhancement based methods, which distill beneficial…

Computer Vision and Pattern Recognition · Computer Science 2020-09-17 Lijian Lin , Haosheng Chen , Honglun Zhang , Jun Liang , Yu Li , Ying Shan , Hanzi Wang

Euclidean representation learning methods have achieved promising results in image fusion tasks, which can be attributed to their clear advantages in handling with linear space. However, data collected from a realistic scene usually has a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Huan Kang , Hui Li , Tianyang Xu , Xiao-Jun Wu , Rui Wang , Chunyang Cheng , Josef Kittler