中文
相关论文

相关论文: Spatio-channel Attention Blocks for Cross-modal Cr…

200 篇论文

In this paper, we propose a novel architecture for multi-modal speech and text input. We combine pretrained speech and text encoders using multi-headed cross-modal attention and jointly fine-tune on the target problem. The resultant…

计算与语言 · 计算机科学 2022-04-21 Karan Singla , Daniel Pressel , Ryan Price , Bhargav Srinivas Chinnari , Yeon-Jun Kim , Srinivas Bangalore

Leveraging complementary relationships across modalities has recently drawn a lot of attention in multimodal emotion recognition. Most of the existing approaches explored cross-attention to capture the complementary relationships across the…

计算机视觉与模式识别 · 计算机科学 2024-07-02 G Rajasekhar , Jahangir Alam

Neural networks can be used in video coding to improve chroma intra-prediction. In particular, usage of fully-connected networks has enabled better cross-component prediction with respect to traditional linear models. Nonetheless,…

图像与视频处理 · 电气工程与系统科学 2020-06-30 Marc Górriz , Saverio Blasi , Alan F. Smeaton , Noel E. O'Connor , Marta Mrak

Understanding the movement patterns of objects (e.g., humans and vehicles) in a city is essential for many applications, including city planning and management. This paper proposes a method for predicting future city-wide crowd flows by…

机器学习 · 计算机科学 2023-10-05 Chung Park , Junui Hong , Cheonbok Park , Taesan Kim , Minsung Choi , Jaegul Choo

Image alignment, also known as image registration, is a critical block used in many computer vision problems. One of the key factors in alignment is efficiency, as inefficient aligners can cause significant overhead to the overall problem.…

计算机视觉与模式识别 · 计算机科学 2022-06-02 Bahri Batuhan Bilecen , Alparslan Fisne , Mustafa Ayazoglu

Clustering is fundamental for gaining insights from complex networks, and spectral clustering (SC) is a popular approach. Conventional SC focuses on second-order structures (e.g., edges connecting two nodes) without direct consideration of…

机器学习 · 计算机科学 2018-12-27 Yan Ge , Haiping Lu , Pan Peng

Detection-based methods have been viewed unfavorably in crowd analysis due to their poor performance in dense crowds. However, we argue that the potential of these methods has been underestimated, as they offer crucial information for crowd…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Shaokai Wu , Fengyu Yang

The integration of RGB and thermal data can significantly improve semantic segmentation performance in wild environments for field robots. Nevertheless, multi-source data processing (e.g. Transformer-based approaches) imposes significant…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Xiaodong Guo , Zi'ang Lin , Luwen Hu , Zhihong Deng , Tong Liu , Wujie Zhou

In action recognition, although the combination of spatio-temporal videos and skeleton features can improve the recognition performance, a separate model and balancing feature representation for cross-modal data are required. To solve these…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Dasom Ahn , Sangwon Kim , Hyunsu Hong , Byoung Chul Ko

Cross-modal retrieval aims to retrieve data in one modality by a query in another modality, which has been a very interesting research issue in the field of multimedia, information retrieval, and computer vision, and database. Most existing…

多媒体 · 计算机科学 2021-05-06 Donghuo Zeng , Yi Yu , Keizo Oyama

Recently, many attention-based deep neural networks have emerged and achieved state-of-the-art performance in environmental sound classification. The essence of attention mechanism is assigning contribution weights on different parts of…

音频与语音处理 · 电气工程与系统科学 2020-11-06 You Wang , Chuyao Feng , David V. Anderson

Human action video recognition has recently attracted more attention in applications such as video security and sports posture correction. Popular solutions, including graph convolutional networks (GCNs) that model the human skeleton as a…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Zhendong Liu , Haifeng Xia , Tong Guo , Libo Sun , Ming Shao , Siyu Xia

The community detection problem on multilayer networks have drawn much interest. When the nodal covariates ar also present, few work has been done to integrate information from both sources. To leverage the multilayer networks and the…

统计方法学 · 统计学 2025-03-13 Da Zhao , Wanjie Wang , Jialiang Li

Social group activity recognition is a challenging task extended from group activity recognition, where social groups must be recognized with their activities and group members. Existing methods tackle this task by leveraging region…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Masato Tamura

With the rapid advances in high-throughput sequencing technologies, the focus of survival analysis has shifted from examining clinical indicators to incorporating genomic profiles with pathological images. However, existing methods either…

图像与视频处理 · 电气工程与系统科学 2023-09-25 Fengtao Zhou , Hao Chen

The proliferation of advanced mobile terminals opened up a new crowdsourcing avenue, spatial crowdsourcing, to utilize the crowd potential to perform real-world tasks. In this work, we study a new type of spatial crowdsourcing, called…

数据库 · 计算机科学 2020-10-30 Ting Wang , Xike Xie , Xin Cao , Torben Bach Pedersen , Yang Wang , Mingjun Xiao

Hyperspectral image (HSI) and LiDAR data joint classification is a challenging task. Existing multi-source remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on…

图像与视频处理 · 电气工程与系统科学 2025-03-11 Junyan Lin , Feng Gap , Lin Qi , Junyu Dong , Qian Du , Xinbo Gao

Scaling methods have long been utilized to simplify and cluster high-dimensional data. However, the general latent spaces across all predefined groups derived from these methods sometimes do not fall into researchers' interest regarding…

社会与信息网络 · 计算机科学 2023-06-02 Takanori Fujiwara , Tzu-Ping Liu

Non-local attention module has been proven to be crucial for image restoration. Conventional non-local attention processes features of each layer separately, so it risks missing correlation between features among different layers. To…

图像与视频处理 · 电气工程与系统科学 2023-04-21 Yancheng Wang , Ning Xu , Yingzhen Yang

Self-attention mechanism recently achieves impressive advancement in Natural Language Processing (NLP) and Image Processing domains. And its permutation invariance property makes it ideally suitable for point cloud processing. Inspired by…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Xian-Feng Han , Zhang-Yue He , Jia Chen , Guo-Qiang Xiao