English
Related papers

Related papers: MPN: Multimodal Parallel Network for Audio-Visual …

200 papers

With the rapid expansion of e-commerce, more consumers have become accustomed to making purchases via livestreaming. Accurately identifying the products being sold by salespeople, i.e., livestreaming product retrieval (LPR), poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Xiaowan Hu , Yiyi Chen , Yan Li , Minquan Wang , Haoqian Wang , Quan Chen , Han Li , Peng Jiang

Temporal Action Localization (TAL) aims to identify actions' start, end, and class labels in untrimmed videos. While recent advancements using transformer networks and Feature Pyramid Networks (FPN) have enhanced visual feature recognition…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Edward Fish , Jon Weinbren , Andrew Gilbert

Weakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Jie Fu , Junyu Gao , Changsheng Xu

We address the problem of video moment localization with natural language, i.e. localizing a video segment described by a natural language sentence. While most prior work focuses on grounding the query as a whole, temporal dependencies and…

Multimedia · Computer Science 2019-08-13 Songyang Zhang , Jinsong Su , Jiebo Luo

Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Lei Ke , Xia Li , Martin Danelljan , Yu-Wing Tai , Chi-Keung Tang , Fisher Yu

Sound events often occur in unstructured environments where they exhibit wide variations in their frequency content and temporal structure. Convolutional neural networks (CNN) are able to extract higher level features that are invariant to…

Machine Learning · Computer Science 2017-05-31 Emre Çakır , Giambattista Parascandolo , Toni Heittola , Heikki Huttunen , Tuomas Virtanen

Classification and identification of the materials lying over or beneath the Earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS) and have garnered a growing concern owing to the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Danfeng Hong , Lianru Gao , Naoto Yokoya , Jing Yao , Jocelyn Chanussot , Qian Du , Bing Zhang

Answering questions according to multi-modal context is a challenging problem as it requires a deep integration of different data sources. Existing approaches only employ partial interactions among data sources in one attention hop. In this…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Anran Wang , Anh Tuan Luu , Chuan-Sheng Foo , Hongyuan Zhu , Yi Tay , Vijay Chandrasekhar

Real-world networks often exist with multiple views, where each view describes one type of interaction among a common set of nodes. For example, on a video-sharing network, while two user nodes are linked if they have common favorite videos…

Machine Learning · Computer Science 2021-04-27 Sezin Kircali Ata , Yuan Fang , Min Wu , Jiaqi Shi , Chee Keong Kwoh , Xiaoli Li

We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an…

Multimedia · Computer Science 2017-01-05 Naoya Takahashi , Michael Gygli , Luc Van Gool

For object detection, how to address the contradictory requirement between feature map resolution and receptive field on high-resolution inputs still remains an open question. In this paper, to tackle this issue, we build a novel…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Junxu Cao , Qi Chen , Jun Guo , Ruichao Shi

Temporal localization remains an important challenge in video understanding. In this work, we present our solution to the 3rd YouTube-8M Video Understanding Challenge organized by Google Research. Participants were required to build a…

Computer Vision and Pattern Recognition · Computer Science 2019-11-19 Lijun Zhang , Srinath Nizampatnam , Ahana Gangopadhyay , Marcos V. Conde

Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long videos with clean and simple audio context, short videos have…

Multimedia · Computer Science 2025-04-10 Wuyang Liu , Yi Chai , Yongpeng Yan , Yanzhen Ren

With the proliferation of imaging sensors, the volume of multi-modal imagery far exceeds the ability of human analysts to adequately consume and exploit it. Full motion video (FMV) possesses the extra challenge of containing large amounts…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Marc Bosch , Joseph Nassar , Benjamin Ortiz , Brendan Lammers , David Lindenbaum , John Wahl , Robert Mangum , Margaret Smith

Millimeter-wave radars are being increasingly integrated into commercial vehicles to support new advanced driver-assistance systems by enabling robust and high-performance object detection, localization, as well as recognition - a key…

Signal Processing · Electrical Eng. & Systems 2022-05-02 Xiangyu Gao , Guanbin Xing , Sumit Roy , Hui Liu

Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional methods usually focus…

Artificial Intelligence · Computer Science 2024-07-12 Jinxing Zhou , Dan Guo , Yuxin Mao , Yiran Zhong , Xiaojun Chang , Meng Wang

Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep…

Long-term time series forecasting plays an important role in various real-world scenarios. Recent deep learning methods for long-term series forecasting tend to capture the intricate patterns of time series by decomposition-based or…

Machine Learning · Computer Science 2023-06-13 Xing Wang , Zhendong Wang , Kexin Yang , Junlan Feng , Zhiyan Song , Chao Deng , Lin zhu

In recent years, encoder-decoder networks have focused on expanding receptive fields and incorporating multi-scale context to capture global features for objects of varying sizes. However, as networks deepen, they often discard fine spatial…

Image and Video Processing · Electrical Eng. & Systems 2024-09-20 Xiaogang Du , Dongxin Gu , Tao Lei , Yipeng Jiao , Yibin Zou

In the field of computer vision, recent works show that a pure MLP architecture mainly stacked by fully-connected layers can achieve competing performance with CNN and transformer. An input image of vision MLP is usually split into multiple…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Yehui Tang , Kai Han , Jianyuan Guo , Chang Xu , Yanxi Li , Chao Xu , Yunhe Wang