English
Related papers

Related papers: A Cocktail-Party Benchmark: Multi-Modal dataset an…

200 papers

Online learning is a rapidly growing industry. However, a major doubt about online learning is whether students are as engaged as they are in face-to-face classes. An engagement recognition system can notify the instructors about the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Chi-hsuan Wu , Shih-yang Liu , Xijie Huang , Xingbo Wang , Rong Zhang , Luca Minciullo , Wong Kai Yiu , Kenny Kwan , Kwang-Ting Cheng

Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues…

Computation and Language · Computer Science 2025-06-11 Divyaksh Shukla , Ritesh Baviskar , Dwijesh Gohil , Aniket Tiwari , Atul Shree , Ashutosh Modi

This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Vimal Manohar , Szu-Jui Chen , Zhiqi Wang , Yusuke Fujita , Shinji Watanabe , Sanjeev Khudanpur

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clustering-based speaker diarization system to handle overlapped…

Sound · Computer Science 2022-02-11 Chen Shen , Yi Liu , Wenzhi Fan , Bin Wang , Shixue Wen , Yao Tian , Jun Zhang , Jingsheng Yang , Zejun Ma

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from noise, speech from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-25 Darius Petermann , Gordon Wichern , Zhong-Qiu Wang , Jonathan Le Roux

Recently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations…

Sound · Computer Science 2022-10-12 Fan Yu , Shiliang Zhang , Pengcheng Guo , Yuhao Liang , Zhihao Du , Yuxiao Lin , Lei Xie

Emotion recognition is a crucial task for human conversation understanding. It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions. As a typical solution, the global- and the local…

Computation and Language · Computer Science 2024-01-31 Cam-Van Thi Nguyen , Anh-Tuan Mai , The-Son Le , Hai-Dang Kieu , Duc-Trong Le

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AVSR models are often…

Sound · Computer Science 2025-06-04 Thai-Binh Nguyen , Ngoc-Quan Pham , Alexander Waibel

Multimodal intent recognition is a significant task for understanding human language in real-world multimodal scenes. Most existing intent recognition methods have limitations in leveraging the multimodal information due to the restrictions…

Artificial Intelligence · Computer Science 2023-02-09 Hanlei Zhang , Hua Xu , Xin Wang , Qianrui Zhou , Shaojie Zhao , Jiayan Teng

Emulating the human ability to solve the cocktail party problem, i.e., focus on a source of interest in a complex acoustic scene, is a long standing goal of audio source separation research. Much of this research investigates separating…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-15 Darius Petermann , Gordon Wichern , Aswin Shanmugam Subramanian , Zhong-Qiu Wang , Jonathan Le Roux

Person Re-Identification (ReID) has several challenges in real-world surveillance systems due to clothing changes (CCReID) and the need for maintaining continual learning (LReID). Previous existing methods either develop models specifically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Robert Long , Rongxin Jiang , Mingrui Yan

Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content. In the era of rapidly growing…

Computation and Language · Computer Science 2025-05-20 Xingyu Li , Chen Gong , Guohong Fu

There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Minkuk Kim , Hyeon Bae Kim , Jinyoung Moon , Jinwoo Choi , Seong Tae Kim

Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong, a benchmark for natural multi-party code-switching…

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Arjun R. Akula , Song-Chun Zhu

Human Multimodal Language Understanding (MLU) aims to infer human intentions by integrating related cues from heterogeneous modalities. Existing works predominantly follow a ``learning to attend" paradigm, which maximizes mutual information…

Computation and Language · Computer Science 2025-09-29 Menghua Jiang , Yuncheng Jiang , Haifeng Hu , Sijie Mai

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Multi-modal Large Language Models (MLLM) often find it difficult…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Xiaoyi Bao , Siyang Sun , Shuailei Ma , Kecheng Zheng , Yuxin Guo , Guosheng Zhao , Yun Zheng , Xingang Wang

Conventional chatbots focus on two-party response generation, which simplifies the real dialogue scene. In this paper, we strive toward a novel task of Response Generation on Multi-Party Chatbot (RGMPC), where the generated responses…

Computation and Language · Computer Science 2019-10-30 Cao Liu , Kang Liu , Shizhu He , Zaiqing Nie , Jun Zhao

Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual and audio representations. We human visual perception could…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Shaobo Min , Qi Dai , Hongtao Xie , Chuang Gan , Yongdong Zhang , Jingdong Wang

Event coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facilitate resolution when…