中文
相关论文

相关论文: STHG: Spatial-Temporal Heterogeneous Graph Learnin…

200 篇论文

The DIarization of SPeaker and LAnguage in Conversational Environments (DISPLACE) 2024 challenge is the second in the series of DISPLACE challenges, which involves tasks of speaker diarization (SD) and language diarization (LD) on a…

The onset of long-form egocentric datasets such as Ego4D and EPIC-Kitchens presents a new challenge for the task of Temporal Sentence Grounding (TSG). Compared to traditional benchmarks on which this task is evaluated, these datasets offer…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Kevin Flanagan , Dima Damen , Michael Wray

This paper describes system setup of our submission to speaker diarisation track (Track 4) of VoxCeleb Speaker Recognition Challenge 2020. Our diarisation system consists of a well-trained neural network based speech enhancement model as…

声音 · 计算机科学 2020-10-26 Renyu Wang , Ruilin Tong , Yu Ting Yeung , Xiao Chen

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025…

In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which models spatiotemporal visual-linguistic dependencies with a…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Zihang Lin , Chaolei Tan , Jian-Fang Hu , Zhi Jin , Tiancai Ye , Wei-Shi Zheng

Recent advances in Graph Neural Networks (GNNs) have revolutionized graph-structured data modeling, yet traditional GNNs struggle with complex heterogeneous structures prevalent in real-world scenarios. Despite progress in handling…

机器学习 · 计算机科学 2025-01-07 Zongwei Li , Lianghao Xia , Hua Hua , Shijie Zhang , Shuangyang Wang , Chao Huang

The LEAP submission for DIHARD-III challenge is described in this paper. The proposed system is composed of a speech bandwidth classifier, and diarization systems fine-tuned for narrowband and wideband speech separately. We use an…

音频与语音处理 · 电气工程与系统科学 2021-06-15 Prachi Singh , Rajat Varma , Venkat Krishnamohan , Srikanth Raj Chetupalli , Sriram Ganapathy

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Yisen Feng , Haoyu Zhang , Qiaohui Chu , Meng Liu , Weili Guan , Yaowei Wang , Liqiang Nie

Everyday communication is dynamic and multisensory, often involving shifting attention, overlapping speech and visual cues. Yet, most neural attention tracking studies are still limited to highly controlled lab settings, using clean, often…

信号处理 · 电气工程与系统科学 2026-01-22 Johanna Wilroth , Oskar Keding , Martin A. Skoglund , Maria Sandsten , Martin Enqvist , Emina Alickovic

In this paper, we present the submitted system for the third DIHARD Speech Diarization Challenge from the DKU-Duke-Lenovo team. Our system consists of several modules: voice activity detection (VAD), segmentation, speaker embedding…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Weiqing Wang , Qingjian Lin , Danwei Cai , Lin Yang , Ming Li

Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization approaches consider acoustic information only, which result in…

计算与语言 · 计算机科学 2023-05-23 Luyao Cheng , Siqi Zheng , Zhang Qinglin , Hui Wang , Yafeng Chen , Qian Chen

This paper introduces the third DIHARD challenge, the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational…

音频与语音处理 · 电气工程与系统科学 2020-12-04 Neville Ryant , Kenneth Church , Christopher Cieri , Jun Du , Sriram Ganapathy , Mark Liberman

Supervised approaches for learning spatio-temporal scene graphs (STSG) from video are greatly hindered due to their reliance on STSG-annotated videos, which are labor-intensive to construct at scale. Is it feasible to instead use readily…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Jiani Huang , Ziyang Li , Mayur Naik , Ser-Nam Lim

We present a novel approach to Speaker Diarization (SD) by leveraging text-based methods focused on Sentence-level Speaker Change Detection within dialogues. Unlike audio-based SD systems, which are often challenged by audio quality and…

计算与语言 · 计算机科学 2025-06-16 Peilin Wu , Jinho D. Choi

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

多媒体 · 计算机科学 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

Conventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multi-step approach generates speaker assignments…

声音 · 计算机科学 2023-02-27 Prachi Singh , Amrit Kaul , Sriram Ganapathy

With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this…

声音 · 计算机科学 2022-08-09 A Kishore Kumar , Shefali Waldekar , Md Sahidullah , Goutam Saha

This paper presents a novel method to predict future human activities from partially observed RGB-D videos. Human activity prediction is generally difficult due to its non-Markovian property and the rich context between human and…

计算机视觉与模式识别 · 计算机科学 2017-08-04 Siyuan Qi , Siyuan Huang , Ping Wei , Song-Chun Zhu

In multi-lingual societies, where multiple languages are spoken in a small geographic vicinity, informal conversations often involve mix of languages. Existing speech technologies may be inefficient in extracting information from such…

音频与语音处理 · 电气工程与系统科学 2024-01-04 Shikha Baghel , Shreyas Ramoji , Somil Jain , Pratik Roy Chowdhuri , Prachi Singh , Deepu Vijayasenan , Sriram Ganapathy

This paper describes our solution for the Diarization of Speaker and Language in Conversational Environments Challenge (Displace 2023). We used a combination of VAD for finding segfments with speech, Resnet architecture based CNN for…

计算与语言 · 计算机科学 2024-06-25 Ali Aliyev