English
Related papers

Related papers: Lhotse: a speech data representation library for t…

200 papers

Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fixed encoder can bridge on its own. The…

Artificial Intelligence · Computer Science 2026-05-29 Vaishali Senthil , Ashutosh Hathidara , Sebastian Schreiber

Training large vision-language models requires extensive, high-quality image-text pairs. Existing web-scraped datasets, however, are noisy and lack detailed image descriptions. To bridge this gap, we introduce PixelProse, a comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Vasu Singla , Kaiyu Yue , Sukriti Paul , Reza Shirkavand , Mayuka Jayawardhana , Alireza Ganjdanesh , Heng Huang , Abhinav Bhatele , Gowthami Somepalli , Tom Goldstein

Recently, deep learning models have been widely applied in program understanding tasks, and these models achieve state-of-the-art results on many benchmark datasets. A major challenge of deep learning for program understanding is that the…

Software Engineering · Computer Science 2024-01-02 Wenhan Wang , Yanzhou Li , Anran Li , Jian Zhang , Wei Ma , Yang Liu

Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Gerald Schwiebert , Cornelius Weber , Leyuan Qu , Henrique Siqueira , Stefan Wermter

This thesis presents new methods for unsupervised learning of distributed representations of words and entities from text and knowledge bases. The first algorithm presented in the thesis is a multi-view algorithm for learning…

Computation and Language · Computer Science 2019-06-14 Pushpendre Rastogi

This paper presents a new approach to fine-tuning OpenAI's Whisper model for low-resource languages by introducing a novel data generation method that converts sentence-level data into a long-form corpus, using Swiss German as a case study.…

Computation and Language · Computer Science 2025-04-23 Vincenzo Timmel , Claudio Paonessa , Reza Kakooee , Manfred Vogel , Daniel Perruchoud

Language Models (LMs) struggle with linguistic understanding at the discourse level, even though discourse patterns such as coherence, cohesion, and narrative flow are prevalent in their pre-training data. To improve the discourse…

Computation and Language · Computer Science 2026-02-17 Zachary Bamberger , Ofek Glick , Chaim Baskin , Yonatan Belinkov

Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the…

Human-Computer Interaction · Computer Science 2024-08-23 Silvan Mertes , Daksitha Withanage Don , Otto Grothe , Johanna Kuch , Ruben Schlagowski , Elisabeth André

As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives, evaluating this…

Computation and Language · Computer Science 2025-10-14 Pietro Bernardelle , Leon Fröhling , Stefano Civelli , Gianluca Demartini

Multi-turn dialogue reading comprehension aims to teach machines to read dialogue contexts and solve tasks such as response selection and answering questions. The major challenges involve noisy history contexts and especial prerequisites of…

Computation and Language · Computer Science 2021-02-11 Zhuosheng Zhang , Junlong Li , Hai Zhao

This paper introduces a novel Russian speech dataset called Golos, a large corpus suitable for speech research. The dataset mainly consists of recorded audio files manually annotated on the crowd-sourcing platform. The total duration of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-21 Nikolay Karpov , Alexander Denisenko , Fedor Minkin

We introduce UltraSuite, a curated repository of ultrasound and acoustic data, collected from recordings of child speech therapy sessions. This release includes three data collections, one from typically developing children and two from…

Computation and Language · Computer Science 2019-07-02 Aciel Eshky , Manuel Sam Ribeiro , Joanne Cleland , Korin Richmond , Zoe Roxburgh , James Scobbie , Alan Wrench

We present OctNet, a representation for deep learning with sparse 3D data. In contrast to existing models, our representation enables 3D convolutional networks which are both deep and high resolution. Towards this goal, we exploit the…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Gernot Riegler , Ali Osman Ulusoy , Andreas Geiger

We introduce Shennong, a Python toolbox and command-line utility for speech features extraction. It implements a wide range of well-established state of art algorithms including spectro-temporal filters such as Mel-Frequency Cepstral…

Computation and Language · Computer Science 2023-02-09 Mathieu Bernard , Maxime Poli , Julien Karadayi , Emmanuel Dupoux

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational…

Sound · Computer Science 2023-05-04 Jinlong Xue , Yayue Deng , Fengping Wang , Ya Li , Yingming Gao , Jianhua Tao , Jianqing Sun , Jiaen Liang

We study the coarse-grained selection module in retrieval-based chatbot. Coarse-grained selection is a basic module in a retrieval-based chatbot, which constructs a rough candidate set from the whole database to speed up the interaction…

Computation and Language · Computer Science 2020-12-21 Tian Lan , Xian-Ling Mao , Xiaoyan Gao , Wei Wei , Heyan Huang

In high-noise environments such as factories, subways, and busy streets, capturing clear speech is challenging. Throat microphones can offer a solution because of their inherent noise-suppression capabilities; however, the passage of sound…

Sound · Computer Science 2026-04-23 Yunsik Kim , Yonghun Song , Yoonyoung Chung

This paper present a strong data mining method based on rough set, which can realize feature selection, classification and knowledge representation at the same time. Rough set has good interpretability, and is a popular method for feature…

Machine Learning · Computer Science 2022-01-13 Shuyin Xia , Xinyu Bai , Guoyin Wang , Deyu Meng , Xinbo Gao , Zizhong Chen , Elisabeth Giem

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-01 Shilong Wu

In this paper, we are interested in exploiting textual and acoustic data of an utterance for the speech emotion classification task. The baseline approach models the information from audio and text independently using two deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-02 Seunghyun Yoon , Seokhyun Byun , Subhadeep Dey , Kyomin Jung