中文
相关论文

相关论文: TRESTLE: Toolkit for Reproducible Execution of Spe…

200 篇论文

NeurST is an open-source toolkit for neural speech translation. The toolkit mainly focuses on end-to-end speech translation, which is easy to use, modify, and extend to advanced speech translation research and products. NeurST aims at…

计算与语言 · 计算机科学 2021-06-16 Chengqi Zhao , Mingxuan Wang , Qianqian Dong , Rong Ye , Lei Li

Various robustness evaluation methodologies from different perspectives have been proposed for different natural language processing (NLP) tasks. These methods have often focused on either universal or task-specific generalization…

In recent years there has been a burgeoning interest in the use of computational methods to distinguish between elicited speech samples produced by patients with dementia, and those from healthy controls. The difference between perplexity…

计算与语言 · 计算机科学 2020-06-30 Trevor Cohen , Serguei Pakhomov

The literature on concept formation has demonstrated that humans are capable of learning concepts incrementally, with a variety of attribute types, and in both supervised and unsupervised settings. Many models of concept formation focus on…

人工智能 · 计算机科学 2024-10-15 Christopher J. MacLellan , Erik Harpstead , Vincent Aleven , Kenneth R. Koedinger

This study introduces \textbf{InteractEval}, a framework that integrates human expertise and Large Language Models (LLMs) using the Think-Aloud (TA) method to generate attributes for checklist-based text evaluation. By combining human…

计算与语言 · 计算机科学 2025-02-21 SeongYeub Chu , JongWoo Kim , MunYong Yi

Speech language models (SpeechLMs) process and generate acoustic data only, without textual supervision. In this work, we propose TWIST, a method for training SpeechLMs using a warm-start from a pretrained textual language models. We show…

We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a…

计算与语言 · 计算机科学 2023-07-11 Gati Aher , Rosa I. Arriaga , Adam Tauman Kalai

Objectives: While Large Language Models (LLMs) have been widely used to assist clinicians and support patients, no existing work has explored dialogue systems for standard diagnostic interviews and assessments. This study aims to bridge the…

计算与语言 · 计算机科学 2025-05-01 Sichang Tu , Abigail Powers , Stephen Doogan , Jinho D. Choi

In this paper, we introduce TextBrewer, an open-source knowledge distillation toolkit designed for natural language processing. It works with different neural network models and supports various kinds of supervised learning tasks, such as…

计算与语言 · 计算机科学 2020-12-14 Ziqing Yang , Yiming Cui , Zhipeng Chen , Wanxiang Che , Ting Liu , Shijin Wang , Guoping Hu

In this paper, we present WEST(WE Speech Toolkit), a speech toolkit based on a large language model (LLM) for speech understanding, generation, and interaction. There are three key features of WEST: 1) Fully LLM-based: Standing on the…

计算与语言 · 计算机科学 2025-10-30 Binbin Zhang , Chengdong Liang , Shuai Wang , Xuelong Geng , Zhao Guo , Haoyu Li , Hao Yin , Xipeng Yang , Pengshen Zhang , Changwei Ma , Lei Xie

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Shuai Wang , Ke Zhang , Shaoxiong Lin , Junjie Li , Xuefei Wang , Meng Ge , Jianwei Yu , Yanmin Qian , Haizhou Li

During the last few years, spoken language technologies have known a big improvement thanks to Deep Learning. However Deep Learning-based algorithms require amounts of data that are often difficult and costly to gather. Particularly,…

声音 · 计算机科学 2019-01-15 Noé Tits , Kevin El Haddad , Thierry Dutoit

Accurate far-field speech datasets are critical for tasks such as automatic speech recognition (ASR), dereverberation, speech enhancement, and source separation. However, current datasets are limited by the trade-off between acoustic…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Sarabeth S. Mullins , Georg Götz , Eric Bezzam , Steven Zheng , Daniel Gert Nielsen

In this paper, we define and apply representational stability analysis (ReStA), an intuitive way of analyzing neural language models. ReStA is a variant of the popular representational similarity analysis (RSA) in cognitive neuroscience.…

人工智能 · 计算机科学 2019-06-06 Samira Abnar , Lisa Beinborn , Rochelle Choenni , Willem Zuidema

Large Language Models (LLMs) falter in multi-step interactions -- often hallucinating, repeating actions, or misinterpreting user corrections -- due to reliance on linear, unstructured context. This fragility stems from the lack of…

人工智能 · 计算机科学 2025-05-27 Ye Ye

Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and…

Machine learning-based behavioral models rely on features extracted from audio-visual recordings. The recordings are processed using open-source tools to extract speech features for classification models. These tools often lack validation…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Tahiya Chowdhury , Veronica Romero

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ…

音频与语音处理 · 电气工程与系统科学 2023-12-15 Yifan Yang , Feiyu Shen , Chenpeng Du , Ziyang Ma , Kai Yu , Daniel Povey , Xie Chen

Large language models (LLMs) are increasingly used in the social sciences to simulate human behavior, based on the assumption that they can generate realistic, human-like text. Yet this assumption remains largely untested. Existing…

计算与语言 · 计算机科学 2025-11-26 Nicolò Pagan , Petter Törnberg , Christopher A. Bail , Anikó Hannák , Christopher Barrie

We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal…

‹ 上一页 1 2 3 10 下一页 ›