English
Related papers

Related papers: A Study on Zero-shot Non-intrusive Speech Assessme…

200 papers

Diagnosing language disorders associated with autism is a complex challenge, often hampered by the subjective nature and variability of traditional assessment methods. Traditional diagnostic methods not only require intensive human effort…

Computation and Language · Computer Science 2024-12-02 Chuanbo Hu , Wenqi Li , Mindi Ruan , Xiangxu Yu , Shalaka Deshpande , Lynn K. Paul , Shuo Wang , Xin Li

Trained on 680,000 hours of massive speech data, Whisper is a multitasking, multilingual speech foundation model demonstrating superior performance in automatic speech recognition, translation, and language identification. However, its…

Sound · Computer Science 2024-07-16 Li Zhang , Ning Jiang , Qing Wang , Yue Li , Quan Lu , Lei Xie

Recently, large pretrained language models have demonstrated strong language understanding capabilities. This is particularly reflected in their zero-shot and in-context learning abilities on downstream tasks through prompting. To assess…

Computation and Language · Computer Science 2023-08-21 Mutian He , Philip N. Garner

In this work, we induce character-level noise in various forms when fine-tuning BERT to enable zero-shot cross-lingual transfer to unseen dialects and languages. We fine-tune BERT on three sentence-level classification tasks and evaluate…

Computation and Language · Computer Science 2023-04-03 Aarohi Srivastava , David Chiang

We present an efficient end-to-end approach for holistic Automatic Speaking Assessment (ASA) of multi-part second-language tests, developed for the 2025 Speak & Improve Challenge. Our system's main novelty is the ability to process all four…

Computation and Language · Computer Science 2025-10-07 Nhan Phan , Anusha Porwal , Yaroslav Getman , Ekaterina Voskoboinik , Tamás Grósz , Mikko Kurimo

This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-21 Siyin Wang , Chao-Han Huck Yang , Ji Wu , Chao Zhang

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Large Language Models (LLMs) are demonstrating remarkable human like capabilities across diverse domains, including psychological assessment. This study evaluates whether LLMs, specifically GPT-4o and GPT-4o mini, can infer Big Five…

Computation and Language · Computer Science 2025-01-14 Jianfeng Zhu , Ruoming Jin , Karin G. Coifman

Effective extraction and application of linguistic features are central to the enhancement of spoken Language IDentification (LID) performance. With the success of recent large models, such as GPT and Whisper, the potential to leverage such…

Computation and Language · Computer Science 2023-12-19 Peng Shen , Xuguang Lu , Hisashi Kawai

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot…

Computation and Language · Computer Science 2020-12-02 Tu Anh Nguyen , Maureen de Seyssel , Patricia Rozé , Morgane Rivière , Evgeny Kharitonov , Alexei Baevski , Ewan Dunbar , Emmanuel Dupoux

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning…

Computation and Language · Computer Science 2024-10-29 OpenAI , : , Aaron Hurst , Adam Lerer , Adam P. Goucher , Adam Perelman , Aditya Ramesh , Aidan Clark , AJ Ostrow , Akila Welihinda , Alan Hayes , Alec Radford , Aleksander Mądry , Alex Baker-Whitcomb , Alex Beutel , Alex Borzunov , Alex Carney , Alex Chow , Alex Kirillov , Alex Nichol , Alex Paino , Alex Renzin , Alex Tachard Passos , Alexander Kirillov , Alexi Christakis , Alexis Conneau , Ali Kamali , Allan Jabri , Allison Moyer , Allison Tam , Amadou Crookes , Amin Tootoochian , Amin Tootoonchian , Ananya Kumar , Andrea Vallone , Andrej Karpathy , Andrew Braunstein , Andrew Cann , Andrew Codispoti , Andrew Galu , Andrew Kondrich , Andrew Tulloch , Andrey Mishchenko , Angela Baek , Angela Jiang , Antoine Pelisse , Antonia Woodford , Anuj Gosalia , Arka Dhar , Ashley Pantuliano , Avi Nayak , Avital Oliver , Barret Zoph , Behrooz Ghorbani , Ben Leimberger , Ben Rossen , Ben Sokolowsky , Ben Wang , Benjamin Zweig , Beth Hoover , Blake Samic , Bob McGrew , Bobby Spero , Bogo Giertler , Bowen Cheng , Brad Lightcap , Brandon Walkin , Brendan Quinn , Brian Guarraci , Brian Hsu , Bright Kellogg , Brydon Eastman , Camillo Lugaresi , Carroll Wainwright , Cary Bassin , Cary Hudson , Casey Chu , Chad Nelson , Chak Li , Chan Jun Shern , Channing Conger , Charlotte Barette , Chelsea Voss , Chen Ding , Cheng Lu , Chong Zhang , Chris Beaumont , Chris Hallacy , Chris Koch , Christian Gibson , Christina Kim , Christine Choi , Christine McLeavey , Christopher Hesse , Claudia Fischer , Clemens Winter , Coley Czarnecki , Colin Jarvis , Colin Wei , Constantin Koumouzelis , Dane Sherburn , Daniel Kappler , Daniel Levin , Daniel Levy , David Carr , David Farhi , David Mely , David Robinson , David Sasaki , Denny Jin , Dev Valladares , Dimitris Tsipras , Doug Li , Duc Phong Nguyen , Duncan Findlay , Edede Oiwoh , Edmund Wong , Ehsan Asdar , Elizabeth Proehl , Elizabeth Yang , Eric Antonow , Eric Kramer , Eric Peterson , Eric Sigler , Eric Wallace , Eugene Brevdo , Evan Mays , Farzad Khorasani , Felipe Petroski Such , Filippo Raso , Francis Zhang , Fred von Lohmann , Freddie Sulit , Gabriel Goh , Gene Oden , Geoff Salmon , Giulio Starace , Greg Brockman , Hadi Salman , Haiming Bao , Haitang Hu , Hannah Wong , Haoyu Wang , Heather Schmidt , Heather Whitney , Heewoo Jun , Hendrik Kirchner , Henrique Ponde de Oliveira Pinto , Hongyu Ren , Huiwen Chang , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian O'Connell , Ian Osband , Ian Silber , Ian Sohl , Ibrahim Okuyucu , Ikai Lan , Ilya Kostrikov , Ilya Sutskever , Ingmar Kanitscheider , Ishaan Gulrajani , Jacob Coxon , Jacob Menick , Jakub Pachocki , James Aung , James Betker , James Crooks , James Lennon , Jamie Kiros , Jan Leike , Jane Park , Jason Kwon , Jason Phang , Jason Teplitz , Jason Wei , Jason Wolfe , Jay Chen , Jeff Harris , Jenia Varavva , Jessica Gan Lee , Jessica Shieh , Ji Lin , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joanne Jang , Joaquin Quinonero Candela , Joe Beutler , Joe Landers , Joel Parish , Johannes Heidecke , John Schulman , Jonathan Lachman , Jonathan McKay , Jonathan Uesato , Jonathan Ward , Jong Wook Kim , Joost Huizinga , Jordan Sitkin , Jos Kraaijeveld , Josh Gross , Josh Kaplan , Josh Snyder , Joshua Achiam , Joy Jiao , Joyce Lee , Juntang Zhuang , Justyn Harriman , Kai Fricke , Kai Hayashi , Karan Singhal , Katy Shi , Kavin Karthik , Kayla Wood , Kendra Rimbach , Kenny Hsu , Kenny Nguyen , Keren Gu-Lemberg , Kevin Button , Kevin Liu , Kiel Howe , Krithika Muthukumar , Kyle Luther , Lama Ahmad , Larry Kai , Lauren Itow , Lauren Workman , Leher Pathak , Leo Chen , Li Jing , Lia Guy , Liam Fedus , Liang Zhou , Lien Mamitsuka , Lilian Weng , Lindsay McCallum , Lindsey Held , Long Ouyang , Louis Feuvrier , Lu Zhang , Lukas Kondraciuk , Lukasz Kaiser , Luke Hewitt , Luke Metz , Lyric Doshi , Mada Aflak , Maddie Simens , Madelaine Boyd , Madeleine Thompson , Marat Dukhan , Mark Chen , Mark Gray , Mark Hudnall , Marvin Zhang , Marwan Aljubeh , Mateusz Litwin , Matthew Zeng , Max Johnson , Maya Shetty , Mayank Gupta , Meghan Shah , Mehmet Yatbaz , Meng Jia Yang , Mengchao Zhong , Mia Glaese , Mianna Chen , Michael Janner , Michael Lampe , Michael Petrov , Michael Wu , Michele Wang , Michelle Fradin , Michelle Pokrass , Miguel Castro , Miguel Oom Temudo de Castro , Mikhail Pavlov , Miles Brundage , Miles Wang , Minal Khan , Mira Murati , Mo Bavarian , Molly Lin , Murat Yesildal , Nacho Soto , Natalia Gimelshein , Natalie Cone , Natalie Staudacher , Natalie Summers , Natan LaFontaine , Neil Chowdhury , Nick Ryder , Nick Stathas , Nick Turley , Nik Tezak , Niko Felix , Nithanth Kudige , Nitish Keskar , Noah Deutsch , Noel Bundick , Nora Puckett , Ofir Nachum , Ola Okelola , Oleg Boiko , Oleg Murk , Oliver Jaffe , Olivia Watkins , Olivier Godement , Owen Campbell-Moore , Patrick Chao , Paul McMillan , Pavel Belov , Peng Su , Peter Bak , Peter Bakkum , Peter Deng , Peter Dolan , Peter Hoeschele , Peter Welinder , Phil Tillet , Philip Pronin , Philippe Tillet , Prafulla Dhariwal , Qiming Yuan , Rachel Dias , Rachel Lim , Rahul Arora , Rajan Troll , Randall Lin , Rapha Gontijo Lopes , Raul Puri , Reah Miyara , Reimar Leike , Renaud Gaubert , Reza Zamani , Ricky Wang , Rob Donnelly , Rob Honsby , Rocky Smith , Rohan Sahai , Rohit Ramchandani , Romain Huet , Rory Carmichael , Rowan Zellers , Roy Chen , Ruby Chen , Ruslan Nigmatullin , Ryan Cheu , Saachi Jain , Sam Altman , Sam Schoenholz , Sam Toizer , Samuel Miserendino , Sandhini Agarwal , Sara Culver , Scott Ethersmith , Scott Gray , Sean Grove , Sean Metzger , Shamez Hermani , Shantanu Jain , Shengjia Zhao , Sherwin Wu , Shino Jomoto , Shirong Wu , Shuaiqi , Xia , Sonia Phene , Spencer Papay , Srinivas Narayanan , Steve Coffey , Steve Lee , Stewart Hall , Suchir Balaji , Tal Broda , Tal Stramer , Tao Xu , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Cunninghman , Thomas Degry , Thomas Dimson , Thomas Raoux , Thomas Shadwell , Tianhao Zheng , Todd Underwood , Todor Markov , Toki Sherbakov , Tom Rubin , Tom Stasi , Tomer Kaftan , Tristan Heywood , Troy Peterson , Tyce Walters , Tyna Eloundou , Valerie Qi , Veit Moeller , Vinnie Monaco , Vishal Kuo , Vlad Fomenko , Wayne Chang , Weiyi Zheng , Wenda Zhou , Wesam Manassra , Will Sheu , Wojciech Zaremba , Yash Patil , Yilei Qian , Yongjik Kim , Youlong Cheng , Yu Zhang , Yuchen He , Yuchen Zhang , Yujia Jin , Yunxing Dai , Yury Malkov

Recent studies have demonstrated promising performance of ChatGPT and GPT-4 on several medical domain tasks. However, none have assessed its performance using a large-scale real-world electronic health record database, nor have evaluated…

Computation and Language · Computer Science 2023-07-18 Jingqing Zhang , Kai Sun , Akshay Jagadeesh , Mahta Ghahfarokhi , Deepa Gupta , Ashok Gupta , Vibhor Gupta , Yike Guo

This paper assesses the potential for the large language models (LLMs) GPT-4 and GPT-3.5 to aid in deriving insight from education feedback surveys. Exploration of LLM use cases in education has focused on teaching and learning, with less…

Computation and Language · Computer Science 2024-06-28 Michael J. Parker , Caitlin Anderson , Claire Stone , YeaRim Oh

Voice assistants have become an essential tool for people with various disabilities because they enable complex phone- or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Colin Lea , Zifang Huang , Dhruv Jain , Lauren Tooley , Zeinab Liaghat , Shrinath Thelapurath , Leah Findlater , Jeffrey P. Bigham

Whisper is a multitask and multilingual speech model covering 99 languages. It yields commendable automatic speech recognition (ASR) results in a subset of its covered languages, but the model still underperforms on a non-negligible number…

Computation and Language · Computer Science 2025-12-02 Thomas Palmeira Ferraz , Marcely Zanon Boito , Caroline Brun , Vassilina Nikoulina

Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS models. Conventional machine unlearning is insufficient in this context, as zero-shot TTS can…

Speech intelligibility evaluation for hearing-impaired (HI) listeners is essential for assessing hearing aid performance, traditionally relying on listening tests or intrusive methods like HASPI. However, these methods require clean…

Sound · Computer Science 2025-09-23 Boxuan Cao , Linkai Li , Hanlin Yu , Changgeng Mo , Haoshuai Zhou , Shan Xiang Wang

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

Speech severity evaluation is becoming increasingly important as the economic burden of speech disorders grows. Current speech severity models often struggle with generalization, learning dataset-specific acoustic cues rather than…

Sound · Computer Science 2025-10-02 Bence Mark Halpern , Tomoki Toda

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been underexplored.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-20 Leying Zhang , Wangyou Zhang , Chenda Li , Yanmin Qian