English
Related papers

Related papers: Analyzing of MOS and Codec Selection for Voice ove…

200 papers

Speech synthesis quality prediction has made remarkable progress with the development of supervised and self-supervised learning (SSL) MOS predictors but some aspects related to the data are still unclear and require further study. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-27 Alessandro Ragano , Emmanouil Benetos , Michael Chinen , Helard B. Martinez , Chandan K. A. Reddy , Jan Skoglund , Andrew Hines

Universities and Institutions these days' deals with issues related to with assessment of large number of students. Various evaluation methods have been adopted by examiners in different institutions to examining the ability of an…

Software Engineering · Computer Science 2012-02-17 Ahmed Barnawi , Abdulrahman H. Altalhi , M. Rizwan Jameel Qureshi , Asif Irshad Khan

Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-04 Jozef Coldenhoff , Niclas Granqvist , Milos Cernak

Network functions virtualization provides opportunities to design, deploy, and manage networking services. It utilizes cloud computing virtualization services that run on high-volume servers, switches, and storage hardware to virtualize…

Networking and Internet Architecture · Computer Science 2017-10-02 Ahmadreza Montazerolghaem , Mohammad Hossein Yaghmaee , Alberto Leon-Garcia , Mahmoud Naghibzadeh , Farzad Tashtarian

We present the CodecMOS-Accent dataset, a mean opinion score (MOS) benchmark designed to evaluate neural audio codec (NAC) models and the large language model (LLM)-based text-to-speech (TTS) models trained upon them, especially across…

Sound · Computer Science 2026-03-17 Wen-Chin Huang , Nicholas Sanders , Erica Cooper

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep…

Computation and Language · Computer Science 2016-11-29 Brian Patton , Yannis Agiomyrgiannakis , Michael Terry , Kevin Wilson , Rif A. Saurous , D. Sculley

This paper considers a cell-free massive multipleinput multiple-output (MIMO) integrated sensing and communication (ISAC) system, where distributed MIMO access points (APs) are used to jointly serve the communication users and detect the…

Information Theory · Computer Science 2023-10-16 Mohamed Elfiatoure , Mohammadali Mohammadi , Hien Quoc Ngo , Michail Matthaiou

In many VoIP systems, Voice Activity Detection (VAD) is often used on VoIP traffic to suppress packets of silence in order to reduce the bandwidth consumption of phone calls. Unfortunately, although VoIP traffic is fully encrypted and…

Cryptography and Security · Computer Science 2023-06-02 Roy Laurens , Edo Christianto , Bruce Caulkins , Cliff C. Zou

In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the…

Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Sewade Ogun , Vincent Colotte , Emmanuel Vincent

In this paper, we propose a novel approach for accurate detection of the vowel onset points (VOPs). VOP is the instant at which the vowel begins in the speech signal. Precise identification of VOPs is important for various speech…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-26 Kumud Tripathi , K. Sreenivasa Rao

Since 2016, all of four major U.S. operators have rolled out nationwide Wi-Fi calling services. They are projected to surpass VoLTE (Voice over LTE) and other VoIP services in terms of mobile IP voice usage minutes in 2018. They enable…

Cryptography and Security · Computer Science 2018-11-30 Tian Xie , Guan-Hua Tu , Bangjie Yin , Chi-Yu Li , Chunyi Peng , Mi Zhang , Hui Liu , Xiaoming Liu

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the…

This Ph.D. thesis focuses on developing a system for high-quality speech synthesis and voice conversion. Vocoder-based speech analysis, manipulation, and synthesis plays a crucial role in various kinds of statistical parametric speech…

Sound · Computer Science 2021-01-26 Mohammed Salah Al-Radhi

Teaching is one of the most important factors affecting any education system. Many research efforts have been conducted to facilitate the presentation modes used by instructors in classrooms as well as provide means for students to review…

Machine Learning · Computer Science 2012-01-16 Marian George , Moustafa Youssef

Turn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world…

Robotics · Computer Science 2025-07-15 Koji Inoue , Yuki Okafuji , Jun Baba , Yoshiki Ohira , Katsuya Hyodo , Tatsuya Kawahara

Wimax (Worldwide Interoperability for Microwave Access) is a promising technology which can offer high speed voice, video and data service up to the customer end. The aim of this paper is the performance evaluation of an Wimax system under…

Performance · Computer Science 2009-10-06 Md. Ashraful Islam , Riaz Uddin Mondal , Md. Zahid Hasan

An efficient speech to text converter for mobile application is presented in this work. The prime motive is to formulate a system which would give optimum performance in terms of complexity, accuracy, delay and memory requirements for…

Computation and Language · Computer Science 2013-07-23 R. Sandanalakshmi , P. Abinaya Viji , M. Kiruthiga , M. Manjari , M. Sharina

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

The proposed model aims to develop a speech recognition technology for hearing, speech, or cognitively disabled people. All the available technology in the field of speech recognition doesn't come with an interface for communication for…

Sound · Computer Science 2024-07-23 Ritabrata Roy Choudhury