English
Related papers

Related papers: Dhvani: A Weakly-supervised Phonemic Error Detecti…

200 papers

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Ruibo Fu , Tao Wang , Longbiao Wang , Jianwu Dang

Named Entity Recognition (NER) is a useful component in Natural Language Processing (NLP) applications. It is used in various tasks such as Machine Translation, Summarization, Information Retrieval, and Question-Answering systems. The…

Speech systems are sensitive to accent variations. This is especially challenging in the Indian context, with an abundance of languages but a dearth of linguistic studies characterising pronunciation variations. The growing number of L2…

Computation and Language · Computer Science 2022-12-20 Shelly Jain , Priyanshi Pal , Anil Vuppala , Prasanta Ghosh , Chiranjeevi Yarra

Script identification and text recognition are some of the major domains in the application of Artificial Intelligence. In this era of digitalization, the use of digital note-taking has become a common practice. Still, conventional methods…

Artificial Intelligence · Computer Science 2023-08-14 Sidhantha Poddar , Rohan Gupta

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-22 Xinhao Mei , Chutong Meng , Haohe Liu , Qiuqiang Kong , Tom Ko , Chengqi Zhao , Mark D. Plumbley , Yuexian Zou , Wenwu Wang

We propose a novel method that uses convolutional neural networks (CNNs) for feature extraction. Not just limited to conventional spatial domain representation, we use multilevel 2D discrete Haar wavelet transform, where image…

Computer Vision and Pattern Recognition · Computer Science 2018-01-08 Soumya Ukil , Swarnendu Ghosh , Sk Md Obaidullah , K. C. Santosh , Kaushik Roy , Nibaran Das

In today's digital world language technology has gained importance. Several softwares, have been developed and are available in the field of computational linguistics. Such tools play a crucial role in making classical language texts easily…

Computation and Language · Computer Science 2022-01-06 Swaraja Salaskar , Diptesh Kanojia , Malhar Kulkarni

While improvements have been made in automatic speech recognition performance over the last several years, machines continue to have significantly lower performance on accented speech than humans. In addition, the most significant…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-13 Xiangyun Chu , Elizabeth Combs , Amber Wang , Michael Picheny

Recent advancements in machine learning have significantly improved speech recognition, but recognizing speech from non-fluent or accented speakers remains a challenge. Previous efforts, relying on rule-based pronunciation patterns, have…

Computation and Language · Computer Science 2025-06-04 Anna Seo Gyeong Choi , Jonghyeon Park , Myungwoo Oh

In this report, we summarize the integrated multilingual audio processing pipeline developed by our team for the inaugural NCIIPC Startup India AI GRAND CHALLENGE, addressing Problem Statement 06: Language-Agnostic Speaker Identification…

Recent advancements in the field of computer vision with the help of deep neural networks have led us to explore and develop many existing challenges that were once unattended due to the lack of necessary technologies. Hand Sign/Gesture…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Shahjalal Ahmed , Md. Rafiqul Islam , Jahid Hassan , Minhaz Uddin Ahmed , Bilkis Jamal Ferdosi , Sanjay Saha , Md. Shopon

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this…

We announce the initial release of "Airavata," an instruction-tuned LLM for Hindi. Airavata was created by fine-tuning OpenHathi with diverse, instruction-tuning Hindi datasets to make it better suited for assistive tasks. Along with the…

Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of…

Computation and Language · Computer Science 2024-11-14 Shayekh Bin Islam , Ridwanul Hasan Tanvir , Sihat Afnan

With the rise of online abuse, the NLP community has begun investigating the use of neural architectures to generate counterspeech that can "counter" the vicious tone of such abusive speech and dilute/ameliorate their rippling effect over…

Computation and Language · Computer Science 2024-02-13 Mithun Das , Saurabh Kumar Pandey , Shivansh Sethi , Punyajoy Saha , Animesh Mukherjee

Graph Convolutional Networks (GCN) have achieved state-of-art results on single text classification tasks like sentiment analysis, emotion detection, etc. However, the performance is achieved by testing and reporting on resource-rich…

Computation and Language · Computer Science 2022-05-04 Mounika Marreddy , Subba Reddy Oota , Lakshmi Sireesha Vakada , Venkata Charan Chinni , Radhika Mamidi

Many consumer speech recognition systems are not tuned for people with speech disabilities, resulting in poor recognition and user experience, especially for severe speech differences. Recent studies have emphasized interest in personalized…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-12 Colin Lea , Dianna Yee , Jaya Narain , Zifang Huang , Lauren Tooley , Jeffrey P. Bigham , Leah Findlater

Computer-Assisted Pronunciation Training (CAPT) systems employ automatic measures of pronunciation quality, such as the goodness of pronunciation (GOP) metric. GOP relies on forced alignments, which are prone to labeling and segmentation…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Aditya Kamlesh Parikh , Cristian Tejedor-Garcia , Catia Cucchiarini , Helmer Strik

People commonly communicate in English, Arabic, and Bengali spoken languages through various mediums. However, deaf and hard-of-hearing individuals primarily use body language and sign language to express their needs and achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Md Hadiuzzaman , Mohammed Sowket Ali , Tamanna Sultana , Abdur Raj Shafi , Abu Saleh Musa Miah , Jungpil Shin

In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Amirsina Torfi , Jeremy Dawson , Nasser M. Nasrabadi
‹ Prev 1 8 9 10 Next ›