中文
相关论文

相关论文: Extracting Different Levels of Speech Information …

200 篇论文

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Prashanth Gurunath Shivakumar , Jari Kolehmainen , Aditya Gourav , Yi Gu , Ankur Gandhe , Ariya Rastrow , Ivan Bulyko

Electroencephalography (EEG) is widely used to study human brain dynamics, yet its quantitative information capacity remains unclear. Here, we combine information theory and synthetic forward modeling to estimate the mutual information…

信息论 · 计算机科学 2025-10-22 Ishir Rao

Decoding the human brain has been a hallmark of neuroscientists and Artificial Intelligence researchers alike. Reconstruction of visual images from brain Electroencephalography (EEG) signals has garnered a lot of interest due to its…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Prajwal Singh , Dwip Dalal , Gautam Vashishtha , Krishna Miyapuram , Shanmuganathan Raman

Speech cognition bears potential application as a brain computer interface that can improve the quality of life for the otherwise communication impaired people. While speech and resting state EEG are popularly studied, here we attempt to…

机器学习 · 计算机科学 2020-10-13 Rini A Sharon , Hema A Murthy

The aim of this paper is to investigate the benefit of combining both language and acoustic modelling for speaker diarization. Although conventional systems only use acoustic features, in some scenarios linguistic data contain high…

音频与语音处理 · 电气工程与系统科学 2025-01-31 Miquel India , Javier Hernando , José A. R. Fonollosa

For EEG-based drowsiness recognition, it is desirable to use subject-independent recognition since conducting calibration on each subject is time-consuming. In this paper, we propose a novel Convolutional Neural Network (CNN)-Long…

神经与进化计算 · 计算机科学 2021-12-22 Jian Cui , Zirui Lan , Tianhu Zheng , Yisi Liu , Olga Sourina , Lipo Wang , Wolfgang Müller-Wittig

In recent years there have been many deep learning approaches towards the multi-speaker source separation problem. Most use Long Short-Term Memory - Recurrent Neural Networks (LSTM-RNN) or Convolutional Neural Networks (CNN) to model the…

机器学习 · 计算机科学 2019-12-20 Jeroen Zegers , Hugo Van hamme

Speech Emotion Recognition (SER) affective technology enables the intelligent embedded devices to interact with sensitivity. Similarly, call centre employees recognise customers' emotions from their pitch, energy, and tone of voice so as to…

声音 · 计算机科学 2023-12-19 David Hason Rudd , Huan Huo , Guandong Xu

Hierarchical Multiscale LSTM (Chung et al., 2016a) is a state-of-the-art language model that learns interpretable structure from character-level input. Such models can provide fertile ground for (cognitive) computational linguistics…

计算与语言 · 计算机科学 2018-07-11 Ákos Kádár , Marc-Alexandre Côté , Grzegorz Chrupała , Afra Alishahi

In this paper, we are interested in exploiting textual and acoustic data of an utterance for the speech emotion classification task. The baseline approach models the information from audio and text independently using two deep neural…

音频与语音处理 · 电气工程与系统科学 2019-12-02 Seunghyun Yoon , Seokhyun Byun , Subhadeep Dey , Kyomin Jung

Language models (LMs) for text data have been studied extensively for their usefulness in language generation and other downstream tasks. However, language modelling purely in the speech domain is still a relatively unexplored topic, with…

计算与语言 · 计算机科学 2021-11-02 Anurag Katakkar , Alan W Black

Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, \emph{e.g.,} band power, connectivity, complexity, and more. Modern EEG foundation models bypass this catalog, learn directly from…

人工智能 · 计算机科学 2026-05-15 Ling Tang , Qian Chen , Jilin Mei , Houshi Xu , Quanshi Zhang , Jing Shao , Na Zou , Xia Hu , Dongrui Liu

In this paper, we propose a novel architecture for direct extractive speech-to-speech summarization, ESSumm, which is an unsupervised model without dependence on intermediate transcribed text. Different from previous methods with text…

音频与语音处理 · 电气工程与系统科学 2022-09-16 Jun Wang

Decoding visual representations from human brain activity has emerged as a thriving research domain, particularly in the context of brain-computer interfaces. Our study presents an innovative method that employs to classify and reconstruct…

信号处理 · 电气工程与系统科学 2023-09-15 Matteo Ferrante , Tommaso Boccato , Stefano Bargione , Nicola Toschi

Due to the limitations in the accuracy and robustness of current electroencephalogram (EEG) classification algorithms, applying motor imagery (MI) for practical Brain-Computer Interface (BCI) applications remains challenging. This paper…

人机交互 · 计算机科学 2023-12-21 Shiwei Cheng , Yuejiang Hao

Being able to analyze and interpret signal coming from electroencephalogram (EEG) recording can be of high interest for many applications including medical diagnosis and Brain-Computer Interfaces. Indeed, human experts are today able to…

人工智能 · 计算机科学 2007-05-23 Nizar Kerkeni , Frederic Alexandre , Mohamed Hedi Bedoui , Laurent Bougrain , Mohamed Dogui

This work introduces approaches to assessing phrase breaks in ESL learners' speech using pre-trained language models (PLMs) and large language models (LLMs). There are two tasks: overall assessment of phrase break for a speech clip and…

计算与语言 · 计算机科学 2023-06-09 Zhiyi Wang , Shaoguang Mao , Wenshan Wu , Yan Xia , Yan Deng , Jonathan Tien

The representations generated by many models of language (word embeddings, recurrent neural networks and transformers) correlate to brain activity recorded while people read. However, these decoding results are usually based on the brain's…

计算与语言 · 计算机科学 2020-10-16 Maryam Hashemzadeh , Greta Kaufeld , Martha White , Andrea E. Martin , Alona Fyshe

Previous initial research has already been carried out to propose speech-based BCI using brain signals (e.g. non-invasive EEG and invasive sEEG / ECoG), but there is a lack of combined methods that investigate non-invasive brain,…

医学物理 · 物理学 2023-10-19 Tamás Gábor Csapó , Frigyes Viktor Arthur , Péter Nagy , Ádám Boncz

Recognizing emotional signals in speech has a significant impact on enhancing the effectiveness of human-computer interaction (HCI). This study introduces EmoAugNet, a hybrid deep learning framework, that incorporates Long Short-Term Memory…

声音 · 计算机科学 2025-08-11 Durjoy Chandra Paul , Gaurob Saha , Md Amjad Hossain
‹ 上一页 1 8 9 10 下一页 ›