English
Related papers

Related papers: A DenseNet-based method for decoding auditory spat…

200 papers

TASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Kyle Min , Jason J. Corso

We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs a gated…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Linya Fu , Yu Liu , Zhijie Liu , Zedong Yang , Zhong-Qiu Wang , Youfu Li , He Kong

Ultrasound imaging is a prevalent diagnostic tool known for its simplicity and non-invasiveness. However, its inherent characteristics often introduce substantial noise, posing considerable challenges for automated lesion or organ…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Ling Zhou , Runtian Yuan , Yi Liu , Yuejie Zhang , Rui Feng , Shang Gao

In a multi-speaker "cocktail party" scenario, a listener can selectively attend to a speaker of interest. Studies into the human auditory attention network demonstrate cortical entrainment to speech envelopes resulting in highly correlated…

Signal Processing · Electrical Eng. & Systems 2023-07-18 Richard Gall , Deniz Kocanaogullari , Murat Akcakaya , Deniz Erdogmus , Rajkumar Kubendran

Breast ultrasound imaging is a valuable tool for early breast cancer detection, but automated tumor segmentation is challenging due to inherent noise, variations in scale of lesions, and fuzzy boundaries. To address these challenges, we…

Image and Video Processing · Electrical Eng. & Systems 2025-06-23 Muhammad Azeem Aslam , Asim Naveed , Nisar Ahmed

Auditory attention is a selective type of hearing in which people focus their attention intentionally on a specific source of a sound or spoken words whilst ignoring or inhibiting other auditory stimuli. In some sense, the auditory…

Machine Learning · Computer Science 2021-10-26 Mahak Kothari , Shreyansh Joshi , Adarsh Nandanwar , Aadetya Jaiswal , Veeky Baths

Carrying conversations in multi-sound environments is one of the more challenging tasks, since the sounds overlap across time and frequency making it difficult to understand a single sound source. One proposed approach to help isolate an…

Machine Learning · Computer Science 2024-10-25 Seyed Ali Alavi Bajestan , Mark Pitt , Donald S. Williamson

We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to encode absolute frame positions by exploiting limited acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-17 Pawel Swietojanski , Xinwei Li , Mingbin Xu , Takaaki Hori , Dogan Can , Xiaodan Zhuang

Recent advances in reconstructing speech envelopes from Electroencephalogram (EEG) signals have enabled continuous auditory attention decoding (AAD) in multi-speaker environments. Most Deep Neural Network (DNN)-based envelope reconstruction…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Yayun Liang , Yuanming Zhang , Fei Chen , Jing Lu , Zhibin Lin

We propose a novel method for Acoustic Event Detection (AED). In contrast to speech, sounds coming from acoustic events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time…

Sound · Computer Science 2016-12-09 Naoya Takahashi , Michael Gygli , Beat Pfister , Luc Van Gool

The decoding of electroencephalography (EEG) signals allows access to user intentions conveniently, which plays an important role in the fields of human-machine interaction. To effectively extract sufficient characteristics of the…

Human-Computer Interaction · Computer Science 2024-09-06 Hongqi Li , Haodong Zhang , Yitong Chen

Emotion recognition based on EEG (electroencephalography) has been widely used in human-computer interaction, distance education and health care. However, the conventional methods ignore the adjacent and symmetrical characteristics of EEG…

Signal Processing · Electrical Eng. & Systems 2021-08-30 Xiangwen Deng , Junlin Zhu , Shangming Yang

Speech enhancement is widely used as a front-end to improve the speech quality in many audio systems, while it is hard to extract the target speech in multi-talker conditions without prior information on the speaker identity. It was shown…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-26 Jie Zhang , Qing-Tian Xu , Zhen-Hua Ling , Haizhou Li

Electroencephalogram (EEG)-based emotion decoding can objectively quantify people's emotional state and has broad application prospects in human-computer interaction and early detection of emotional disorders. Recently emerging deep…

Human-Computer Interaction · Computer Science 2024-11-08 Xinke Shen , Runmin Gan , Kaixuan Wang , Shuyi Yang , Qingzhu Zhang , Quanying Liu , Dan Zhang , Sen Song

Lung cancer has been one of the major threats across the world with the highest mortalities. Computer-aided detection (CAD) can help in early detection and thus can help increase the survival rate. Accurate lung parenchyma segmentation (to…

Image and Video Processing · Electrical Eng. & Systems 2025-09-18 Muhammad Abdullah , Furqan Shaukat

Image super-resolution is a challenging task and has attracted increasing attention in research and industrial communities. In this paper, we propose a novel end-to-end Attention-based DenseNet with Residual Deconvolution named as ADRD. In…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Zhuangzi Li

Recognition of electroencephalographic (EEG) signals highly affect the efficiency of non-invasive brain-computer interfaces (BCIs). While recent advances of deep-learning (DL)-based EEG decoders offer improved performances, the development…

Machine Learning · Computer Science 2022-10-06 Yue-Ting Pan , Jing-Lun Chou , Chun-Shu Wei

Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is…

Multimedia · Computer Science 2025-06-24 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-08 Davide Berghi , Adrian Hilton , Philip J. B. Jackson

Student attention is an indispensable input for uncovering their goals, intentions, and interests, which prove to be invaluable for a multitude of research areas, ranging from psychology to interactive systems. However, most existing…

Human-Computer Interaction · Computer Science 2023-11-07 Dhruv Verma , Sejal Bhalla , S. V. Sai Santosh , Saumya Yadav , Aman Parnami , Jainendra Shukla