English
Related papers

Related papers: Attention-Based Acoustic Feature Fusion Network fo…

200 papers

Convolutional neural networks (CNNs) and their variations have shown effectiveness in facial expression recognition (FER). However, they face challenges when dealing with high computational complexity and multi-view head poses in real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Ali Ezati , Mohammadreza Dezyani , Rajib Rana , Roozbeh Rajabi , Ahmad Ayatollahi

In acoustic scene classification (ASC), acoustic features play a crucial role in the extraction of scene information, which can be stored over different time scales. Moreover, the limited size of the dataset may lead to a biased model with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Hangting Chen , Zuozhen Liu , Zongming Liu , Pengyuan Zhang

Music emotion recognition (MER), a sub-task of music information retrieval (MIR), has developed rapidly in recent years. However, the learning of affect-salient features remains a challenge. In this paper, we propose an end-to-end…

Sound · Computer Science 2022-07-01 Zi Huang , Shulei Ji , Zhilan Hu , Chuangjian Cai , Jing Luo , Xinyu Yang

With the acceleration of the pace of work and life, people have to face more and more pressure, which increases the possibility of suffering from depression. However, many patients may fail to get a timely diagnosis due to the serious…

Signal Processing · Electrical Eng. & Systems 2021-06-02 Lang He , Mingyue Niu , Prayag Tiwari , Pekka Marttinen , Rui Su , Jiewei Jiang , Chenguang Guo , Hongyu Wang , Songtao Ding , Zhongmin Wang , Wei Dang , Xiaoying Pan

Medical image segmentation, a crucial task in computer vision, facilitates the automated delineation of anatomical structures and pathologies, supporting clinicians in diagnosis, treatment planning, and disease monitoring. Notably,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Fuchen Zheng , Xinyi Chen , Xuhang Chen , Haolun Li , Xiaojiao Guo , Weihuang Liu , Chi-Man Pun , Shoujun Zhou

We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an…

Multimedia · Computer Science 2017-01-05 Naoya Takahashi , Michael Gygli , Luc Van Gool

Intelligent monitoring systems and affective computing applications have emerged in recent years to enhance healthcare. Examples of these applications include assessment of affective states such as Major Depressive Disorder (MDD). MDD…

Human-Computer Interaction · Computer Science 2020-11-19 Alice Othmani , Daoud Kadoch , Kamil Bentounes , Emna Rejaibi , Romain Alfred , Abdenour Hadid

Twitter is currently a popular online social media platform which allows users to share their user-generated content. This publicly-generated user data is also crucial to healthcare technologies because the discovered patterns would hugely…

Machine Learning · Computer Science 2021-05-25 Hamad Zogan , Imran Razzak , Shoaib Jameel , Guandong Xu

In this paper, the dual-optical attention fusion crowd head point counting model (TAPNet) is proposed to address the problem of the difficulty of accurate counting in complex scenes such as crowd dense occlusion and low light in crowd…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Fei Zhou , Yi Li , Mingqing Zhu

We present our preliminary work to determine if patient's vocal acoustic, linguistic, and facial patterns could predict clinical ratings of depression severity, namely Patient Health Questionnaire depression scale (PHQ-8). We proposed a…

Computer Vision and Pattern Recognition · Computer Science 2017-12-01 Aven Samareh , Yan Jin , Zhangyang Wang , Xiangyu Chang , Shuai Huang

Semantic segmentation of high-resolution remote sensing images plays a crucial role in land-use monitoring and urban planning. Recent remarkable progress in deep learning-based methods makes it possible to generate satisfactory segmentation…

Image and Video Processing · Electrical Eng. & Systems 2025-04-04 Feng Gao , Miao Fu , Jingchao Cao , Junyu Dong , Qian Du

Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Designing a lightweight…

Sound · Computer Science 2024-01-26 Samuel Pegg , Kai Li , Xiaolin Hu

With the increasing popularity of convolutional neural networks (CNNs), recent works on face-based age estimation employ these networks as the backbone. However, state-of-the-art CNN-based methods treat each facial region equally, thus…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Haoyi Wang , Victor Sanchez , Chang-Tsun Li

Limited accessibility to neurological care leads to underdiagnosed Parkinson's Disease (PD), preventing early intervention. Existing AI-based PD detection methods primarily focus on unimodal analysis of motor or speech tasks, overlooking…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Md Saiful Islam , Tariq Adnan , Jan Freyberg , Sangwu Lee , Abdelrahman Abdelkader , Meghan Pawlik , Cathe Schwartz , Karen Jaffe , Ruth B. Schneider , E Ray Dorsey , Ehsan Hoque

Recent studies show that depression can be partially reflected from human facial attributes. Since facial attributes have various data structure and carry different information, existing approaches fail to specifically consider the optimal…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Mingzhe Chen , Xi Xiao , Bin Zhang , Xinyu Liu , Runiu Lu

Recently, deep convolutional neural network (CNN) have been widely used in image restoration and obtained great success. However, most of existing methods are limited to local receptive field and equal treatment of different types of…

Image and Video Processing · Electrical Eng. & Systems 2021-01-26 Yucheng Hang , Qingmin Liao , Wenming Yang , Yupeng Chen , Jie Zhou

During psychiatric assessment, clinicians observe not only what patients report, but important nonverbal signs such as tone, speech rate, fluency, responsiveness, and body language. Weighing and integrating these different information…

Machine Learning · Computer Science 2025-12-19 Agnes Norbury , George Fairs , Alexandra L. Georgescu , Matthew M. Nour , Emilia Molimpakis , Stefano Goria

In this paper, we present a novel deep fusion architecture for audio classification tasks. The multi-channel model presented is formed using deep convolution layers where different acoustic features are passed through each channel. To…

Sound · Computer Science 2018-11-05 Gaurav Bhatt , Akshita Gupta , Aditya Arora , Balasubramanian Raman

Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone…

Sound · Computer Science 2025-07-08 Nhan Duc Thanh Nguyen , Huy Phan , Simon Geirnaert , Kaare Mikkelsen , Preben Kidmose

Depression is one of the most common mental illness problems, and the symptoms shown by patients are not consistent, making it difficult to diagnose in the process of clinical practice and pathological research. Although researchers hope…

Computers and Society · Computer Science 2024-10-08 Xiaohang Xu , Hao Peng , Lichao Sun , Md Zakirul Alam Bhuiyan , Lianzhong Liu , Lifang He