English
Related papers

Related papers: Sensing the Breath: A Multimodal Singing Tutoring …

200 papers

The surging popularity of home assistants and their voice user interface (VUI) have made them an ideal central control hub for smart home devices. However, current form factors heavily rely on VUI, which poses accessibility and usability…

Sound · Computer Science 2025-09-05 Yin Li , Rohan Reddy , Cheng Zhang , Rajalakshmi Nandakumar

In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice. Our approach is based on determining if the lips motion and the voice in a video are synchronised or not, depending on their…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Venkatesh S. Kadandale , Juan F. Montesinos , Gloria Haro

As communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack…

Human-Computer Interaction · Computer Science 2023-09-12 Zeyuan Huang , Qiang He , Kevin Maher , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Sheng-feng Qin , Yong-Jin Liu , Hongan Wang

Robotic assistance presents an opportunity to benefit the lives of many people with physical disabilities, yet accurately sensing the human body and tracking human motion remain difficult for robots. We present a multidimensional capacitive…

Robotics · Computer Science 2019-05-28 Zackory Erickson , Henry M. Clever , Vamsee Gangaram , Greg Turk , C. Karen Liu , Charles C. Kemp

This paper presents the soma design process of creating Body Electric: a novel interface for the capture and use of biofeedback signals and physiological changes generated in the body by breathing, during singing. This NIME design is…

Human-Computer Interaction · Computer Science 2023-11-15 Kelsey Cotton , Pedro Sanches , Vasiliki Tsaknaki , Pavel Karpashevich

In this project, we have developed a sign language tutor that lets users learn isolated signs by watching recorded videos and by trying the same signs. The system records the user's video and analyses it. If the sign is recognized, both…

We propose MoodNet - A Deep Convolutional Neural Network based architecture to effectively predict the emotion associated with a piece of music given its audio and lyrical content.We evaluate different architectures consisting of varying…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-15 Aniruddha Bhattacharya , K. V. Kadambari

Breathing rate is a vital health metric that is an invaluable indicator of the overall health of a person. In recent years, the non-contact measurement of health signals such as breathing rate has been a huge area of development, with a…

Image and Video Processing · Electrical Eng. & Systems 2023-11-16 Robyn Maxwell , Timothy Hanley , Dara Golden , Adara Andonie , Joseph Lemley , Ashkan Parsi

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

This paper presents the development of a multi-sensor user interface to facilitate the instruction of arc welding tasks. Traditional methods to acquire hand-eye coordination skills are typically conducted through one-to-one instruction…

Human-Computer Interaction · Computer Science 2022-06-29 Hoi-Yin Lee , Peng Zhou , Anqing Duan , Jiangliu Wang , Victor Wu , David Navarro-Alarcon

In the Massive Open Online Courses (MOOC) learning scenario, the semantic information of instructional videos has a crucial impact on learners' emotional state. Learners mainly acquire knowledge by watching instructional videos, and the…

Multimedia · Computer Science 2024-04-12 Yuan Zhang , Xiaomei Tao , Hanxu Ai , Tao Chen , Yanling Gan

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-26 Yin-Jyun Luo , Chin-Chen Hsu , Kat Agres , Dorien Herremans

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

In this paper, we touch on the problem of markerless multi-modal human motion capture especially for string performance capture which involves inherently subtle hand-string contacts and intricate movements. To fulfill this goal, we first…

In this paper, we study an approach to multimodal person verification using audio, visual, and thermal modalities. The combination of audio and visual modalities has already been shown to be effective for robust person verification. From…

Computer Vision and Pattern Recognition · Computer Science 2022-03-07 Madina Abdrakhmanova , Saniya Abushakimova , Yerbolat Khassanov , Huseyin Atakan Varol

The most common sensing modalities found in a robot perception system are vision and touch, which together can provide global and highly localized data for manipulation. However, these sensing modalities often fail to adequately capture the…

Robotics · Computer Science 2022-04-20 Jessica Yin , Gregory M. Campbell , James Pikul , Mark Yim

Robot-assisted neurological surgery is receiving growing interest due to the improved dexterity, precision, and control of surgical tools, which results in better patient outcomes. However, such systems often limit surgeons' natural sensory…

Signal Processing · Electrical Eng. & Systems 2025-08-13 Zacharias Chen , Alexa Cristelle Cahilig , Sarah Dias , Prithu Kolar , Ravi Prakash , Patrick J. Codd

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

Respiratory rate (RR) monitoring is integral to understanding physical and mental health and tracking fitness. Existing studies have demonstrated the feasibility of RR monitoring under specific user conditions (e.g., while remaining still,…

Human-Computer Interaction · Computer Science 2024-07-10 Yang Liu , Kayla-Jade Butkow , Jake Stuchbury-Wass , Adam Pullin , Dong Ma , Cecilia Mascolo

Singing voice synthesis (SVS) is a task that aims to generate audio signals according to musical scores and lyrics. With its multifaceted nature concerning music and language, producing singing voices indistinguishable from that of human…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-07 Yin-Ping Cho , Fu-Rong Yang , Yung-Chuan Chang , Ching-Ting Cheng , Xiao-Han Wang , Yi-Wen Liu