English
Related papers

Related papers: Human Silhouette and Skeleton Video Synthesis thro…

200 papers

Video capture is the most extensively utilized human perception source due to its intuitively understandable nature. A desired video capture often requires multiple environmental conditions such as ample ambient-light, unobstructed space,…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Xiang Li , Rabih Younes

Radio frequency (RF) signals have been proved to be flexible for human silhouette segmentation (HSS) under complex environments. Existing studies are mainly based on a one-shot approach, which lacks a coherent projection ability from the RF…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Penghui Wen , Kun Hu , Dong Yuan , Zhiyuan Ning , Changyang Li , Zhiyong Wang

Human silhouette segmentation, which is originally defined in computer vision, has achieved promising results for understanding human activities. However, the physical limitation makes existing systems based on optical cameras suffer from…

Computer Vision and Pattern Recognition · Computer Science 2022-01-26 Zhi Wu , Dongheng Zhang , Chunyang Xie , Cong Yu , Jinbo Chen , Yang Hu , Yan Chen

Wi-Fi channel state information (CSI) has emerged as a plausible modality for sensing different human activities as a function of modulations in the wireless signal that travels between wireless devices. Until now, most research has taken a…

Information Theory · Computer Science 2019-02-05 Mohammed Alloulah , Anton Isopoussu , Chulhong Min , Fahim Kawsar

Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to help the visually impaired people to better perceive the…

Sound · Computer Science 2021-03-19 Hailong Ning , Xiangtao Zheng , Yuan Yuan , Xiaoqiang Lu

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Ultrasound (US) is widely used for its advantages of real-time imaging, radiation-free and portability. In clinical practice, analysis and diagnosis often rely on US sequences rather than a single image to obtain dynamic anatomical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Jiamin Liang , Xin Yang , Yuhao Huang , Kai Liu , Xinrui Zhou , Xindi Hu , Zehui Lin , Huanjia Luo , Yuanji Zhang , Yi Xiong , Dong Ni

This paper demonstrates human synthesis based on the Radio Frequency (RF) signals, which leverages the fact that RF signals can record human movements with the signal reflections off the human body. Different from existing RF sensing works…

Multimedia · Computer Science 2021-12-08 Cong Yu , Zhi Wu , Dongheng Zhang , Zhi Lu , Yang Hu , Yan Chen

Wi-Fi sensing is gaining momentum as a non-intrusive and privacy-preserving alternative to vision-based systems for human identification. However, person identification through wireless signals, particularly without user motion, remains…

In this paper, we propose a novel approach to convert given speech audio to a photo-realistic speaking video of a specific person, where the output video has synchronized, realistic, and expressive rich body dynamics. We achieve this by…

Computer Vision and Pattern Recognition · Computer Science 2020-10-12 Miao Liao , Sibo Zhang , Peng Wang , Hao Zhu , Xinxin Zuo , Ruigang Yang

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

Human activity recognition (HAR) has been playing an increasingly important role in various domains such as healthcare, security monitoring, and metaverse gaming. Though numerous HAR methods based on computer vision have been developed to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-01 Jianfei Yang , Shijie Tang , Yuecong Xu , Yunjiao Zhou , Lihua Xie

WiFi channel state information (CSI) has emerged as a plausible modality for sensing different human vital signs, i.e. respiration and body motion, as a function of modulated wireless signals that travel between WiFi devices. Although a…

Signal Processing · Electrical Eng. & Systems 2020-03-23 Kamran Ali , Mohammed Alloulah , Fahim Kawsar , Alex X. Liu

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

Sound · Computer Science 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

While fulfilling communication tasks, wireless signals can also be used to sense the environment. Among various types of sensing media, WiFi signals offer advantages such as widespread availability, low hardware cost, and strong robustness…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Ruijing Liu , Cunhua Pan , Jiaming Zeng , Hong Ren , Kezhi Wang , Lei Kong , Jiangzhou Wang

Objects in an environment affect electromagnetic waves. While this effect varies across frequencies, there exists a correlation between them, and a model with enough capacity can capture this correlation between the measurements in…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Mohammad Hadi Kefayati , Vahid Pourahmadi , Hassan Aghaeinia

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Ji-Hoon Kim , Jeongsoo Choi , Jaehun Kim , Chaeyoung Jung , Joon Son Chung

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Super-resolution (SR) aims to enhance the quality of low-resolution images and has been widely applied in medical imaging. We found that the design principles of most existing methods are influenced by SR tasks based on real-world images…

Image and Video Processing · Electrical Eng. & Systems 2025-11-14 Feiyang Jia , Zhineng Chen , Ziying Song , Lin Liu , Caiyan Jia
‹ Prev 1 2 3 10 Next ›