中文
相关论文

相关论文: A Challenging Benchmark of Anime Style Recognition

200 篇论文

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Olivier Siohan

Recent breakthroughs in Automatic Speech Recognition (ASR) have enabled fully automated Alzheimer's Disease (AD) detection using ASR transcripts. Nonetheless, the impact of ASR errors on AD detection remains poorly understood. This paper…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Yin-Long Liu , Rui Feng , Jia-Xin Chen , Yi-Ming Wang , Jia-Hong Yuan , Zhen-Hua Ling

Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to grammatical errors, disfluency, and other…

计算与语言 · 计算机科学 2020-04-10 Junwei Liao , Sefik Emre Eskimez , Liyang Lu , Yu Shi , Ming Gong , Linjun Shou , Hong Qu , Michael Zeng

High Resolution (HR) medical images provide rich anatomical structure details to facilitate early and accurate diagnosis. In MRI, restricted by hardware capacity, scan time, and patient cooperation ability, isotropic 3D HR image acquisition…

图像与视频处理 · 电气工程与系统科学 2022-12-01 Qing Wu , Yuwei Li , Yawen Sun , Yan Zhou , Hongjiang Wei , Jingyi Yu , Yuyao Zhang

Identifying individuals in unconstrained video settings is a valuable yet challenging task in biometric analysis due to variations in appearances, environments, degradations, and occlusions. In this paper, we present ShARc, a multimodal…

计算机视觉与模式识别 · 计算机科学 2023-10-25 Haidong Zhu , Wanrong Zheng , Zhaoheng Zheng , Ram Nevatia

Cosplay has grown from its origins at fan conventions into a billion-dollar global dress phenomenon. To facilitate imagination and reinterpretation from animated images to real garments, this paper presents an automatic costume image…

计算机视觉与模式识别 · 计算机科学 2020-08-27 Koya Tango , Marie Katsurai , Hayato Maki , Ryosuke Goto

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Pengcheng Guo , He Wang , Bingshen Mu , Ao Zhang , Peikun Chen

In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at…

End-to-end automatic speech recognition (ASR) systems are increasingly popular due to their relative architectural simplicity and competitive performance. However, even though the average accuracy of these systems may be high, the…

计算与语言 · 计算机科学 2021-09-14 Chao-Han Huck Yang , Linda Liu , Ankur Gandhe , Yile Gu , Anirudh Raju , Denis Filimonov , Ivan Bulyko

This paper proposes an Any-time super-Resolution Method (ARM) to tackle the over-parameterized single image super-resolution (SISR) models. Our ARM is motivated by three observations: (1) The performance of different image patches varies…

计算机视觉与模式识别 · 计算机科学 2022-07-20 Bohong Chen , Mingbao Lin , Kekai Sheng , Mengdan Zhang , Peixian Chen , Ke Li , Liujuan Cao , Rongrong Ji

Face Super-Resolution (SR) is a domain-specific super-resolution problem. The specific facial prior knowledge could be leveraged for better super-resolving face images. We present a novel deep end-to-end trainable Face Super-Resolution…

计算机视觉与模式识别 · 计算机科学 2017-11-30 Yu Chen , Ying Tai , Xiaoming Liu , Chunhua Shen , Jian Yang

Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational…

计算与语言 · 计算机科学 2024-09-19 Gaurav Maheshwari , Dmitry Ivanov , Théo Johannet , Kevin El Haddad

Automatic Speech Recognition (ASR) systems can be trained to achieve remarkable performance given large amounts of manually transcribed speech, but large labeled data sets can be difficult or expensive to acquire for all languages of…

计算与语言 · 计算机科学 2022-03-22 Hanan Aldarmaki , Asad Ullah , Nazar Zaki

Scribble colors based line art colorization is a challenging computer vision problem since neither greyscale values nor semantic information is presented in line arts, and the lack of authentic illustration-line art training pairs also…

计算机视觉与模式识别 · 计算机科学 2018-08-13 Yuanzheng Ci , Xinzhu Ma , Zhihui Wang , Haojie Li , Zhongxuan Luo

Remote sensing image retrieval(RSIR), which aims to efficiently retrieve data of interest from large collections of remote sensing data, is a fundamental task in remote sensing. Over the past several decades, there has been significant…

计算机视觉与模式识别 · 计算机科学 2018-07-24 Weixun Zhou , Shawn Newsam , Congmin Li , Zhenfeng Shao

Language understanding in speech-based systems have attracted much attention in recent years with the growing demand for voice interface applications. However, the robustness of natural language understanding (NLU) systems to errors…

计算与语言 · 计算机科学 2022-03-17 Lingyun Feng , Jianwei Yu , Deng Cai , Songxiang Liu , Haitao Zheng , Yan Wang

With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image-voice retrieval provides a new insight. This paper aims to…

多媒体 · 计算机科学 2022-01-05 Hailong Ning , Bin Zhao , Yuan Yuan

Single image super resolution (SISR) is to reconstruct a high resolution image from a single low resolution image. The SISR task has been a very attractive research topic over the last two decades. In recent years, convolutional neural…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Bingzhe Wu , Haodong Duan , Zhichao Liu , Guangyu Sun

Common measures of accuracy used to assess the performance of automatic speech recognition (ASR) systems, as well as human transcribers, conflate multiple sources of error. Stylistic differences, such as verbatim vs non-verbatim, can play a…

计算与语言 · 计算机科学 2024-09-06 Annika Heuser , Tyler Kendall , Miguel del Rio , Quinten McNamara , Nishchal Bhandari , Corey Miller , Migüel Jetté

Reference-based Super Resolution (RefSR) improves upon Single Image Super Resolution (SISR) by leveraging high-quality reference images to enhance texture fidelity and visual realism. However, a critical limitation of existing RefSR…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Jiaqi Yan , Shuning Xu , Xiangyu Chen , Dell Zhang , Jiantao Zhou , Jie Tang , Gangshan Wu , Jie Liu