中文
相关论文

相关论文: A Challenging Benchmark of Anime Style Recognition

200 篇论文

Encouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with few-shot examples. However, this feature sharing mechanism…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Binghao Liu , Yao Ding , Jianbin Jiao , Xiangyang Ji , Qixiang Ye

It is well known that many machine learning systems demonstrate bias towards specific groups of individuals. This problem has been studied extensively in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR).…

音频与语音处理 · 电气工程与系统科学 2021-11-22 Chunxi Liu , Michael Picheny , Leda Sarı , Pooja Chitkara , Alex Xiao , Xiaohui Zhang , Mark Chou , Andres Alvarado , Caner Hazirbas , Yatharth Saraf

Automatic speech recognition (ASR) allows a natural and intuitive interface for robotic educational applications for children. However there are a number of challenges to overcome to allow such an interface to operate robustly in realistic…

Recent advances in image generation, particularly diffusion models, have significantly lowered the barrier for creating sophisticated forgeries, making image manipulation detection and localization (IMDL) increasingly challenging. While…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Chenyang Zhu , Xing Zhang , Yuyang Sun , Ching-Chun Chang , Isao Echizen

End-to-end automatic speech recognition (ASR) can achieve promising performance with large-scale training data. However, it is known that domain mismatch between training and testing data often leads to a degradation of recognition…

声音 · 计算机科学 2021-06-10 Wenxin Hou , Jindong Wang , Xu Tan , Tao Qin , Takahiro Shinozaki

This paper observes the application of the Compressive Sensing in reconstruction of the under-sampled iris images. Iris recognition represents form of biometric identification whose usage in real applications is growing. Compressive Sensing…

图像与视频处理 · 电气工程与系统科学 2019-02-11 Radoje Darmanovic , Tamara Bulatovic , Seid Salkovic

Although the deep integration of the Automatic Speech Recognition (ASR) system with Large Language Models (LLMs) has significantly improved accuracy, the deployment of such systems in low-latency streaming scenarios remains challenging. In…

声音 · 计算机科学 2026-03-13 Yinfeng Xia , Jian Tang , Junfeng Hou , Gaopeng Xu , Haitao Yao

We show how to learn a map that takes a content code, derived from a face image, and a randomly chosen style code to an anime image. We derive an adversarial loss from our simple and effective definitions of style and content. This…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Min Jin Chong , David Forsyth

\textbf{Objectives}: We aimed to investigate how errors from automatic speech recognition (ASR) systems affect dementia classification accuracy, specifically in the ``Cookie Theft'' picture description task. We aimed to assess whether…

计算与语言 · 计算机科学 2024-01-12 Changye Li , Weizhe Xu , Trevor Cohen , Serguei Pakhomov

In recent years research has been producing an important effort to encode the digital image content. Most of the adopted paradigms only focus on local features and lack in information about location and relationships between them. To fill…

图像与视频处理 · 电气工程与系统科学 2021-07-14 Mario Manzo , Simone Pellino

Fine-tuning pretrained language models (LMs) is a popular approach to automatic speech recognition (ASR) error detection during post-processing. While error detection systems often take advantage of statistical language archetypes captured…

计算与语言 · 计算机科学 2021-08-05 Seongmin Park , Dongchan Shin , Sangyoun Paik , Subong Choi , Alena Kazakova , Jihwa Lee

Abstract Visual Reasoning (AVR) problems are commonly used to approximate human intelligence. They test the ability of applying previously gained knowledge, experience and skills in a completely new setting, which makes them particularly…

人工智能 · 计算机科学 2023-02-27 Mikołaj Małkiński , Jacek Mańdziuk

Automatic speech recognition (ASR) systems generate real-time transcriptions but often miss nuances that human interpreters capture. While ASR is useful in many contexts, interpreters-who already use ASR tools such as Dragon-add critical…

声音 · 计算机科学 2025-10-15 Carlos Arriaga , Alejandro Pozo , Javier Conde , Alvaro Alonso

Low-resolution text images are often seen in natural scenes such as documents captured by mobile phones. Recognizing low-resolution text images is challenging because they lose detailed content information, leading to poor recognition…

计算机视觉与模式识别 · 计算机科学 2020-08-04 Wenjia Wang , Enze Xie , Xuebo Liu , Wenhai Wang , Ding Liang , Chunhua Shen , Xiang Bai

The combination of Large Language Models (LLM) and Automatic Speech Recognition (ASR), when deployed on edge devices (called edge ASR-LLM), can serve as a powerful personalized assistant to enable audio-based interaction for users. Compared…

The notion of visual similarity is essential for computer vision, and in applications and studies revolving around vector embeddings of images. However, the scarcity of benchmark datasets poses a significant hurdle in exploring how these…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Tillmann Ohm , Andres Karjus , Mikhail Tamm , Maximilian Schich

We study the video super-resolution (SR) problem for facilitating video analytics tasks, e.g. action recognition, instead of for visual quality. The popular action recognition methods based on convolutional networks, exemplified by…

计算机视觉与模式识别 · 计算机科学 2020-03-13 Haochen Zhang , Dong Liu , Zhiwei Xiong

Arbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Hangwei Chen , Feng Shao , Xiongli Chai , Yuese Gu , Qiuping Jiang , Xiangchao Meng , Yo-Sung Ho

Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. While a wealth of spoken content is accessible in television…

音频与语音处理 · 电气工程与系统科学 2026-02-26 Cheng-Yeh Yang , Chien-Chun Wang , Li-Wei Chen , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

Artistic text recognition is an extremely challenging task with a wide range of applications. However, current scene text recognition methods mainly focus on irregular text while have not explored artistic text specifically. The challenges…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Xudong Xie , Ling Fu , Zhifei Zhang , Zhaowen Wang , Xiang Bai