English
Related papers

Related papers: ASTRA: Enhancing Multi-Subject Generation with Ret…

200 papers

There are many excellent solutions in image restoration.However, most methods require on training separate models to restore images with different types of degradation.Although existing all-in-one models effectively address multiple types…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Jiawei Mao , Juncheng Wu , Yuyin Zhou , Xuesong Yin , Yuanqi Chang

Retrieval-augmented generation (RAG) typically treats retrieval and generation as separate systems. We ask whether an attention-based encoder-decoder can instead retrieve directly from its own internal representations. We introduce INTRA…

Machine Learning · Computer Science 2026-05-11 Elad Hoffer , Yochai Blau , Edan Kinderman , Ron Banner , Daniel Soudry , Boris Ginsburg

State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While…

Artificial Intelligence · Computer Science 2025-10-17 Supriti Sinhamahapatra , Jan Niehues

The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical tasks that require unified slide-level reasoning and interpretable linkage to clinically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Tianyang Wang , Ziyu Su , Abdul Rehman Akbar , Usama Sajjad , Lina Gokhale , Charles Rabolli , Wei Chen , Anil Parwani , Muhammad Khalid Khan Niazi

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Foivos Paraperas Papantoniou , Stathis Galanakis , Rolandos Alexandros Potamias , Bernhard Kainz , Stefanos Zafeiriou

Stance detection is an important task, supporting many downstream tasks such as discourse parsing and modeling the propagation of fake news, rumors, and science denial. In this paper, we propose a novel framework for stance detection. Our…

Computation and Language · Computer Science 2021-12-21 Ron Korenblum Pick , Vladyslav Kozhukhov , Dan Vilenchik , Oren Tsur

Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encoders has recently been shown to provide a powerful inductive bias, improving both convergence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Loukas Sfountouris , Giannis Daras , Paris Giampouras

This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Xulong Zhang , Xiaoyang Qu , Haoxiang Shi , Chunguang Xiao , Jianzong Wang

Image classification models tend to make decisions based on peripheral attributes of data items that have strong correlation with a target variable (i.e., dataset bias). These biased models suffer from the poor generalization capability…

Machine Learning · Computer Science 2021-10-26 Jungsoo Lee , Eungyeup Kim , Juyoung Lee , Jihyeon Lee , Jaegul Choo

Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric face templates and hand-designed blendshape spaces. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Matan Levy , Gavriel Habib , Issar Tzachor , Dvir Samuel , Rami Ben-Ari , Nir Darshan , Or Litany , Dani Lischinski

The remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chenghao Zhang , Gaofeng Meng , Bin Fan , Kun Tian , Zhaoxiang Zhang , Shiming Xiang , Chunhong Pan

Reference Audio-Visual Segmentation (Ref-AVS) tasks challenge models to precisely locate sounding objects by integrating visual, auditory, and textual cues. Existing methods often lack genuine semantic understanding, tending to memorize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Ziyang Luo , Nian Liu , Fahad Shahbaz Khan , Junwei Han

Developing a reliable and practical face recognition system is a long-standing goal in computer vision research. Existing literature suggests that pixel-wise face alignment is the key to achieve high-accuracy face recognition. By assuming a…

Computer Vision and Pattern Recognition · Computer Science 2015-01-21 Yuting Zhang , Kui Jia , Yueming Wang , Gang Pan , Tsung-Han Chan , Yi Ma

Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis. While prior methods emphasize identity preservation from a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Guocheng Gordon Qian , Daniil Ostashev , Egor Nemchinov , Avihay Assouline , Sergey Tulyakov , Kuan-Chieh Jackson Wang , Kfir Aberman

Person re-identification (Re-ID) often faces challenges due to variations in human poses and camera viewpoints, which significantly affect the appearance of individuals across images. Existing datasets frequently lack diversity and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Inès Hyeonsu Kim , Woojeong Jin , Soowon Son , Junyoung Seo , Seokju Cho , JeongYeol Baek , Byeongwon Lee , JoungBin Lee , Seungryong Kim

Text-to-image diffusion models are increasingly developed through open-source reuse and repeated downstream fine-tuning, where reused checkpoints are difficult to verify and thus more susceptible to hidden backdoor behaviors. In such…

Cryptography and Security · Computer Science 2026-05-20 Kai Wang , Jiale Zhang , Chengcheng Zhu , Chuang Ma , Songze Li

In this work, we explore how a strategic selection of camera movements can facilitate the task of 6D multi-object pose estimation in cluttered scenarios while respecting real-world constraints important in robotics and augmented reality…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Juil Sock , Guillermo Garcia-Hernando , Tae-Kyun Kim

Introducing Entity-Aspect Sentiment Triplet Extraction (EASTE), a novel Aspect-Based Sentiment Analysis (ABSA) task which extends Target-Aspect-Sentiment Detection (TASD) by separating aspect categories (e.g., food#quality) into pre-defined…

Computation and Language · Computer Science 2024-07-08 Vorakit Vorakitphan , Milos Basic , Guilhaume Leroy Meline

Self-supervised representation learning has gained increasing attention for strong generalization ability without relying on paired datasets. However, it has not been explored sufficiently for facial representation. Self-supervised facial…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Ruian He , Zhen Xing , Weimin Tan , Bo Yan

In recent years, learning-based methods have achieved significant advancements in multi-exposure image fusion. However, two major stumbling blocks hinder the development, including pixel misalignment and inefficient inference. Reliance on…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Zhu Liu , Jinyuan Liu , Guanyao Wu , Zihang Chen , Xin Fan , Risheng Liu