中文
相关论文

相关论文: Diagnosing and Correcting Concept Omission in Mult…

200 篇论文

Recovering degraded low-resolution text images is challenging, especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. Ensuring both text fidelity and style realness is crucial for…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Yuzhe Zhang , Jiawei Zhang , Hao Li , Zhouxia Wang , Luwei Hou , Dongqing Zou , Liheng Bian

Existing approaches for controlling text-to-image diffusion models, while powerful, do not allow for explicit 3D object-centric control, such as precise control of object orientation. In this work, we address the problem of multi-object…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Rishubh Parihar , Vaibhav Agrawal , Sachidanand VS , R. Venkatesh Babu

Recent text-to-image models have achieved impressive results in generating high-quality images. However, when tasked with multi-concept generation creating images that contain multiple characters or objects, existing methods often suffer…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Yang Zhang , Rui Zhang , Xuecheng Nie , Haochen Li , Jikun Chen , Yifan Hao , Xin Zhang , Luoqi Liu , Ling Li

In this study, a modular, data-free pipeline for multi-label intention recognition is proposed for agentic AI applications in transportation. Unlike traditional intent recognition systems that depend on large, annotated corpora and often…

机器学习 · 计算机科学 2025-11-06 Xiaocai Zhang , Hur Lim , Ke Wang , Zhe Xiao , Jing Wang , Kelvin Lee , Xiuju Fu , Zheng Qin

Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correlation between visual and textual domains in open concepts…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Tuan-Anh Vu , Duc Thanh Nguyen , Qing Guo , Nhat Chung , Binh-Son Hua , Ivor W. Tsang , Sai-Kit Yeung

Open World Object Detection(OWOD) addresses realistic scenarios where unseen object classes emerge, enabling detectors trained on known classes to detect unknown objects and incrementally incorporate the knowledge they provide. While…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Sunoh Lee , Minsik Jeon , Jihong Min , Junwon Seo

When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading Optical Character Recognition (OCR) tokens, which requires…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Zhaokai Wang , Renda Bao , Qi Wu , Si Liu

Large scale multiple-input multiple-output (MIMO) or Massive MIMO is one of the pivotal technologies for future wireless networks. However, the performance of massive MIMO systems heavily relies on accurate channel estimation. While the…

信号处理 · 电气工程与系统科学 2020-02-25 Parna Sabeti , Arman Farhang , Irene Macaluso , Nicola Marchetti , Linda Doyle

Vision-language models (VLMs) are often deployed on text-only inputs, although they are trained with images. We find that removing the vision modality causes large drops in accuracy and severe miscalibration, and the model does not behave…

计算与语言 · 计算机科学 2026-05-14 Mingyeong Kim , Jungwon Choi , Chaeyun Jang , Juho Lee

Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concepts can be replaced through text conditioning (e.g., a dog to…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Chao Huang , Susan Liang , Yunlong Tang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control,…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Haoteng Li , Zhao Yang , Zezhong Qian , Gongpeng Zhao , Yuqi Huang , Jun Yu , Huazheng Zhou , Longjun Liu

Text-to-image diffusion models suffer from the risk of generating outdated, copyrighted, incorrect, and biased content. While previous methods have mitigated the issues on a small scale, it is essential to handle them simultaneously in…

计算机视觉与模式识别 · 计算机科学 2024-10-14 Tianwei Xiong , Yue Wu , Enze Xie , Yue Wu , Zhenguo Li , Xihui Liu

This paper considers a $K$-user single-input-single-output interference channel with inter-symbol interference (ISI), in which the channel coefficients are assumed to be linear time-invariant with finite-length impulse response. The primary…

信息论 · 计算机科学 2016-09-09 Namyoon Lee

With the development of artificial intelligence (AI) techniques, implementing AI-based techniques to improve wireless transceivers becomes an emerging research topic. Within this context, AI-based channel characterization and estimation…

信号处理 · 电气工程与系统科学 2025-10-29 Yuzhi Yang , Sen Yan , Weijie Zhou , Brahim Mefgouda , Ridong Li , Zhaoyang Zhang , Mérouane Debbah

This paper focuses on wireless multiple-input multiple-output (MIMO)-orthogonal frequency division multiplex (OFDM) receivers. Traditional wireless receivers have relied on mathematical modeling and Bayesian inference, achieving remarkable…

信号处理 · 电气工程与系统科学 2026-01-30 Yuzhi Yang , Omar Alhussein , Atefeh Arani , Zhaoyang Zhang , Mérouane Debbah

As large-scale diffusion models continue to advance, they excel at producing high-quality images but often generate unwanted content, such as sexually explicit or violent content. Existing methods for concept removal generally guide the…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Lingyun Zhang , Yu Xie , Yanwei Fu , Ping Chen

Watermarking is an important mechanism for provenance and copyright protection of diffusion-generated images. Training-free methods, exemplified by Gaussian Shading, embed watermarks into the initial noise of diffusion models with…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Yuwei Chen , Zhenliang He , Jia Tang , Meina Kan , Shiguang Shan

Text-guided image generation has advanced rapidly with large-scale diffusion models, yet achieving precise stylization with visual exemplars remains difficult. Existing approaches often depend on task-specific retraining or expensive…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Yingying Deng , Xiangyu He , Fan Tang , Weiming Dong , Xucheng Yin

The miniaturization of thermal sensors for mobile platforms inherently limits their spatial resolution and textural fidelity, leading to blurry and less informative images. Existing thermal super-resolution (SR) methods can be grouped into…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Minchong Chen , Xiaoyun Yuan , Junzhe Wan , Jianing Zhang , Jun Zhang

Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full article supports.…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Fanxiao Li , Jiaying Wu , Tingchao Fu , Dayang Li , Herun Wan , Wei Zhou , Min-Yen Kan