中文
相关论文

相关论文: An Interpretable X-ray Style Transfer via Trainabl…

200 篇论文

Capturing different intensity and directions of light rays at the same scene Light field (LF) can encode the 3D scene cues into a 4D LF image which has a wide range of applications (i.e. post-capture refocusing and depth sensing). LF image…

图像与视频处理 · 电气工程与系统科学 2024-09-27 Zhongxin Yu , Liang Chen , Zhiyun Zeng , Kunping Yang , Shaofei Luo , Shaorui Chen , Cheng Zhong

Interpretability in machine learning models is important in high-stakes decisions, such as whether to order a biopsy based on a mammographic exam. Mammography poses important challenges that are not present in other computer vision tasks:…

Deep implicit functions (DIFs), as a kind of 3D shape representation, are becoming more and more popular in the 3D vision community due to their compactness and strong representation power. However, unlike polygon mesh-based templates, it…

计算机视觉与模式识别 · 计算机科学 2021-05-14 Zerong Zheng , Tao Yu , Qionghai Dai , Yebin Liu

Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Yihe Deng , Pan Lu , Fan Yin , Ziniu Hu , Sheng Shen , Quanquan Gu , James Zou , Kai-Wei Chang , Wei Wang

The Learned Primal Dual (LPD) method has shown promising results in various tomographic reconstruction modalities, particularly under challenging acquisition restrictions such as limited viewing angles or a limited number of views. We…

图像与视频处理 · 电气工程与系统科学 2026-01-01 Sean Breckling , Matthew Swan , Keith D. Tan , Derek Wingard , Brandon Baldonado , Yoohwan Kim , Ju-Yeon Jo , Evan Scott , Jordan Pillow

A visual-language model (VLM) pre-trained on natural images and text pairs poses a significant barrier when applied to medical contexts due to domain shift. Yet, adapting or fine-tuning these VLMs for medical use presents considerable…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Aisha Urooj Khan , John Garrett , Tyler Bradshaw , Lonie Salkowski , Jiwoong Jason Jeong , Amara Tariq , Imon Banerjee

With a growing demand for the search by image, many works have studied the task of fashion instance-level image retrieval (FIR). Furthermore, the recent works introduce a concept of fashion attribute manipulation (FAM) which manipulates a…

计算机视觉与模式识别 · 计算机科学 2019-07-12 Minchul Shin , Sanghyuk Park , Taeksoo Kim

How to efficiently transform large language models (LLMs) into instruction followers is recently a popular research direction, while training LLM for multi-modal reasoning remains less explored. Although the recent LLaMA-Adapter…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Peng Gao , Jiaming Han , Renrui Zhang , Ziyi Lin , Shijie Geng , Aojun Zhou , Wei Zhang , Pan Lu , Conghui He , Xiangyu Yue , Hongsheng Li , Yu Qiao

This paper presents a systematic study on developing multi-template machine learning (ML) surrogate models and applying them to the inverse design of transformers (XFMRs) in radio-frequency integrated circuits (RFICs). Our study starts with…

机器学习 · 计算机科学 2025-12-01 Houbo He , Yizhou Xu , Lei Xia , Yaolong Hu , Fan Cai , Taiyun Chi

Generative AI models provide a wide range of tools capable of performing complex tasks in a fraction of the time it would take a human. Among these, Large Language Models (LLMs) stand out for their ability to generate diverse texts, from…

计算与语言 · 计算机科学 2024-10-07 Baldomero R. Árbol , Dan Casas

Humans learn language via multi-modal knowledge. However, due to the text-only pre-training scheme, most existing pre-trained language models (PLMs) are hindered from the multi-modal information. To inject visual knowledge into PLMs,…

计算与语言 · 计算机科学 2024-02-19 Xinyun Zhang , Haochen Tan , Han Wu , Bei Yu

We introduce Lavender, a simple supervised fine-tuning (SFT) method that boosts the performance of advanced vision-language models (VLMs) by leveraging state-of-the-art image generation models such as Stable Diffusion. Specifically,…

机器学习 · 计算机科学 2025-05-27 Chen Jin , Ryutaro Tanno , Amrutha Saseendran , Tom Diethe , Philip Teare

Image style transfer is an underdetermined problem, where a large number of solutions can satisfy the same constraint (the content and style). Although there have been some efforts to improve the diversity of style transfer by introducing…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Zhizhong Wang , Lei Zhao , Haibo Chen , Lihong Qiu , Qihang Mo , Sihuan Lin , Wei Xing , Dongming Lu

The recent work Local Implicit Image Function (LIIF) and subsequent Implicit Neural Representation (INR) based works have achieved remarkable success in Arbitrary-Scale Super-Resolution (ASSR) by using MLP to decode Low-Resolution (LR)…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Zongyao He , Zhi Jin

Photorealistic stylization aims to transfer the style of a reference photo onto a content photo in a natural fashion, such that the stylized image looks like a real photo taken by a camera. State-of-the-art methods stylize the image locally…

计算机视觉与模式识别 · 计算机科学 2021-07-07 Ying Qu , Zhenzhou Shao , Hairong Qi

In near-field extremely large-scale multiple-input multiple-output (XL-MIMO) systems, spherical wavefront propagation expands the traditional beam codebook into the joint angular-distance domain, rendering conventional beam training…

信号处理 · 电气工程与系统科学 2026-03-18 Mengyuan Li , Qianfan Lu , Jiachen Tian , Hongjun Hu , Yu Han , Xiao Li , Chao-kai Wen , Shi Jin

Photo enhancement plays a crucial role in augmenting the visual aesthetics of a photograph. In recent years, photo enhancement methods have either focused on enhancement performance, producing powerful models that cannot be deployed on edge…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Feng Zhang , Haoyou Deng , Zhiqiang Li , Lida Li , Bin Xu , Qingbo Lu , Zisheng Cao , Minchen Wei , Changxin Gao , Nong Sang , Xiang Bai

This paper proposes a novel method for Text Style Transfer (TST) based on parameter-efficient fine-tuning of Large Language Models (LLMs). Addressing the scarcity of parallel corpora that map between styles, the study employs roundtrip…

计算与语言 · 计算机科学 2026-02-17 Ruoxi Liu , Philipp Koehn

While filtered back projection (FBP) is still the method of choice for fast tomographic reconstruction, its performance degrades noticeably in the presence of noise, incomplete sampling, or non-standard scan geometries. We propose a…

数值分析 · 数学 2026-02-16 Hamid Fathi , Alexander Skorikov , Tristan van Leeuwen

Image-to-Image translation models can help mitigate various challenges inherent to medical image acquisition. Latent diffusion models (LDMs) leverage efficient learning in compressed latent space and constitute the core of state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Junhyeok Lee , Hyunwoong Kim , Hyungjin Chung , Heeseong Eom , Joon Jang , Chul-Ho Sohn , Kyu Sung Choi