English
Related papers

Related papers: Seedream 3.0 Technical Report

200 papers

Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision-language models face two…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Pengcheng Shi , Minghui Zhang , Kehan Song , Jiaqi Liu , Yun Gu , Xinglin Zhang

Observing that Semantic features learned in an image classification task and Appearance features learned in a similarity matching task complement each other, we build a twofold Siamese network, named SA-Siam, for real-time object tracking.…

Computer Vision and Pattern Recognition · Computer Science 2018-02-27 Anfeng He , Chong Luo , Xinmei Tian , Wenjun Zeng

This work aims to produce translations that convey source language content at a formality level that is appropriate for a particular audience. Framing this problem as a neural sequence-to-sequence task ideally requires training triplets…

Computation and Language · Computer Science 2019-12-02 Xing Niu , Marine Carpuat

Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Team Seedance , Heyi Chen , Siyan Chen , Xin Chen , Yanfei Chen , Ying Chen , Zhuo Chen , Feng Cheng , Tianheng Cheng , Xinqi Cheng , Xuyan Chi , Jian Cong , Jing Cui , Qinpeng Cui , Qide Dong , Junliang Fan , Jing Fang , Zetao Fang , Chengjian Feng , Han Feng , Mingyuan Gao , Yu Gao , Dong Guo , Qiushan Guo , Boyang Hao , Qingkai Hao , Bibo He , Qian He , Tuyen Hoang , Ruoqing Hu , Xi Hu , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Donglei Ji , Siqi Jiang , Wei Jiang , Yunpu Jiang , Zhuo Jiang , Ashley Kim , Jianan Kong , Zhichao Lai , Shanshan Lao , Yichong Leng , Ai Li , Feiya Li , Gen Li , Huixia Li , JiaShi Li , Liang Li , Ming Li , Shanshan Li , Tao Li , Xian Li , Xiaojie Li , Xiaoyang Li , Xingxing Li , Yameng Li , Yifu Li , Yiying Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Zhiqiang Liang , Wang Liao , Yalin Liao , Heng Lin , Kengyu Lin , Shanchuan Lin , Xi Lin , Zhijie Lin , Feng Ling , Fangfang Liu , Gaohong Liu , Jiawei Liu , Jie Liu , Jihao Liu , Shouda Liu , Shu Liu , Sichao Liu , Songwei Liu , Xin Liu , Xue Liu , Yibo Liu , Zikun Liu , Zuxi Liu , Junlin Lyu , Lecheng Lyu , Qian Lyu , Han Mu , Xiaonan Nie , Jingzhe Ning , Xitong Pan , Yanghua Peng , Lianke Qin , Xueqiong Qu , Yuxi Ren , Kai Shen , Guang Shi , Lei Shi , Yan Song , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Yan Sun , Zeyu Sun , Wenjing Tang , Yaxue Tang , Zirui Tao , Feng Wang , Furui Wang , Jinran Wang , Junkai Wang , Ke Wang , Kexin Wang , Qingyi Wang , Rui Wang , Sen Wang , Shuai Wang , Tingru Wang , Weichen Wang , Xin Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Ziyu Wang , Guoqiang Wei , Wanru Wei , Di Wu , Guohong Wu , Hanjie Wu , Jian Wu , Jie Wu , Ruolan Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Liang Xiang , Fei Xiao , XueFeng Xiao , Pan Xie , Shuangyi Xie , Shuang Xu , Jinlan Xue , Shen Yan , Bangbang Yang , Ceyuan Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yang Yang , Yihang Yang , ZhiXian Yang , Ziyan Yang , Songting Yao , Yifan Yao , Zilyu Ye , Bowen Yu , Jian Yu , Chujie Yuan , Linxiao Yuan , Sichun Zeng , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Chuntao Zhang , Heng Zhang , Jingjie Zhang , Kuo Zhang , Liang Zhang , Liying Zhang , Manlin Zhang , Ting Zhang , Weida Zhang , Xiaohe Zhang , Xinyan Zhang , Yan Zhang , Yuan Zhang , Zixiang Zhang , Fengxuan Zhao , Huating Zhao , Yang Zhao , Hao Zheng , Jianbin Zheng , Xiaozheng Zheng , Yangyang Zheng , Yijie Zheng , Jiexin Zhou , Jiahui Zhu , Kuan Zhu , Shenhan Zhu , Wenjia Zhu , Benhui Zou , Feilong Zuo

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Yushi Lan , Yihang Luo , Fangzhou Hong , Shangchen Zhou , Honghua Chen , Zhaoyang Lyu , Shuai Yang , Bo Dai , Chen Change Loy , Xingang Pan

The completion, extension, and generation of 3D semantic scenes are an interrelated set of capabilities that are useful for robotic navigation and exploration. Existing approaches seek to decouple these problems and solve them one-off.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xujia Zhang , Brendan Crowe , Christoffer Heckman

This paper contributes to the "BraTS 2024 Brain MR Image Synthesis Challenge" and presents a conditional Wavelet Diffusion Model (cWDM) for directly solving a paired image-to-image translation task on high-resolution volumes. While deep…

Image and Video Processing · Electrical Eng. & Systems 2024-11-27 Paul Friedrich , Alicia Durrer , Julia Wolleb , Philippe C. Cattin

Large language models (LLMs) exhibit exceptional performance across a wide range of tasks; however, their token-by-token autoregressive generation process significantly hinders inference speed. Speculative decoding presents a promising…

Computation and Language · Computer Science 2025-03-04 Kai Lv , Honglin Guo , Qipeng Guo , Xipeng Qiu

Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusion priors, optimized…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Ling Yang , Zixiang Zhang , Junlin Han , Bohan Zeng , Runjia Li , Philip Torr , Wentao Zhang

In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of VideoLLaMA3 is vision-centric. The meaning of "vision-centric" is two-fold: the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Boqiang Zhang , Kehan Li , Zesen Cheng , Zhiqiang Hu , Yuqian Yuan , Guanzheng Chen , Sicong Leng , Yuming Jiang , Hang Zhang , Xin Li , Peng Jin , Wenqi Zhang , Fan Wang , Lidong Bing , Deli Zhao

A very recent trend in generative modeling is building 3D-aware generators from 2D image collections. To induce the 3D bias, such models typically rely on volumetric rendering, which is expensive to employ at high resolutions. During the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Ivan Skorokhodov , Sergey Tulyakov , Yiqun Wang , Peter Wonka

In the realm of text-to-3D generation, utilizing 2D diffusion models through score distillation sampling (SDS) frequently leads to issues such as blurred appearances and multi-faced geometry, primarily due to the intrinsically noisy nature…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Pengsheng Guo , Hans Hao , Adam Caccavale , Zhongzheng Ren , Edward Zhang , Qi Shan , Aditya Sankar , Alexander G. Schwing , Alex Colburn , Fangchang Ma

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Chunyu Xie , Bin Wang , Fanjing Kong , Jincheng Li , Dawei Liang , Ji Ao , Dawei Leng , Yuhui Yin

Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promising approach to enhance the reasoning quality of Large Language Models (LLMs). However,…

Computation and Language · Computer Science 2026-01-07 Lingyin Zhang , Jun Gao , Xiaoxue Ren , Ziqiang Cao

Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious…

Computational Engineering, Finance, and Science · Computer Science 2025-12-03 Yuan Wang , Shujian Gao , Jiaxiang Liu , Songtao Jiang , Haoxiang Xia , Xiaotian Zhang , Zhaolu Kang , Yemin Wang , Zuozhu Liu

In recent times, automatic text-to-3D content creation has made significant progress, driven by the development of pretrained 2D diffusion models. Existing text-to-3D methods typically optimize the 3D representation to ensure that the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Yiwei Ma , Yijun Fan , Jiayi Ji , Haowei Wang , Xiaoshuai Sun , Guannan Jiang , Annan Shu , Rongrong Ji

Recent one image to 3D generation methods commonly adopt Score Distillation Sampling (SDS). Despite the impressive results, there are multiple deficiencies including multi-view inconsistency, over-saturated and over-smoothed textures, as…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Junwu Zhang , Zhenyu Tang , Yatian Pang , Xinhua Cheng , Peng Jin , Yida Wei , Munan Ning , Li Yuan

In this paper, we propose VistaDream a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Haiping Wang , Yuan Liu , Ziwei Liu , Wenping Wang , Zhen Dong , Bisheng Yang

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Andreas Blattmann , Robin Rombach , Huan Ling , Tim Dockhorn , Seung Wook Kim , Sanja Fidler , Karsten Kreis

We present Stable Video Diffusion - a latent video diffusion model for high-resolution, state-of-the-art text-to-video and image-to-video generation. Recently, latent diffusion models trained for 2D image synthesis have been turned into…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Andreas Blattmann , Tim Dockhorn , Sumith Kulal , Daniel Mendelevitch , Maciej Kilian , Dominik Lorenz , Yam Levi , Zion English , Vikram Voleti , Adam Letts , Varun Jampani , Robin Rombach
‹ Prev 1 8 9 10 Next ›