中文
相关论文

相关论文: Multi-Agent Synergy-Driven Iterative Visual Narrat…

200 篇论文

Automatically generating presentations from documents is a challenging task that requires accommodating content quality, visual appeal, and structural coherence. Existing methods primarily focus on improving and evaluating the content…

人工智能 · 计算机科学 2025-02-24 Hao Zheng , Xinyan Guan , Hao Kong , Jia Zheng , Weixiang Zhou , Hongyu Lin , Yaojie Lu , Ben He , Xianpei Han , Le Sun

Current one-pass 3D scene synthesis methods often suffer from spatial hallucinations, such as collisions, due to a lack of deliberative reasoning. To bridge this gap, we introduce SceneReVis, a vision-grounded self-reflection framework that…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Yang Zhao , Shizhao Sun , Meisheng Zhang , Yingdong Shi , Xubo Yang , Jiang Bian

Event-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily.…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Qiang Qu , Xiaoming Chen , Yuk Ying Chung , Yiran Shen

In today's fast-paced world, effective presentations have become an essential tool for communication in both online and offline meetings. The crafting of a compelling presentation requires significant time and effort, from gathering key…

计算与语言 · 计算机科学 2025-01-17 Tushar Aggarwal , Aarohi Bhand

Today's image generation systems are capable of producing realistic and high-quality images. However, user prompts often contain ambiguities, making it difficult for these systems to interpret users' potential intentions. Consequently,…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Yuheng Feng , Yangfan He , Yinghui Xia , Tianyu Shi , Jun Wang , Jinsong Yang

This paper presents a tri-agent cross-validation framework for analyzing stability and explainability in multi-model large language systems. The architecture integrates three heterogeneous LLMs-used for semantic generation, analytical…

计算与语言 · 计算机科学 2026-01-15 Toshiyuki Shigemura

Meetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content. To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and…

计算与语言 · 计算机科学 2023-10-24 Haochen Tan , Han Wu , Wei Shao , Xinyun Zhang , Mingjie Zhan , Zhaohui Hou , Ding Liang , Linqi Song

Iterative data generation and model re-training can effectively align large language models(LLMs) to human preferences. The process of data sampling is crucial, as it significantly influences the success of policy improvement. Repeated…

计算与语言 · 计算机科学 2024-10-07 Hai Ye , Hwee Tou Ng

Rendering photo-realistic novel-view images of complex scenes has been a long-standing challenge in computer graphics. In recent years, great research progress has been made on enhancing rendering quality and accelerating rendering speed in…

图形学 · 计算机科学 2024-02-21 Tiansong Zhou , Yebin Liu , Xuangeng Chu , Chengkun Cao , Changyin Zhou , Fei Yu , Yu Li

The promotion of academic papers has become an important means of enhancing research visibility. However, existing automated methods struggle limited storytelling, insufficient aesthetic quality, and constrained self-adjustment, making it…

计算与语言 · 计算机科学 2025-10-23 Chengzhi Liu , Yuzhe Yang , Kaiwen Zhou , Zhen Zhang , Yue Fan , Yanan Xie , Peng Qi , Xin Eric Wang

Presentation generation requires deep content research, coherent visual design, and iterative refinement based on observation. However, existing presentation agents often rely on predefined workflows and fixed templates. To address this, we…

人工智能 · 计算机科学 2026-04-21 Hao Zheng , Guozhao Mo , Xinru Yan , Qianhao Yuan , Wenkai Zhang , Xuanang Chen , Yaojie Lu , Hongyu Lin , Xianpei Han , Le Sun

In a conventional supervised learning setting, a machine learning model has access to examples of all object classes that are desired to be recognized during the inference stage. This results in a fixed model that lacks the flexibility to…

计算机视觉与模式识别 · 计算机科学 2020-01-27 Jathushan Rajasegaran , Munawar Hayat , Salman Khan , Fahad Shahbaz Khan , Ling Shao , Ming-Hsuan Yang

Paired image-text data with subtle variations in-between (e.g., people holding surfboards vs. people holding shovels) hold the promise of producing Vision-Language Models with proper compositional understanding. Synthesizing such training…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Haoxin Li , Boyang Li

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Hongji Yang , Yucheng Zhou , Wencheng Han , Runzhou Tao , Zhongying Qiu , Jianfei Yang , Jianbing Shen

Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or factoid-based. However, many real-world applications demand…

计算与语言 · 计算机科学 2025-08-29 Rohan Phanse , Yijie Zhou , Kejian Shi , Wencai Zhang , Yixin Liu , Yilun Zhao , Arman Cohan

Evidence synthesis has advanced through improved reporting standards, bias assessment tools, and analytic methods, but current workflows remain limited by a single-layer structure in which conceptual, methodological, and procedural…

统计方法学 · 统计学 2025-12-11 Hung Kuan Lee

Sequential reasoning in agent systems has been significantly advanced by large language models (LLMs), yet existing approaches face limitations. Reflection-driven reasoning relies solely on knowledge in pretrained models, limiting…

机器学习 · 计算机科学 2024-10-23 Chen Yang , Chenyang Zhao , Quanquan Gu , Dongruo Zhou

Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive…

音频与语音处理 · 电气工程与系统科学 2024-01-26 Heming Wang , Eric W. Healy , DeLiang Wang

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zhengwentai Sun , Keru Zheng , Chenghong Li , Hongjie Liao , Xihe Yang , Heyuan Li , Yihao Zhi , Shuliang Ning , Shuguang Cui , Xiaoguang Han

Although chain-of-thought reasoning and reinforcement learning (RL) have driven breakthroughs in NLP, their integration into generative vision models remains underexplored. We introduce ReasonGen-R1, a two-stage framework that first imbues…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Yu Zhang , Yunqi Li , Yifan Yang , Rui Wang , Yuqing Yang , Dai Qi , Jianmin Bao , Dongdong Chen , Chong Luo , Lili Qiu
‹ 上一页 1 2 3 10 下一页 ›