中文
相关论文

相关论文: CineVision: An Interactive Pre-visualization Story…

200 篇论文

Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in professional production workflows such as storyboard…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Hongji Yang , Songlian Li , Yucheng Zhou , Xiaotong Zhao , Alan Zhao , Chengzhong Xu , Jianbing Shen

Indoor scene synthesis has become increasingly important with the rise of Embodied AI, which requires 3D environments that are not only visually realistic but also physically plausible and functionally diverse. While recent approaches have…

图形学 · 计算机科学 2025-10-28 Yandan Yang , Baoxiong Jia , Shujie Zhang , Siyuan Huang

Interior design is crucial in creating aesthetically pleasing and functional indoor spaces. However, developing and editing interior design concepts requires significant time and expertise. We propose Virtual Interior DESign (VIDES) system…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Minh-Hien Le , Chi-Bien Chu , Khanh-Duy Le , Tam V. Nguyen , Minh-Triet Tran , Trung-Nghia Le

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Sridhar S , Nithin A , Shakeel Rifath , Vasantha Raj

Audio description (AD) makes video content accessible to millions of blind and low vision (BLV) users. However, creating high-quality AD involves a trade-off between the precision of human-crafted descriptions and the efficiency of…

人机交互 · 计算机科学 2025-08-05 Maryam Cheema , Sina Elahimanesh , Samuel Martin , Pooyan Fazli , Hasti Seifi

Generative AI is reshaping product design practices through "vibe coding," where product team members express intent in natural language and AI translates it into functional prototypes and code. Despite rapid adoption, little research has…

人机交互 · 计算机科学 2026-05-04 Jie Li , Youyang Hou , Laura Lin , Ruihao Zhu , Hancheng Cao , Abdallah El Ali

Creators struggle to edit long-form, narrative-rich videos not because of UI complexity, but due to the cognitive demands of searching, storyboarding, and sequencing hours of footage. Existing transcript- or embedding-based methods fall…

人工智能 · 计算机科学 2025-09-30 Zihan Ding , Xinyi Wang , Junlong Chen , Per Ola Kristensson , Junxiao Shen

Generative AI has greatly transformed creative work in various domains, such as screenwriting. To understand this transformation, prior research often focused on capturing a snapshot of human-AI co-creation practice at a specific moment,…

人机交互 · 计算机科学 2026-02-09 Yuying Tang , Jiayi Zhou , Haotian Li , Xing Xie , Xiaojuan Ma , Huamin Qu

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-to-audio (VT2A), the current formulation faces three key…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Bingxuan Li , Yiming Cui , Yicheng He , Yiwei Wang , Shu Zhang , Longyin Wen , Yulei Niu

OpenKinoAI is an open source framework for post-production of ultra high definition video which makes it possible to emulate professional multiclip editing techniques for the case of single camera recordings. OpenKinoAI includes tools for…

多媒体 · 计算机科学 2020-11-11 Rémi Ronfard , Rémi Colin de Verdière

We present Vinci, a vision-language system designed to provide real-time, comprehensive AI assistance on portable devices. At its core, Vinci leverages EgoVideo-VL, a novel model that integrates an egocentric vision foundation model with a…

Biomedical researchers face increasing challenges in navigating millions of publications in diverse domains. Traditional search engines typically return articles as ranked text lists, offering little support for global exploration or…

信息检索 · 计算机科学 2026-01-29 Huan He , Xueqing Peng , Yutong Xie , Qijia Liu , Chia-Hsuan Chang , Lingfei Qian , Brian Ondov , Qiaozhu Mei , Hua Xu

Innovative HealthTech teams develop Artificial Intelligence (AI) systems in contexts where ethical expectations and organizational priorities must be balanced under severe resource constraints. While Responsible AI practices are expected to…

人机交互 · 计算机科学 2026-03-02 Svitlana Surodina , Sinem Görücü , Lili Golmohammadi , Emelia Delaney , Rita Borgo

The target of automatic video summarization is to create a short skim of the original long video while preserving the major content/events. There is a growing interest in the integration of user queries into video summarization or…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Guande Wu , Jianzhe Lin , Claudio T. Silva

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplified workflow where…

计算与语言 · 计算机科学 2025-11-19 Rui Liu , Yuan Zhao , Zhenqi Jia

Conversational recommender systems support users in accomplishing recommendation-related goals via multi-turn conversations. To better model dynamically changing user preferences and provide the community with a reusable development…

信息检索 · 计算机科学 2020-09-09 Javeria Habib , Shuo Zhang , Krisztian Balog

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

What if the patterns hidden within dialogue reveal more about communication than the words themselves? We introduce Conversational DNA, a novel visual language that treats any dialogue -- whether between humans, between human and AI, or…

人机交互 · 计算机科学 2025-08-12 Baihan Lin

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Instruction-based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although reference-guided editing…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Yiqi Lin , Guoqiang Liang , Ziyun Zeng , Zechen Bai , Yanzhe Chen , Mike Zheng Shou
‹ 上一页 1 8 9 10 下一页 ›