中文
相关论文

相关论文: Sora: A Review on Background, Technology, Limitati…

200 篇论文

With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Rui Sun , Yumin Zhang , Tejal Shah , Jiahao Sun , Shuoying Zhang , Wenqi Li , Haoran Duan , Bo Wei , Rajiv Ranjan

Text-to-video generative AI models such as Sora OpenAI have the potential to disrupt multiple industries. In this paper, we report a qualitative social media analysis aiming to uncover people's perceived impact of and concerns about Sora's…

计算机与社会 · 计算机科学 2024-06-19 Kyrie Zhixuan Zhou , Abhinav Choudhry , Ece Gumusel , Madelyn Rose Sanfilippo

The advent of text-to-video generation models has revolutionized content creation as it produces high-quality videos from textual prompts. However, concerns regarding inherent biases in such models have prompted scrutiny, particularly…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Mohammad Nadeem , Shahab Saquib Sohail , Erik Cambria , Björn W. Schuller , Amir Hussain

The rapid advancement of Generative AI (Gen-AI) is transforming Human-Computer Interaction (HCI), with significant implications across various sectors. This study investigates the public's perception of Sora OpenAI, a pioneering Gen-AI…

计算机与社会 · 计算机科学 2024-03-27 Reza Hadi Mogavi , Derrick Wang , Joseph Tu , Hilda Hadan , Sabrina A. Sgandurra , Pan Hui , Lennart E. Nacke

The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential…

Text-to-video generation has made significant strides, but replicating the capabilities of advanced systems like OpenAI Sora remains challenging due to their closed-source nature. Existing open-source methods struggle to achieve comparable…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Zhengqing Yuan , Yixin Liu , Yihan Cao , Weixiang Sun , Haolong Jia , Ruoxi Chen , Zhaoxu Li , Bin Lin , Li Yuan , Lifang He , Chi Wang , Yanfang Ye , Lichao Sun

Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Zangwei Zheng , Xiangyu Peng , Tianji Yang , Chenhui Shen , Shenggui Li , Hongxin Liu , Yukun Zhou , Tianyi Li , Yang You

General world models represent a crucial pathway toward achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications ranging from virtual environments to decision-making systems. Recently, the…

High-quality video generation, encompassing text-to-video (T2V), image-to-video (I2V), and video-to-video (V2V) generation, holds considerable significance in content creation to benefit anyone express their inherent creativity in new ways…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Ailing Zeng , Yuhang Yang , Weidong Chen , Wei Liu

OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. However, the ability of video generation models to discover such laws purely from visual data without human…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Bingyi Kang , Yang Yue , Rui Lu , Zhijie Lin , Yang Zhao , Kaixin Wang , Gao Huang , Jiashi Feng

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Faraz Waseem , Muhammad Shahzad

The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena. Despite its growing popularity, there is a lack of…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Xuanyi Li , Daquan Zhou , Chenxu Zhang , Shaodong Wei , Qibin Hou , Ming-Ming Cheng

We present On-device Sora, the first model training-free solution for diffusion-based on-device text-to-video generation that operates efficiently on smartphone-grade devices. To address the challenges of diffusion-based text-to-video…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Bosung Kim , Kyuhwan Lee , Isu Jeong , Jungmin Cheon , Yeojin Lee , Seulki Lee

We present On-device Sora, the first model training-free solution for diffusion-based on-device text-to-video generation that operates efficiently on smartphone-grade devices. To address the challenges of diffusion-based text-to-video…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Bosung Kim , Kyuhwan Lee , Isu Jeong , Jungmin Cheon , Yeojin Lee , Seulki Lee

Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content. However, it remains a formidable challenge pertaining to maintaining temporal consistency and ensuring action…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Deshun Yang , Luhui Hu , Yu Tian , Zihao Li , Chris Kelly , Bang Yang , Cindy Yang , Yuexian Zou

As AI-generated video platforms rapidly advance, ethical challenges such as copyright infringement emerge. This study examines how users make sense of AI-generated videos on OpenAI's Sora by conducting a qualitative content analysis of user…

人机交互 · 计算机科学 2025-12-08 Bohui Shen , Shrikar Bhatta , Alex Ireebanije , Zexuan Liu , Abhinav Choudhry , Ece Gumusel , Kyrie Zhixuan Zhou

The December 2024 release of OpenAI's Sora, a powerful video generation model driven by natural language prompts, highlights a growing convergence between large language models (LLMs) and video synthesis. As these multimodal systems evolve…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Misora Sugiyama , Hirokatsu Kataoka

Video generation models have achieved remarkable progress in the past year. The quality of AI video continues to improve, but at the cost of larger model size, increased data quantity, and greater demand for training compute. In this…

We introduce Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs. Our project comprises multiple components…

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, along with other text-to-video diffusion models, is highly reliant on…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Wenhao Wang , Yi Yang
‹ 上一页 1 2 3 10 下一页 ›