English
Related papers

Related papers: Beyond Play and Pause: Turning GPT-4o Spatial Weak…

200 papers

Text-to-video models have demonstrated impressive capabilities in producing diverse and captivating video content, showcasing a notable advancement in generative AI. However, these models generally lack fine-grained control over motion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tuna Han Salih Meral , Hidir Yesiltepe , Connor Dunlop , Pinar Yanardag

We present a new data-driven video inpainting method for recovering missing regions of video frames. A novel deep learning architecture is proposed which contains two sub-networks: a temporal structure inference network and a spatial detail…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Chuan Wang , Haibin Huang , Xiaoguang Han , Jue Wang

Prompt tuning, a parameter- and data-efficient transfer learning paradigm that tunes only a small number of parameters in a model's input space, has become a trend in the vision community since the emergence of large vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Yuhang Zang , Wei Li , Kaiyang Zhou , Chen Huang , Chen Change Loy

The integration of Large Language Models (LLMs), especially ChatGPT, into education is poised to revolutionize students' learning experiences by introducing innovative conversational learning methodologies. To empower students to fully…

Human-Computer Interaction · Computer Science 2024-09-18 Zixin Chen , Jiachen Wang , Meng Xia , Kento Shigyo , Dingdong Liu , Rong Zhang , Huamin Qu

Recent advances in text-to-video generation have harnessed the power of diffusion models to create visually compelling content conditioned on text prompts. However, they usually encounter high computational costs and often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Jiaxi Lv , Yi Huang , Mingfu Yan , Jiancheng Huang , Jianzhuang Liu , Yifan Liu , Yafei Wen , Xiaoxin Chen , Shifeng Chen

Finding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Ye Liu , Siyuan Li , Yang Wu , Chang Wen Chen , Ying Shan , Xiaohu Qie

Existing unsupervised video-to-video translation methods fail to produce translated videos which are frame-wise realistic, semantic information preserving and video-level consistent. In this work, we propose UVIT, a novel unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2020-04-15 Kangning Liu , Shuhang Gu , Andres Romero , Radu Timofte

People increasingly use videos on the Web as a source for learning. To support this way of learning, researchers and developers are continuously developing tools, proposing guidelines, analyzing data, and conducting experiments. However, it…

Multimedia · Computer Science 2023-08-15 Evelyn Navarrete , Andreas Nehring , Sascha Schanze , Ralph Ewerth , Anett Hoppe

This paper explores integrating microlearning strategies into university curricula, particularly in computer science education, to counteract the decline in class attendance and engagement in US universities after COVID. As students…

Learning sensorimotor control policies from high-dimensional images crucially relies on the quality of the underlying visual representations. Prior works show that structured latent space such as visual keypoints often outperforms…

Machine Learning · Computer Science 2021-06-15 Boyuan Chen , Pieter Abbeel , Deepak Pathak

Visuomotor policies often suffer from perceptual challenges, where visual differences between training and evaluation environments degrade policy performance. Policies relying on state estimations, like 6D pose, require task-specific…

Robotics · Computer Science 2025-10-07 Yunchu Zhang , Shubham Mittal , Zhengyu Zhang , Liyiming Ke , Siddhartha Srinivasa , Abhishek Gupta

Cross-domain Click-Through Rate prediction aims to tackle the data sparsity and the cold start problems in online advertising systems by transferring knowledge from source domains to a target domain. Most existing methods rely on…

Artificial Intelligence · Computer Science 2025-07-08 Wei Xu , Haoran Li , Baoyuan Ou , Lai Xu , Yingjie Qin , Ruilong Su , Ruiwen Xu

As data from IoT (Internet of Things) sensors become ubiquitous, state-of-the-art machine learning algorithms face many challenges on directly using sensor data. To overcome these challenges, methods must be designed to learn directly from…

Computer Vision and Pattern Recognition · Computer Science 2020-09-18 Qiong Liu , Yanxia Zhang

Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition. In this paper, we seek a deeper understanding of self-attention for temporal modeling in videos. We…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Bo He , Xitong Yang , Zuxuan Wu , Hao Chen , Ser-Nam Lim , Abhinav Shrivastava

Restoring images distorted by atmospheric turbulence is a ubiquitous problem in long-range imaging applications. While existing deep-learning-based methods have demonstrated promising results in specific testing conditions, they suffer from…

Image and Video Processing · Electrical Eng. & Systems 2023-12-12 Xingguang Zhang , Zhiyuan Mao , Nicholas Chimitt , Stanley H. Chan

Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term video dependency with self-attention. Unfortunately, they…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Limin Wang , Yu Qiao

Semi-supervised and unsupervised systems provide operators with invaluable support and can tremendously reduce the operators load. In the light of the necessity to process large volumes of video data and provide autonomous decisions, this…

Machine Learning · Statistics 2017-09-20 Olga Isupova , Danil Kuzin , Lyudmila Mihaylova

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning…

Computation and Language · Computer Science 2024-10-29 OpenAI , : , Aaron Hurst , Adam Lerer , Adam P. Goucher , Adam Perelman , Aditya Ramesh , Aidan Clark , AJ Ostrow , Akila Welihinda , Alan Hayes , Alec Radford , Aleksander Mądry , Alex Baker-Whitcomb , Alex Beutel , Alex Borzunov , Alex Carney , Alex Chow , Alex Kirillov , Alex Nichol , Alex Paino , Alex Renzin , Alex Tachard Passos , Alexander Kirillov , Alexi Christakis , Alexis Conneau , Ali Kamali , Allan Jabri , Allison Moyer , Allison Tam , Amadou Crookes , Amin Tootoochian , Amin Tootoonchian , Ananya Kumar , Andrea Vallone , Andrej Karpathy , Andrew Braunstein , Andrew Cann , Andrew Codispoti , Andrew Galu , Andrew Kondrich , Andrew Tulloch , Andrey Mishchenko , Angela Baek , Angela Jiang , Antoine Pelisse , Antonia Woodford , Anuj Gosalia , Arka Dhar , Ashley Pantuliano , Avi Nayak , Avital Oliver , Barret Zoph , Behrooz Ghorbani , Ben Leimberger , Ben Rossen , Ben Sokolowsky , Ben Wang , Benjamin Zweig , Beth Hoover , Blake Samic , Bob McGrew , Bobby Spero , Bogo Giertler , Bowen Cheng , Brad Lightcap , Brandon Walkin , Brendan Quinn , Brian Guarraci , Brian Hsu , Bright Kellogg , Brydon Eastman , Camillo Lugaresi , Carroll Wainwright , Cary Bassin , Cary Hudson , Casey Chu , Chad Nelson , Chak Li , Chan Jun Shern , Channing Conger , Charlotte Barette , Chelsea Voss , Chen Ding , Cheng Lu , Chong Zhang , Chris Beaumont , Chris Hallacy , Chris Koch , Christian Gibson , Christina Kim , Christine Choi , Christine McLeavey , Christopher Hesse , Claudia Fischer , Clemens Winter , Coley Czarnecki , Colin Jarvis , Colin Wei , Constantin Koumouzelis , Dane Sherburn , Daniel Kappler , Daniel Levin , Daniel Levy , David Carr , David Farhi , David Mely , David Robinson , David Sasaki , Denny Jin , Dev Valladares , Dimitris Tsipras , Doug Li , Duc Phong Nguyen , Duncan Findlay , Edede Oiwoh , Edmund Wong , Ehsan Asdar , Elizabeth Proehl , Elizabeth Yang , Eric Antonow , Eric Kramer , Eric Peterson , Eric Sigler , Eric Wallace , Eugene Brevdo , Evan Mays , Farzad Khorasani , Felipe Petroski Such , Filippo Raso , Francis Zhang , Fred von Lohmann , Freddie Sulit , Gabriel Goh , Gene Oden , Geoff Salmon , Giulio Starace , Greg Brockman , Hadi Salman , Haiming Bao , Haitang Hu , Hannah Wong , Haoyu Wang , Heather Schmidt , Heather Whitney , Heewoo Jun , Hendrik Kirchner , Henrique Ponde de Oliveira Pinto , Hongyu Ren , Huiwen Chang , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian O'Connell , Ian Osband , Ian Silber , Ian Sohl , Ibrahim Okuyucu , Ikai Lan , Ilya Kostrikov , Ilya Sutskever , Ingmar Kanitscheider , Ishaan Gulrajani , Jacob Coxon , Jacob Menick , Jakub Pachocki , James Aung , James Betker , James Crooks , James Lennon , Jamie Kiros , Jan Leike , Jane Park , Jason Kwon , Jason Phang , Jason Teplitz , Jason Wei , Jason Wolfe , Jay Chen , Jeff Harris , Jenia Varavva , Jessica Gan Lee , Jessica Shieh , Ji Lin , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joanne Jang , Joaquin Quinonero Candela , Joe Beutler , Joe Landers , Joel Parish , Johannes Heidecke , John Schulman , Jonathan Lachman , Jonathan McKay , Jonathan Uesato , Jonathan Ward , Jong Wook Kim , Joost Huizinga , Jordan Sitkin , Jos Kraaijeveld , Josh Gross , Josh Kaplan , Josh Snyder , Joshua Achiam , Joy Jiao , Joyce Lee , Juntang Zhuang , Justyn Harriman , Kai Fricke , Kai Hayashi , Karan Singhal , Katy Shi , Kavin Karthik , Kayla Wood , Kendra Rimbach , Kenny Hsu , Kenny Nguyen , Keren Gu-Lemberg , Kevin Button , Kevin Liu , Kiel Howe , Krithika Muthukumar , Kyle Luther , Lama Ahmad , Larry Kai , Lauren Itow , Lauren Workman , Leher Pathak , Leo Chen , Li Jing , Lia Guy , Liam Fedus , Liang Zhou , Lien Mamitsuka , Lilian Weng , Lindsay McCallum , Lindsey Held , Long Ouyang , Louis Feuvrier , Lu Zhang , Lukas Kondraciuk , Lukasz Kaiser , Luke Hewitt , Luke Metz , Lyric Doshi , Mada Aflak , Maddie Simens , Madelaine Boyd , Madeleine Thompson , Marat Dukhan , Mark Chen , Mark Gray , Mark Hudnall , Marvin Zhang , Marwan Aljubeh , Mateusz Litwin , Matthew Zeng , Max Johnson , Maya Shetty , Mayank Gupta , Meghan Shah , Mehmet Yatbaz , Meng Jia Yang , Mengchao Zhong , Mia Glaese , Mianna Chen , Michael Janner , Michael Lampe , Michael Petrov , Michael Wu , Michele Wang , Michelle Fradin , Michelle Pokrass , Miguel Castro , Miguel Oom Temudo de Castro , Mikhail Pavlov , Miles Brundage , Miles Wang , Minal Khan , Mira Murati , Mo Bavarian , Molly Lin , Murat Yesildal , Nacho Soto , Natalia Gimelshein , Natalie Cone , Natalie Staudacher , Natalie Summers , Natan LaFontaine , Neil Chowdhury , Nick Ryder , Nick Stathas , Nick Turley , Nik Tezak , Niko Felix , Nithanth Kudige , Nitish Keskar , Noah Deutsch , Noel Bundick , Nora Puckett , Ofir Nachum , Ola Okelola , Oleg Boiko , Oleg Murk , Oliver Jaffe , Olivia Watkins , Olivier Godement , Owen Campbell-Moore , Patrick Chao , Paul McMillan , Pavel Belov , Peng Su , Peter Bak , Peter Bakkum , Peter Deng , Peter Dolan , Peter Hoeschele , Peter Welinder , Phil Tillet , Philip Pronin , Philippe Tillet , Prafulla Dhariwal , Qiming Yuan , Rachel Dias , Rachel Lim , Rahul Arora , Rajan Troll , Randall Lin , Rapha Gontijo Lopes , Raul Puri , Reah Miyara , Reimar Leike , Renaud Gaubert , Reza Zamani , Ricky Wang , Rob Donnelly , Rob Honsby , Rocky Smith , Rohan Sahai , Rohit Ramchandani , Romain Huet , Rory Carmichael , Rowan Zellers , Roy Chen , Ruby Chen , Ruslan Nigmatullin , Ryan Cheu , Saachi Jain , Sam Altman , Sam Schoenholz , Sam Toizer , Samuel Miserendino , Sandhini Agarwal , Sara Culver , Scott Ethersmith , Scott Gray , Sean Grove , Sean Metzger , Shamez Hermani , Shantanu Jain , Shengjia Zhao , Sherwin Wu , Shino Jomoto , Shirong Wu , Shuaiqi , Xia , Sonia Phene , Spencer Papay , Srinivas Narayanan , Steve Coffey , Steve Lee , Stewart Hall , Suchir Balaji , Tal Broda , Tal Stramer , Tao Xu , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Cunninghman , Thomas Degry , Thomas Dimson , Thomas Raoux , Thomas Shadwell , Tianhao Zheng , Todd Underwood , Todor Markov , Toki Sherbakov , Tom Rubin , Tom Stasi , Tomer Kaftan , Tristan Heywood , Troy Peterson , Tyce Walters , Tyna Eloundou , Valerie Qi , Veit Moeller , Vinnie Monaco , Vishal Kuo , Vlad Fomenko , Wayne Chang , Weiyi Zheng , Wenda Zhou , Wesam Manassra , Will Sheu , Wojciech Zaremba , Yash Patil , Yilei Qian , Yongjik Kim , Youlong Cheng , Yu Zhang , Yuchen He , Yuchen Zhang , Yujia Jin , Yunxing Dai , Yury Malkov

Drawing is an art that enables people to express their imagination and emotions. However, individuals usually face challenges in drawing, especially when translating conceptual ideas into visually coherent representations and bridging the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Trong-Vu Hoang , Quang-Binh Nguyen , Duy-Nam Ly , Khanh-Duy Le , Tam V. Nguyen , Minh-Triet Tran , Trung-Nghia Le

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Bingqing Zhang , Zhuo Cao , Heming Du , Yang Li , Xue Li , Jiajun Liu , Sen Wang