English
Related papers

Related papers: Myriad People Open Source Software for New Media A…

200 papers

Existing AI-generated dance methods primarily train on motion capture data from solo dance performances, but a critical feature of dance in nearly any genre is the interaction of two or more bodies in space. Moreover, many works at the…

Machine Learning · Computer Science 2025-03-07 Zixuan Wang , Luis Zerkowski , Ilya Vidrin , Mariel Pettee

This paper provides guidance for building and maintaining infrastructure for participatory AI efforts by sharing reflections on building World Wide Dishes (WWD), a bottom-up, community-led image and text dataset of culinary dishes and…

The significance and abundance of data are increasing due to the growing digital data generated from social media, sensors, scholarly literature, patents, different forms of documents published online, databases, product manuals, etc.…

Machine Learning · Computer Science 2022-05-23 Workneh Yilma Ayele

We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt…

In this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source separation and localization, cross-modal correspondences,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Juan F. Montesinos , Olga Slizovskaia , Gloria Haro

The remarkable ease of use of diffusion models for image generation has led to a proliferation of synthetic content online. While these models are often employed for legitimate purposes, they are also used to generate fake images that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Giulia Bertazzini , Daniele Baracchi , Dasara Shullani , Isao Echizen , Alessandro Piva

People often create art by following an artistic workflow involving multiple stages that inform the overall design. If an artist wishes to modify an earlier decision, significant work may be required to propagate this new decision forward…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Hung-Yu Tseng , Matthew Fisher , Jingwan Lu , Yijun Li , Vladimir Kim , Ming-Hsuan Yang

With the advent of open source software, a veritable treasure trove of previously proprietary software development data was made available. This opened the field of empirical software engineering research to anyone in academia. Data that is…

Software Engineering · Computer Science 2022-04-19 Adam Tutko , Austin Z. Henley , Audris Mockus

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting…

Robotics · Computer Science 2025-04-23 Alexander Khazatsky , Karl Pertsch , Suraj Nair , Ashwin Balakrishna , Sudeep Dasari , Siddharth Karamcheti , Soroush Nasiriany , Mohan Kumar Srirama , Lawrence Yunliang Chen , Kirsty Ellis , Peter David Fagan , Joey Hejna , Masha Itkina , Marion Lepert , Yecheng Jason Ma , Patrick Tree Miller , Jimmy Wu , Suneel Belkhale , Shivin Dass , Huy Ha , Arhan Jain , Abraham Lee , Youngwoon Lee , Marius Memmel , Sungjae Park , Ilija Radosavovic , Kaiyuan Wang , Albert Zhan , Kevin Black , Cheng Chi , Kyle Beltran Hatch , Shan Lin , Jingpei Lu , Jean Mercat , Abdul Rehman , Pannag R Sanketi , Archit Sharma , Cody Simpson , Quan Vuong , Homer Rich Walke , Blake Wulfe , Ted Xiao , Jonathan Heewon Yang , Arefeh Yavary , Tony Z. Zhao , Christopher Agia , Rohan Baijal , Mateo Guaman Castro , Daphne Chen , Qiuyu Chen , Trinity Chung , Jaimyn Drake , Ethan Paul Foster , Jensen Gao , Vitor Guizilini , David Antonio Herrera , Minho Heo , Kyle Hsu , Jiaheng Hu , Muhammad Zubair Irshad , Donovon Jackson , Charlotte Le , Yunshuang Li , Kevin Lin , Roy Lin , Zehan Ma , Abhiram Maddukuri , Suvir Mirchandani , Daniel Morton , Tony Nguyen , Abigail O'Neill , Rosario Scalise , Derick Seale , Victor Son , Stephen Tian , Emi Tran , Andrew E. Wang , Yilin Wu , Annie Xie , Jingyun Yang , Patrick Yin , Yunchu Zhang , Osbert Bastani , Glen Berseth , Jeannette Bohg , Ken Goldberg , Abhinav Gupta , Abhishek Gupta , Dinesh Jayaraman , Joseph J Lim , Jitendra Malik , Roberto Martín-Martín , Subramanian Ramamoorthy , Dorsa Sadigh , Shuran Song , Jiajun Wu , Michael C. Yip , Yuke Zhu , Thomas Kollar , Sergey Levine , Chelsea Finn

Creating MIDI music can be a practical challenge. In the past, working with it was difficult and frustrating to all but the most accomplished and determined. Now, however, we are offering a powerful Visual Basic program called MIDI-LAB,…

Software Engineering · Computer Science 2015-03-13 Kai Yang , Xi Zhou

Pedestrian detection and tracking in crowded video sequences have many applications, including autonomous driving, robot navigation and pedestrian flow analysis. However, detecting and tracking pedestrians in high-density crowds face many…

Image and Video Processing · Electrical Eng. & Systems 2025-08-21 Kailai Sun , Xinwei Wang , Shaobo Liu , Qianchuan Zhao , Gao Huang , Chang Liu

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering both, dense multi-view camera captures, and rich facial…

We present a systematic review of 458 papers that report on evaluations in mixed and augmented reality (MR/AR) published in ISMAR, CHI, IEEE VR, and UIST over a span of 11 years (2009-2019). Our goal is to provide guidance for future…

Human-Computer Interaction · Computer Science 2020-10-14 Leonel Merino , Magdalena Schwarzl , Matthias Kraus , Michael Sedlmair , Dieter Schmalstieg , Daniel Weiskopf

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images,…

Computation and Language · Computer Science 2024-10-04 Hossein Aboutalebi , Hwanjun Song , Yusheng Xie , Arshit Gupta , Justin Sun , Hang Su , Igor Shalyminov , Nikolaos Pappas , Siffi Singh , Saab Mansour

Person identification in the wild is very challenging due to great variation in poses, face quality, clothes, makeup and so on. Traditional research, such as face recognition, person re-identification, and speaker recognition, often focuses…

Computer Vision and Pattern Recognition · Computer Science 2019-04-23 Yuanliu Liu , Bo Peng , Peipei Shi , He Yan , Yong Zhou , Bing Han , Yi Zheng , Chao Lin , Jianbin Jiang , Yin Fan , Tingwei Gao , Ganwen Wang , Jian Liu , Xiangju Lu , Danming Xie

Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-source community due to two key challenges: (1) scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Andrew Li , Rahul Thapa , Rahul Chalamala , Qingyang Wu , Kezhen Chen , James Zou

Modern, state-of-the-art deep learning approaches yield human like performance in numerous object detection and classification tasks. The foundation for their success is the availability of training datasets of substantially high quantity,…

This work presents CLIPDraw, an algorithm that synthesizes novel drawings based on natural language input. CLIPDraw does not require any training; rather a pre-trained CLIP language-image encoder is used as a metric for maximizing…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Kevin Frans , L. B. Soros , Olaf Witkowski

The widespread consumer-grade 3D printers and learning resources online enable novices to self-train in remote settings. While troubleshooting plays an essential part of 3D printing, the process remains challenging for many remote novices…

Human-Computer Interaction · Computer Science 2024-02-05 Nahyun Kwon , Tong Sun , Yuyang Gao , Liang Zhao , Xu Wang , Jeeeun Kim , Sungsoo Ray Hong

We introduce MosAIc, an interactive web app that allows users to find pairs of semantically related artworks that span different cultures, media, and millennia. To create this application, we introduce Conditional Image Retrieval (CIR)…

‹ Prev 1 8 9 10 Next ›