中文
相关论文

相关论文: The KIT Motion-Language Dataset

200 篇论文

We propose a novel multi-task learning system that combines appearance and motion cues for a better semantic reasoning of the environment. A unified architecture for joint vehicle detection and motion segmentation is introduced. In this…

计算机视觉与模式识别 · 计算机科学 2018-10-19 Mennatullah Siam , Heba Mahgoub , Mohamed Zahran , Senthil Yogamani , Martin Jagersand , Ahmad El-Sallab

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions…

Social infrastructure and other built environments are increasingly expected to support well-being and community resilience by enabling social interaction. Yet in civil and built-environment research, there is no consistent and…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Cheyu Lin , Katherine A. Flanigan , Sirajum Munir

Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to language variation, thus…

计算与语言 · 计算机科学 2026-03-17 Xinyue Ma , Pol Pastells , Mireia Farrús , Mariona Taulé

Learning a joint language-visual embedding has a number of very appealing properties and can result in variety of practical application, including natural language image/video annotation and search. In this work, we study three different…

计算机视觉与模式识别 · 计算机科学 2016-09-27 Atousa Torabi , Niket Tandon , Leonid Sigal

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Jiaxin Lu , Chun-Hao Paul Huang , Uttaran Bhattacharya , Qixing Huang , Yi Zhou

For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other. Recently, however, the community has realized that…

计算机视觉与模式识别 · 计算机科学 2022-06-06 Yiyi Liao , Jun Xie , Andreas Geiger

We propose a novel large-scale emotional dialogue dataset, consisting of 1M dialogues retrieved from the OpenSubtitles corpus and annotated with 32 emotions and 9 empathetic response intents using a BERT-based fine-grained dialogue emotion…

计算与语言 · 计算机科学 2020-12-29 Anuradha Welivita , Yubo Xie , Pearl Pu

Our goal is to generate realistic human motion from natural language. Modern methods often face a trade-off between model expressiveness and text-to-motion alignment. Some align text and motion latent spaces but sacrifice expressiveness;…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Nefeli Andreou , Xi Wang , Victoria Fernández Abrevaya , Marie-Paule Cani , Yiorgos Chrysanthou , Vicky Kalogeiton

Generating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging research area within the domains of computer vision, graphics, and robotics. Despite significant advancements in…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Ke Fan , Shunlin Lu , Minyue Dai , Runyi Yu , Lixing Xiao , Zhiyang Dou , Junting Dong , Lizhuang Ma , Jingbo Wang

The last decade has witnessed the emergence of massive mobility data sets, such as tracks generated by GPS devices, call detail records, and geo-tagged posts from social media platforms. These data sets have fostered a vast scientific…

物理与社会 · 物理学 2021-06-07 Luca Pappalardo , Filippo Simini , Gianni Barlacchi , Roberto Pellungrini

In the field of robotic manipulation, deep imitation learning is recognized as a promising approach for acquiring manipulation skills. Additionally, learning from diverse robot datasets is considered a viable method to achieve versatility…

机器人学 · 计算机科学 2024-03-20 Heecheol Kim , Yoshiyuki Ohmura , Yasuo Kuniyoshi

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to investigating the…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Yansong Tang , Jiwen Lu , Jie Zhou

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

The study illustrates a first step towards an ongoing work aimed at developing a dataset of dialogues potentially useful for customer service conversation management between humans and AI chatbots. The approach exploits ChatGPT 3.5 to…

人机交互 · 计算机科学 2025-01-03 Alfredo Cuzzocrea , Giovanni Pilato , Pablo Garcia Bringas

Conversational information seeking has evolved rapidly in the last few years with the development of Large Language Models (LLMs), providing the basis for interpreting and responding in a naturalistic manner to user requests. The extended…

信息检索 · 计算机科学 2024-05-07 Mohammad Aliannejadi , Zahra Abbasiantaeb , Shubham Chatterjee , Jeffery Dalton , Leif Azzopardi

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional…

声音 · 计算机科学 2026-01-14 Runchuan Ye , Yixuan Zhou , Renjie Yu , Zijian Lin , Kehan Li , Xiang Li , Xin Liu , Guoyang Zeng , Zhiyong Wu

Large language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large language models are not good at math theorem proving using formal…

计算与语言 · 计算机科学 2025-06-19 Huaiyuan Ying , Zijian Wu , Yihan Geng , Zheng Yuan , Dahua Lin , Kai Chen

With the rapid advancement of Natural Language Processing in recent years, numerous studies have shown that generic summaries generated by Large Language Models (LLMs) can sometimes surpass those annotated by experts, such as journalists,…

In this paper, we introduce a new benchmark dataset for the challenging writing in the air (WiTA) task -- an elaborate task bridging vision and NLP. WiTA implements an intuitive and natural writing method with finger movement for…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Ue-Hwan Kim , Yewon Hwang , Sun-Kyung Lee , Jong-Hwan Kim
‹ 上一页 1 8 9 10 下一页 ›