English
Related papers

Related papers: ChinaOpen: A Dataset for Open-world Multimodal Lea…

200 papers

Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction between image-text pairs, by assuming that there exists strong…

Multimodal interleaved datasets featuring free-form interleaved sequences of images and text are crucial for training frontier large multimodal models (LMMs). Despite the rapid progression of open-source LMMs, there remains a pronounced…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Anas Awadalla , Le Xue , Oscar Lo , Manli Shu , Hannah Lee , Etash Kumar Guha , Matt Jordan , Sheng Shen , Mohamed Awadalla , Silvio Savarese , Caiming Xiong , Ran Xu , Yejin Choi , Ludwig Schmidt

With the development of the Internet, more and more people get accustomed to online shopping. When communicating with customer service, users may express their requirements by means of text, images, and videos, which precipitates the need…

Computation and Language · Computer Science 2021-09-28 Nan Zhao , Haoran Li , Youzheng Wu , Xiaodong He , Bowen Zhou

Non-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques. Given a context, current systems are able to yield a relevant and…

Computation and Language · Computer Science 2020-04-10 Leyang Cui , Yu Wu , Shujie Liu , Yue Zhang , Ming Zhou

Internet memes have gained significant influence in communicating political, psychological, and sociocultural ideas. While memes are often humorous, there has been a rise in the use of memes for trolling and cyberbullying. Although a wide…

Computation and Language · Computer Science 2024-01-19 Prince Jha , Krishanu Maity , Raghav Jain , Apoorv Verma , Sriparna Saha , Pushpak Bhattacharyya

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first…

We introduce mTVR, a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips. The dataset is collected by extending the popular TVR dataset (in English) with paired…

Computation and Language · Computer Science 2021-08-03 Jie Lei , Tamara L. Berg , Mohit Bansal

The topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified…

Computer Vision and Pattern Recognition · Computer Science 2023-02-15 Shizhe Chen , Jia Chen , Qin Jin , Alexander Hauptmann

Compared with the domain-specific model, the vision-language pre-training models (VLPMs) have shown superior performance on downstream tasks with fast fine-tuning process. For example, ERNIE-ViL, Oscar and UNIMO trained VLPMs with a uniform…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Sha Yuan , Shuai Zhao , Jiahong Leng , Zhao Xue , Hanyu Zhao , Peiyu Liu , Zheng Gong , Wayne Xin Zhao , Junyi Li , Jie Tang

Humans express feelings or emotions via different channels. Take language as an example, it entails different sentiments under different visual-acoustic contexts. To precisely understand human intentions as well as reduce the…

Artificial Intelligence · Computer Science 2021-11-17 Ting Wu , Junjie Peng , Wenqiang Zhang , Huiran Zhang , Chuanshuai Ma , Yansong Huang

Recently, significant public efforts have been directed towards developing low-cost models with capabilities akin to ChatGPT, thereby fostering the growth of open-source conversational models. However, there remains a scarcity of…

Computation and Language · Computer Science 2023-04-18 Yunjie Ji , Yan Gong , Yong Deng , Yiping Peng , Qiang Niu , Baochang Ma , Xiangang Li

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

As a kind of new expression elements, Internet memes are popular and extensively used in online chatting scenarios since they manage to make dialogues vivid, moving, and interesting. However, most current dialogue researches focus on…

Computation and Language · Computer Science 2021-09-07 Zhengcong Fei , Zekang Li , Jinchao Zhang , Yang Feng , Jie Zhou

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge capacity and reliability. In this paper, we introduce the…

Text-to-video generative models convert textual prompts into dynamic visual content, offering wide-ranging applications in film production, gaming, and education. However, their real-world performance often falls short of user expectations.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Wenhao Wang , Yi Yang

This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying…

Computation and Language · Computer Science 2023-12-05 Keito Kudo , Haruki Nagasawa , Jun Suzuki , Nobuyuki Shimizu

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing large-scale pre-training datasets for language models, which…

In this article, we create a system called AI-EVL. This is an annotated-based learning system. We extend AI to learning experience. If a user from the main YouTube page browses YouTube videos and a user from the AI-EVL system does the same,…

Information Retrieval · Computer Science 2022-03-22 Faeze Gholamrezaie , Melika Bahman-Abadi , M. B. Ghaznavi-Ghoushchi

Most large language models are fine-tuned using either expensive human-annotated data or GPT-4 generated data which cannot guarantee performance in certain domains. We argue that although the web-crawled data often has formatting errors…

Computation and Language · Computer Science 2024-08-16 Jing Zhou , Chenglin Jiang , Wei Shen , Xiao Zhou , Xiaonan He
‹ Prev 1 3 4 5 6 7 10 Next ›