English
Related papers

Related papers: PTVD: A Large-Scale Plot-Oriented Multimodal Datas…

200 papers

We present ViDRiP-LLaVA, the first large multimodal model (LMM) in computational pathology that integrates three distinct image scenarios, including single patch images, automatically segmented pathology video clips, and manually segmented…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Trinh T. L. Vuong , Jin Tae Kwak

Multimodality has recently gained attention in the medical domain, where imaging or video modalities may be integrated with biomedical signals or health records. Yet, two challenges remain: balancing the contributions of modalities,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Julie Mordacq , Leo Milecki , Maria Vakalopoulou , Steve Oudot , Vicky Kalogeiton

Detecting visual content on language expression has become an emerging topic in the community. However, in the video domain, the existing setting, i.e., spatial-temporal video grounding (STVG), is formulated to only detect one pre-existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Wei Ji , Xiangyan Liu , Yingfei Sun , Jiajun Deng , You Qin , Ammar Nuwanna , Mengyao Qiu , Lina Wei , Roger Zimmermann

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Ning Cheng , You Li , Jing Gao , Bin Fang , Jinan Xu , Wenjuan Han

Video generation models are revolutionizing content creation, with image-to-video models drawing increasing attention due to their enhanced controllability, visual consistency, and practical applications. However, despite their popularity,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Wenhao Wang , Yi Yang

The ability to choose an appropriate camera view among multiple cameras plays a vital role in TV shows delivery. But it is hard to figure out the statistical pattern and apply intelligent processing due to the lack of high-quality training…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Anyi Rao , Xuekun Jiang , Sichen Wang , Yuwei Guo , Zihao Liu , Bo Dai , Long Pang , Xiaoyu Wu , Dahua Lin , Libiao Jin

Automatic Emotion Detection (ED) aims to build systems to identify users' emotions automatically. This field has the potential to enhance HCI, creating an individualised experience for the user. However, ED systems tend to perform poorly on…

Human-Computer Interaction · Computer Science 2023-07-27 Annanda Sousa , Karen Young , Mathieu D'aquin , Manel Zarrouk , Jennifer Holloway

Non-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques. Given a context, current systems are able to yield a relevant and…

Computation and Language · Computer Science 2020-04-10 Leyang Cui , Yu Wu , Shujie Liu , Yue Zhang , Ming Zhou

Recently, a more challenging state tracking task, Audio-Video Scene-Aware Dialogue (AVSD), is catching an increasing amount of attention among researchers. Different from purely text-based dialogue state tracking, the dialogue in AVSD…

Computation and Language · Computer Science 2020-07-21 Xiangyang Mou , Brandyn Sigouin , Ian Steenstra , Hui Su

Among ubiquitous multimodal data in the real world, text is the modality generated by human, while image reflects the physical world honestly. In a visual understanding application, machines are expected to understand images like human.…

Computation and Language · Computer Science 2021-06-15 Pengda Qin , Yuhong Li , Kefeng Deng , Qiang Wu

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Letian Fu , Gaurav Datta , Huang Huang , William Chung-Ho Panitch , Jaimyn Drake , Joseph Ortiz , Mustafa Mukadam , Mike Lambeta , Roberto Calandra , Ken Goldberg

Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spatiotemporal dynamic reasoning. To study this capability gap, we formulate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Chaoyue Li , Yongxue Xu , Jie Feng , Jiayu Ding

Movie screenplay summarization is challenging, as it requires an understanding of long input contexts and various elements unique to movies. Large language models have shown significant advancements in document summarization, but they often…

Computation and Language · Computer Science 2024-08-13 Rohit Saxena , Frank Keller

Charts provide visual representations of data and are widely used for analyzing information, addressing queries, and conveying insights to others. Various chart-related downstream tasks have emerged recently, such as question-answering and…

Computation and Language · Computer Science 2024-03-15 Ahmed Masry , Mehrad Shahmohammadi , Md Rizwan Parvez , Enamul Hoque , Shafiq Joty

Most recent works on sentiment analysis have exploited the text modality. However, millions of hours of video recordings posted on social media platforms everyday hold vital unstructured information that can be exploited to more effectively…

Computation and Language · Computer Science 2021-03-05 Kia Dashtipour , Mandar Gogate , Erik Cambria , Amir Hussain

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Xin Jin , Qianqian Qiao , Yi Lu , Huaye Wang , Heng Huang , Shan Gao , Jianfei Liu , Rui Li

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constructing them directly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Zhengxian Yang , Shengqi Wang , Shi Pan , Hongshuai Li , Haoxiang Wang , Lin Li , Guanjun Li , Zhengqi Wen , Borong Lin , Jianhua Tao , Tao Yu

Social interactions dominate our perceptions of the world and shape our daily behavior by attaching social meaning to acts as simple and spontaneous as gestures, facial expressions, voice, and speech. People mimic and otherwise respond to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Xiang Zhang , Xiaotian Li , Taoyue Wang , Nan Bi , Xin Zhou , Cody Zhou , Zoie Wang , Andrew Yang , Yuming Su , Jeff Cohn , Qiang Ji , Lijun Yin

Depression has been the leading cause of mental-health illness worldwide. Major depressive disorder (MDD), is a common mental health disorder that affects both psychologically as well as physically which could lead to loss of lives. Due to…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Anupama Ray , Siddharth Kumar , Rutvik Reddy , Prerana Mukherjee , Ritu Garg