English
Related papers

Related papers: DetectiumFire: A Comprehensive Multi-modal Dataset…

200 papers

Accurate dialogue description in audiovisual video captioning is crucial for downstream understanding and generation tasks. However, existing models generally struggle to produce faithful dialogue descriptions within audiovisual captions.…

Computation and Language · Computer Science 2026-01-28 Xinlong Chen , Weihong Lin , Jingyun Hua , Linli Yao , Yue Ding , Bozhou Li , Bohan Zeng , Yang Shi , Qiang Liu , Yuanxing Zhang , Pengfei Wan , Liang Wang , Tieniu Tan

The rapid evolution of large language models (LLMs) is transforming artificial intelligence into autonomous research partners, yet a critical gap persists in complex scientific domains such as combustion modeling. Here, practical AI…

Machine Learning · Computer Science 2026-01-06 Ke Xiao , Haoze Zhang , Runze Mao , Han Li , Zhi X. Chen

Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their performance leaps lie more insidious and destructive safety risks, namely deception. Unlike…

Artificial Intelligence · Computer Science 2026-05-28 Sitong Fang , Shiyi Hou , Kaile Wang , Boyuan Chen , Donghai Hong , Jiayi Zhou , Josef Dai , Yaodong Yang , Jiaming Ji

Existing Multimodal Large Language Models (MLLMs) increasingly emphasize complex understanding of various visual elements, including multiple objects, text information, and spatial relations. Their development for comprehensive visual…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Xiaotong Li , Fan Zhang , Haiwen Diao , Yueze Wang , Xinlong Wang , Ling-Yu Duan

In many cyber-physical systems, imaging can be an important but expensive or 'difficult to deploy' sensing modality. One such example is detecting combustion instability using flame images, where deep learning frameworks have demonstrated…

Machine Learning · Computer Science 2021-10-07 Tryambak Gangopadhyay , Vikram Ramanan , Satyanarayanan R Chakravarthy , Soumik Sarkar

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Dark humor in online memes poses unique challenges due to its reliance on implicit, sensitive, and culturally contextual cues. To address the lack of resources and methods for detecting dark humor in multimodal content, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Sai Kartheek Reddy Kasu , Mohammad Zia Ur Rehman , Shahid Shafi Dar , Rishi Bharat Junghare , Dhanvin Sanjay Namboodiri , Nagendra Kumar

With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Cai Yu , Shan Jia , Xiaomeng Fu , Jin Liu , Jiahe Tian , Jiao Dai , Xi Wang , Siwei Lyu , Jizhong Han

We present a new four-pronged approach to build firefighter's situational awareness for the first time in the literature. We construct a series of deep learning frameworks built on top of one another to enhance the safety, efficiency, and…

Computer Vision and Pattern Recognition · Computer Science 2021-11-10 Manish Bhattarai

In recent decades, wildfires, as widespread and extremely destructive natural disasters, have caused tremendous property losses and fatalities, as well as extensive damage to forest ecosystems. Many fire risk assessment projects have been…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Shuchang Shen , Sachith Seneviratne , Xinye Wanyan , Michael Kirley

Multimodal information extraction (IE) tasks have attracted increasing attention because many studies have shown that multimodal information benefits text information extraction. However, existing multimodal IE datasets mainly focus on…

Computation and Language · Computer Science 2024-12-17 Jiang Liu , Bobo Li , Xinran Yang , Na Yang , Hao Fei , Mingyao Zhang , Fei Li , Donghong Ji

In an era of rapidly evolving internet technology, the surge in multimodal content, including videos, has expanded the horizons of online communication. However, the detection of toxic content in this diverse landscape, particularly in…

Artificial Intelligence · Computer Science 2024-07-16 Krishanu Maity , A. S. Poornash , Sriparna Saha , Pushpak Bhattacharyya

To advance foundation Large Language Models (LLMs) for combustion science, this study presents the first end-to-end framework for developing domain-specialized models for the combustion community. The framework comprises an AI-ready…

Computation and Language · Computer Science 2026-03-06 Zonglin Yang , Runze Mao , Tianhao Wu , Han Li , QingGuo Zhou , Zhi X. Chen

Due to climate change and the disruption of ecosystems worldwide, wildfires are increasingly impacting environment, infrastructure, and human lives globally. Additionally, an exacerbating climate crisis means that these losses would…

Machine Learning · Computer Science 2025-09-16 Rohan Tan Bhowmik , Youn Soo Jung , Juan Aguilera , Mary Prunicki , Kari Nadeau

Cross-modal retrieval relies on accurate models to retrieve relevant results for queries across modalities such as image, text, and video. In this paper, we build upon previous work by tackling the difficulty of evaluating models both…

Multimedia · Computer Science 2020-10-20 Tony Zhao , Jaeyoung Choi , Gerald Friedland

Fire is considered one of the most serious threats to human lives which results in a high probability of fatalities. Those severe consequences stem from the heavy smoke emitted from a fire that mostly restricts the visibility of escaping…

Computer Vision and Pattern Recognition · Computer Science 2023-07-11 Truong-Dong Do , Nghe-Nhan Truong , My-Ha Le

An increasingly common expression of online hate speech is multimodal in nature and comes in the form of memes. Designing systems to automatically detect hateful content is of paramount importance if we are to mitigate its undesirable…

Recent advancement of large language models (LLMs) represents a transformational capability at the frontier of artificial intelligence. However, LLMs are generalized models, trained on extensive text corpus, and often struggle to provide…

We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 TsaiChing Ni , ZhenQi Chen , YuanFu Yang

This study proposes an enhanced dual-model YOLOv8 framework for intelligent fire detection and proximity-aware risk assessment, extending conventional vision-based monitoring beyond simple detection to actionable hazard prioritization. The…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Ammar K. AlMhdawi , Nonso Nnamoko , Alaa Mashan Ubaid