中文
相关论文

相关论文: GETReason: Enhancing Image Context Extraction thro…

200 篇论文

One of the essential human skills is the ability to seamlessly build an inner representation of the world. By exploiting this representation, humans are capable of easily finding consensus between visual, auditory and linguistic…

计算与语言 · 计算机科学 2023-05-23 Mihai Masala , Nicolae Cudlenco , Traian Rebedea , Marius Leordeanu

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Whole slide image (WSI) analysis presents significant computational challenges due to the massive number of patches in gigapixel images. While transformer architectures excel at modeling long-range correlations through self-attention, their…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zhengrui Guo , Qichen Sun , Jiabo Ma , Lishuang Feng , Jinzhuo Wang , Hao Chen

The robustness of semantic segmentation on edge cases of traffic scene is a vital factor for the safety of intelligent transportation. However, most of the critical scenes of traffic accidents are extremely dynamic and previously unseen,…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Jiaming Zhang , Kailun Yang , Rainer Stiefelhagen

Text-guided image retrieval is to incorporate conditional text to better capture users' intent. Traditionally, the existing methods focus on minimizing the embedding distances between the source inputs and the targeted image, using the…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Junyang Chen , Hanjiang Lai

There is an increasing requirement for efficient image retargeting techniques to adapt the content to various forms of digital media. With rapid growth of mobile communications and dynamic web page layouts, one often needs to resize the…

图形学 · 计算机科学 2015-08-14 Sukrit Shankar , Pier Luigi Dragotti

Image captioning systems often produce generic descriptions that fail to capture event-level semantics which are crucial for applications like news reporting and digital archiving. We present ReCap, a novel pipeline for event-enriched image…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Thinh-Phuc Nguyen , Thanh-Hai Nguyen , Gia-Huy Dinh , Lam-Huy Nguyen , Minh-Triet Tran , Trung-Nghia Le

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story, where the images should be realistic and keep global consistency across dynamic scenes and characters. Current works face the…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Bowen Li , Thomas Lukasiewicz

Large text-to-image models have achieved astonishing performance in synthesizing diverse and high-quality images guided by texts. With detail-oriented conditioning control, even finer-grained spatial control can be achieved. However, some…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yuhe Liu , Mengxue Kang , Zengchang Qin , Xiangxiang Chu

The rapid progress of visual generative models has made AI-generated images increasingly difficult to distinguish from authentic ones, posing growing risks to social trust and information integrity. This motivates detectors that are not…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Huangsen Cao , Qin Mei , Zhiheng Li , Yuxi Li , Zhan Meng , Ying Zhang , Chen Li , Zhimeng Zhang , Xin Ding , Yongwei Wang , Jing Lyu , Fei Wu

Document-level relation extraction aims to extract relations among entities within a document. Different from sentence-level relation extraction, it requires reasoning over multiple sentences across a document. In this paper, we propose…

计算与语言 · 计算机科学 2020-09-30 Shuang Zeng , Runxin Xu , Baobao Chang , Lei Li

The rapid growth in the volume, variety, and velocity of geospatial data has created data ecosystems that are highly distributed, heterogeneous, and semantically inconsistent. Existing data catalogs, portals, and infrastructures still rely…

人工智能 · 计算机科学 2026-03-25 Ruixiang Liu , Zhenlong Li , Ali Khosravi Kazazi

To understand a document with multiple events, event-event relation extraction (ERE) emerges as a crucial task, aiming to discern how natural events temporally or structurally associate with each other. To achieve this goal, our work…

信息论 · 计算机科学 2024-12-20 Peixin Huang , Xiang Zhao , Minghao Hu , Zhen Tan , Weidong Xiao

Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their correspondence to the world.…

计算与语言 · 计算机科学 2017-10-03 Stephanie Zhou , Alane Suhr , Yoav Artzi

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yushuo Zheng , Zicheng Zhang , Huiyu Duan , Chunyi Li , Zijian Chen , Ziheng Jia , Yue Shi , Ke Gu , Xiongkuo Min , Guangtao Zhai

Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Hammad A. Ayyubi , Christopher Thomas , Lovish Chum , Rahul Lokesh , Long Chen , Yulei Niu , Xudong Lin , Xuande Feng , Jaywon Koo , Sounak Ray , Shih-Fu Chang

Event-based moving object detection is a challenging task, where static background and moving object are mixed together. Typically, existing methods mainly align the background events to the same spatial coordinate system via motion…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Hanyu Zhou , Zhiwei Shi , Hao Dong , Shihan Peng , Yi Chang , Luxin Yan

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought reasoning, and…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Yuxiang Ji , Yong Wang , Ziyu Ma , Yiming Hu , Hailang Huang , Xuecai Hu , Guanhua Chen , Liaoni Wu , Xiangxiang Chu

This paper proposes a framework for representing and reasoning causality between geographic events by introducing the notion of Geo-Situation. This concept links to observational snapshots that represent sets of conditions, and either acts…

人工智能 · 计算机科学 2022-06-29 Shirly Stephen , Wenwen Li , Torsten Hahmann

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Existing spatial-temporal benchmarks mainly focus on egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Qinghongbing Xie , Zhaoyuan Xia , Feng Zhu , Lijun Gong , Ziyue Li , Rui Zhao , Long Zeng