中文
相关论文

相关论文: ViRED: Prediction of Visual Relations in Engineeri…

200 篇论文

In architecture and computer-aided design, wireframes (i.e., line-based models) are widely used as basic 3D models for design evaluation and fast design iterations. However, unlike a full design file, a wireframe model lacks critical…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Yuan Xue , Zihan Zhou , Xiaolei Huang

Data quality assessment process is essential to ensure reliable analytical outcomes. This process depends on human supervision-driven approaches since it is impossible to determine a defect based only on data. Visualization systems belong…

数据库 · 计算机科学 2018-09-27 João Marcelo Borovina Josko , João Eduardo Ferreira

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich…

计算机视觉与模式识别 · 计算机科学 2025-05-27 GuangHao Meng , Sunan He , Jinpeng Wang , Tao Dai , Letian Zhang , Jieming Zhu , Qing Li , Gang Wang , Rui Zhang , Yong Jiang

Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, large-scale pre-training approaches have been developed and…

计算机视觉与模式识别 · 计算机科学 2023-11-10 Cheng Yang , Rui Xu , Ye Guo , Peixiang Huang , Yiru Chen , Wenkui Ding , Zhongyuan Wang , Hong Zhou

Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades. Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Yu Zhao , Jianguo Wei , Zhichao Lin , Yueheng Sun , Meishan Zhang , Min Zhang

Visual-frame prediction is a pixel-dense prediction task that infers future frames from past frames. Lacking of appearance details, low prediction accuracy and high computational overhead are still major problems with current models or…

计算机视觉与模式识别 · 计算机科学 2022-11-16 Chaofan Ling , Junpei Zhong , Weihua Li

Network visualizations are commonly used to analyze relationships in various contexts. To efficiently explore a network visualization, the user needs to quickly navigate to different parts of the network and analyze local details. Recent…

人机交互 · 计算机科学 2023-03-29 Helen H. Huang , Hanspeter Pfister , Yalong Yang

In order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects. We propose…

计算机视觉与模式识别 · 计算机科学 2019-10-11 Linjie Li , Zhe Gan , Yu Cheng , Jingjing Liu

We propose a novel model to address the task of Visual Dialog which exhibits complex dialog structures. To obtain a reasonable answer based on the current question and the dialog history, the underlying semantic dependencies between dialog…

计算机视觉与模式识别 · 计算机科学 2019-05-30 Zilong Zheng , Wenguan Wang , Siyuan Qi , Song-Chun Zhu

The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relations. Visual language…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Qi Liu , Weiying Xue , Yuxiao Wang , Zhenao Wei

This work proposes, implements and tests an immersive framework upon Virtual Reality (VR) for comprehension, knowledge development and learning process assisting an improved perception of complex spatial arrangements in AEC in comparison to…

人机交互 · 计算机科学 2022-09-23 Michael Kraus , Romana Rust , Maximilian Rietschel , Daniel Hall

Recent advancements in vision models have greatly improved their ability to handle complex chart understanding tasks, like chart captioning and question answering. However, it remains challenging to assess how these models process charts.…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Soohyun Lee , Minsuk Chang , Seokhyeon Park , Jinwook Seo

Recognizing instruments' interactions with tissues is essential for building context-aware AI assistants in robotic surgery. Vision-language models (VLMs) have opened a new avenue for surgical perception and achieved better generalization…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Jiajun Cheng , Xiaofan Yu , Subarna Tripathi , Sainan Liu , Shan Lin

One of the most common modes of representing engineering schematics are Piping and Instrumentation diagrams (P&IDs) that describe the layout of an engineering process flow along with the interconnected process equipment. Over the years,…

计算机视觉与模式识别 · 计算机科学 2019-02-01 Rohit Rahul , Shubham Paliwal , Monika Sharma , Lovekesh Vig

Video relation detection forms a new and challenging problem in computer vision, where subjects and objects need to be localized spatio-temporally and a predicate label needs to be assigned if and only if there is an interaction between the…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Shuo Chen , Pascal Mettes , Cees G. M. Snoek

Visual question answering (VQA) is a challenging task to provide an accurate natural language answer given an image and a natural language question about the image. It involves multi-modal learning, i.e., computer vision (CV) and natural…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Luoqian Jiang , Yifan He , Jian Chen

Document Understanding is an evolving field in Natural Language Processing (NLP). In particular, visual and spatial features are essential in addition to the raw text itself and hence, several multimodal models were developed in the field…

计算与语言 · 计算机科学 2024-04-18 Wiam Adnan , Joel Tang , Yassine Bel Khayat Zouggari , Seif Edinne Laatiri , Laurent Lam , Fabien Caspani

Requirements Engineering (RE) is closely tied to other development activities and is at the heart and foundation of every software development process. This makes RE the most data and communication-intensive activity compared to other…

软件工程 · 计算机科学 2017-07-10 Zahra Shakeri Hossein Abad , Alex Shymka , Jenny Le , Noor Hammad , Guenther Ruhe

Visual modes of communication are ubiquitous in modern life --- from maps to data plots to political cartoons. Here we investigate drawing, the most basic form of visual communication. Participants were paired in an online environment to…

其他计算机科学 · 计算机科学 2019-12-17 Judith Fan , Robert Hawkins , Mike Wu , Noah Goodman

Parsing chemical reaction diagrams from scientific literature is challenging due to heterogeneous layouts, intertwined visual elements, and the difficulty of integrating recognition and reasoning. Existing vision-language models advance…

人工智能 · 计算机科学 2026-05-28 Chuang Tang , Chenhao Lin , Yin Xu , Hao Wang , Jinrui Zhou , Xin Li , Mingjun Xiao , Enhong Chen
‹ 上一页 1 8 9 10 下一页 ›