English
Related papers

Related papers: CL-HOI: Cross-Level Human-Object Interaction Disti…

200 papers

Modeling spatial-temporal relations is imperative for recognizing human actions, especially when a human is interacting with objects, while multiple objects appear around the human differently over time. Most existing action recognition…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Muna Almushyti , Frederick W. Li

Large Vision-Language Models (VLMs) have achieved remarkable success across diverse multimodal tasks but remain vulnerable to hallucinations rooted in inherent language bias. Despite recent progress, existing hallucination mitigation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yilin Yang , Zhenghui Guo , Yuke Wang , Omprakash Gnawali , Sheng Di , Chengming Zhang

Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts have developed various attention-based modules for solving…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Siyu Zhang , Yeming Chen , Yaoru Sun , Fang Wang , Haibo Shi , Haoran Wang

In Large Visual Language Models (LVLMs), the efficacy of In-Context Learning (ICL) remains limited by challenges in cross-modal interactions and representation disparities. To overcome these challenges, we introduce a novel Visual…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yucheng Zhou , Xiang Li , Qianning Wang , Jianbing Shen

Human-Object Interaction (HOI) detection is a task to localize humans and objects in an image and predict the interactions in human-object pairs. In real-world scenarios, HOI detection models need systematic generalization, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Kentaro Takemoto , Moyuru Yamada , Tomotake Sasaki , Hisanao Akima

Human-object interaction (HOI) video generation has garnered increasing attention due to its promising applications in digital humans, e-commerce, advertising, and robotics imitation learning. However, existing methods face two critical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Bangya Liu , Xinyu Gong , Zelin Zhao , Ziyang Song , Yulei Lu , Suhui Wu , Jun Zhang , Suman Banerjee , Hao Zhang

We study in this paper the problem of novel human-object interaction (HOI) detection, aiming at improving the generalization ability of the model to unseen scenarios. The challenge mainly stems from the large compositional space of objects…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Yuhang Song , Wenbo Li , Lei Zhang , Jianwei Yang , Emre Kiciman , Hamid Palangi , Jianfeng Gao , C. -C. Jay Kuo , Pengchuan Zhang

Understanding humans from LiDAR point clouds is one of the most critical tasks in autonomous driving due to its close relationships with pedestrian safety, yet it remains challenging in the presence of diverse human-object interactions and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Daniel Sungho Jung , Dohee Cho , Kyoung Mu Lee

The widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Trung Thanh Nguyen , Yasutomo Kawanishi , Vijay John , Takahiro Komamizu , Ichiro Ide

Open-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, eg, CLIP, to inherit its generalizable recognition ability so that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Shenghao Fu , Junkai Yan , Qize Yang , Xihan Wei , Xiaohua Xie , Wei-Shi Zheng

Recovering 3D Human-Object Interaction (HOI) from single color images is challenging due to depth ambiguities, occlusions, and the huge variation in object shape and appearance. Thus, past work requires controlled settings such as known…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Alpár Cseke , Shashank Tripathi , Sai Kumar Dwivedi , Arjun Lakshmipathy , Agniv Chatterjee , Michael J. Black , Dimitrios Tzionas

Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction. However, most benchmarks are non-interactive in that a model produces output without human involvement.…

While diffusion models and large-scale motion datasets have advanced text-driven human motion synthesis, extending these advances to 4D human-object interaction (HOI) remains challenging, mainly due to the limited availability of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Shujia Li , Haiyu Zhang , Xinyuan Chen , Yaohui Wang , Yutong Ban

Improving the visual understanding ability of vision-language models (VLMs) is crucial for enhancing their performance across various tasks. While using multiple pretrained visual experts has shown great promise, it often incurs significant…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Yimu Wang , Mozhgan Nasr Azadani , Sean Sedwards , Krzysztof Czarnecki

Large-scale pre-trained Vision-Language Models (VLMs), such as CLIP, establish the correlation between texts and images, achieving remarkable success on various downstream tasks with fine-tuning. In existing fine-tuning methods, the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Yi Zhang , Ce Zhang , Yushun Tang , Zhihai He

We propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering. Dual-encoder models have a faster inference speed than…

Computation and Language · Computer Science 2022-10-18 Zekun Wang , Wenhui Wang , Haichao Zhu , Ming Liu , Bing Qin , Furu Wei

We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description. This task is complicated by the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Thomas Hanwen Zhu , Ruining Li , Tomas Jakab

The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Andrey Palaev , Adil Khan , Syed M. Ahsan Kazmi

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang

Learning-based methods to understand and model hand-object interactions (HOI) require a large amount of high-quality HOI data. One way to create HOI data is to transfer hand poses from a source object to another based on the objects'…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Qiaochu Wang , Chufeng Xiao , Manfred Lau , Hongbo Fu