中文
相关论文

相关论文: COMPOSER: Compositional Reasoning of Group Activit…

200 篇论文

The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable video synthesis remains challenging due to the large variation…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Xiang Wang , Hangjie Yuan , Shiwei Zhang , Dayou Chen , Jiuniu Wang , Yingya Zhang , Yujun Shen , Deli Zhao , Jingren Zhou

Computer vision has undergone a dramatic revolution in performance, driven in large part through deep features trained on large-scale supervised datasets. However, much of these improvements have focused on static image analysis; video…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Rohit Girdhar , Deva Ramanan

We present PartComposer: a framework for part-level concept learning from single-image examples that enables text-to-image diffusion models to compose novel objects from meaningful components. Existing methods either struggle with…

图形学 · 计算机科学 2025-09-16 Junyu Liu , R. Kenny Jones , Daniel Ritchie

Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability. This work offers a new generation paradigm that allows flexible control of the output image,…

计算机视觉与模式识别 · 计算机科学 2023-02-23 Lianghua Huang , Di Chen , Yu Liu , Yujun Shen , Deli Zhao , Jingren Zhou

Objective: Visual cohort analysis utilizing electronic health record data has become an important tool in clinical assessment of patient outcomes. In this paper, we introduce Composer, a visual analysis tool for orthopedic surgeons to…

人机交互 · 计算机科学 2019-03-20 Jennifer Rogers , Nicholas Spina , Ashley Neese , Rachel Hess , Darrel Brodke , Alexander Lex

The design of a complex system warrants a compositional methodology, i.e., composing simple components to obtain a larger system that exhibits their collective behavior in a meaningful way. We propose an automaton-based paradigm for…

计算机科学中的逻辑 · 计算机科学 2023-02-03 Tobias Kappé , Farhad Arbab , Carolyn Talcott

Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, in practice, composition is often coupled with semantics. As a result, existing methods…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Kai Zou , Zhiwei Zhao , Bin Liu , Nenghai Yu

Transformer architectures have achieved remarkable success across language, vision, and multimodal tasks, and there is growing demand for them to address in-context compositional learning tasks. In these tasks, models solve the target…

机器学习 · 计算机科学 2025-11-26 Wei Chen , Jingxi Yu , Zichen Miao , Qiang Qiu

In this work we present the modular Crowd Simulation Evaluation through Composition framework (CSEC) which provides a quantitative comparison between different pedestrian and crowd simulation approaches. Evaluation is made based on the…

计算机视觉与模式识别 · 计算机科学 2018-11-27 Rob Dupre , Vasileios Argyriou

Humans have the natural ability to recognize actions even if the objects involved in the action or the background are changed. Humans can abstract away the action from the appearance of the objects which is referred to as compositionality…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Ramanathan Rajendiran , Debaditya Roy , Basura Fernando

Previous group activity recognition approaches were limited to reasoning using human relations or finding important subgroups and tended to ignore indispensable group composition and human-object interactions. This absence makes a partial…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Youliang Zhang , Zhuo Zhou , Wenxuan Liu , Danni Xu , Zheng Wang

We present a system that demonstrates how the compositional structure of events, in concert with the compositional structure of language, can interplay with the underlying focusing mechanisms in video action recognition, thereby providing a…

计算机视觉与模式识别 · 计算机科学 2014-05-29 N. Siddharth , Andrei Barbu , Jeffrey Mark Siskind

Analysis of human actions in videos demands understanding complex human dynamics, as well as the interaction between actors and context. However, these interaction relationships usually exhibit large intra-class variations from diverse…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Zhijun Zhang , Xu Zou , Jiahuan Zhou , Sheng Zhong , Ying Wu

Group activity recognition is a crucial yet challenging problem, whose core lies in fully exploring spatial-temporal interactions among individuals and generating reasonable group representations. However, previous methods either model…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Shuaicheng Li , Qianggang Cao , Lingbo Liu , Kunlin Yang , Shinan Liu , Jun Hou , Shuai Yi

Recent advancements in language models have significantly enhanced performance in multiple speech-related tasks. Existing speech language models typically utilize task-dependent prompt tokens to unify various speech tasks in a single model.…

计算与语言 · 计算机科学 2024-02-01 Yihan Wu , Soumi Maiti , Yifan Peng , Wangyou Zhang , Chenda Li , Yuyue Wang , Xihua Wang , Shinji Watanabe , Ruihua Song

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

Human-object interaction (HOI) detection is an important part of understanding human activities and visual scenes. The long-tailed distribution of labeled instances is a primary challenge in HOI detection, promoting research in few-shot and…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Zikun Zhuang , Ruihao Qian , Chi Xie , Shuang Liang

The rise of large-scale multimodal models has paved the pathway for groundbreaking advances in generative modeling and reasoning, unlocking transformative applications in a variety of complex tasks. However, a pressing question that remains…

计算与语言 · 计算机科学 2024-04-19 Semih Yagcioglu , Osman Batur İnce , Aykut Erdem , Erkut Erdem , Desmond Elliott , Deniz Yuret

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

This paper presents ReasonFormer, a unified reasoning framework for mirroring the modular and compositional reasoning process of humans in complex decision-making. Inspired by dual-process theory in cognitive science, the representation…

计算与语言 · 计算机科学 2022-12-08 Wanjun Zhong , Tingting Ma , Jiahai Wang , Jian Yin , Tiejun Zhao , Chin-Yew Lin , Nan Duan
‹ 上一页 1 2 3 10 下一页 ›