中文
相关论文

相关论文: TechCoach: Towards Technical-Point-Aware Descripti…

200 篇论文

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Kumar Ashutosh , Tushar Nagarajan , Georgios Pavlakos , Kris Kitani , Kristen Grauman

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized coaching. Existing AI tools typically provide isolated…

人机交互 · 计算机科学 2025-11-25 Sirui Chen , Jinsong Zhou , Xinli Xu , Xiaoyu Yang , Litao Guo , Ying-Cong Chen

We introduce an object-aware decoder for improving the performance of spatio-temporal representations on ego-centric videos. The key idea is to enhance object-awareness during training by tasking the model to predict hand positions, object…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

Current methods for action recognition primarily rely on deep convolutional networks to derive feature embeddings of visual and motion features. While these methods have demonstrated remarkable performance on standard benchmarks, we are…

计算机视觉与模式识别 · 计算机科学 2020-05-21 Dian Shao , Yue Zhao , Bo Dai , Dahua Lin

Current AI-driven educational systems primarily rely on behavioural analytics, performance metrics, and content-level interactions to model learning. While these approaches provide useful indicators of learner activity, they are…

人机交互 · 计算机科学 2026-05-19 Annie Yuan

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

In human-AI collaboration, a central challenge is deciding whether the AI should handle a task, be deferred to a human expert, or be addressed through collaborative effort. Existing Learning to Defer approaches typically make binary choices…

人工智能 · 计算机科学 2025-05-27 Chengbo He , Bochao Zou , Junliang Xing , Jiansheng Chen , Yuanchun Shi , Huimin Ma

Detecting anomalies in human-related videos is crucial for surveillance applications. Current methods primarily include appearance-based and action-based techniques. Appearance-based methods rely on low-level visual features such as color,…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Chenglizhao Chen , Xinyu Liu , Mengke Song , Luming Li , Xu Yu , Shanchen Pang

Keypoint detection is a pivotal step in 3D reconstruction, whereby sets of (up to) K points are detected in each view of a scene. Crucially, the detected points need to be consistent between views, i.e., correspond to the same 3D point in…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Johan Edstedt , Georg Bökman , Mårten Wadenbäck , Michael Felsberg

Benchmarking autonomous driving planners to align with human judgment remains a critical challenge, as state-of-the-art metrics like the Extended Predictive Driver Model Score (EPDMS) lack context awareness in nuanced scenarios. To address…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Jingyu Song , Zhenxin Li , Shiyi Lan , Xinglong Sun , Nadine Chang , Maying Shen , Joshua Chen , Katherine A. Skinner , Jose M. Alvarez

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 David Gastager , Ghazal Ghazaei , Constantin Patsch

Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewards are primarily confined to guiding the final localization,…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Tao Yang , Qing Zhou , Yanliang Li , Qi Wang

Action Quality Assessment (AQA) is a task that tries to answer how well an action is carried out. While remarkable progress has been achieved, existing works on AQA assume that all the training data are visible for training at one time, but…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yuan-Ming Li , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng

Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods based on saliency produce entangled explanations, making it unclear…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Jongseo Lee , Wooil Lee , Gyeong-Moon Park , Seong Tae Kim , Jinwoo Choi

Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Michele Mazzamuto , Daniele Di Mauro , Gianpiero Francesca , Giovanni Maria Farinella , Antonino Furnari

Interpretation of deep learning remains a very challenging problem. Although the Class Activation Map (CAM) is widely used to interpret deep model predictions by highlighting object location, it fails to provide insight into the salient…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Yuguang Yang , Runtang Guo , Sheng Wu , Yimi Wang , Juan Zhang , Xuan Gong , Baochang Zhang

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the existence of distractors, such as viewpoint changes.…

计算机视觉与模式识别 · 计算机科学 2020-10-01 Xiangxi Shi , Xu Yang , Jiuxiang Gu , Shafiq Joty , Jianfei Cai

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed…

计算机视觉与模式识别 · 计算机科学 2021-11-08 Xuanhan Wang , Xiaojia Chen , Lianli Gao , Lechao Chen , Jingkuan Song

Image captioning has long been a pivotal task in visual understanding, with recent advancements in vision-language models (VLMs) significantly enhancing the ability to generate detailed image captions. However, the evaluation of detailed…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Qinghao Ye , Xianhan Zeng , Fu Li , Chunyuan Li , Haoqi Fan

Temporal action detection aims to locate and classify actions in untrimmed videos. While recent works focus on designing powerful feature processors for pre-trained representations, they often overlook the inherent noise and redundancy…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Xinnan Zhu , Yicheng Zhu , Tixin Chen , Wentao Wu , Yuanjie Dang