中文
相关论文

相关论文: Unsupervised Ego- and Exo-centric Dense Procedural…

200 篇论文

Action recognition is currently one of the top-challenging research fields in computer vision. Convolutional Neural Networks (CNNs) have significantly boosted its performance but rely on fixed-size spatio-temporal windows of analysis,…

计算机视觉与模式识别 · 计算机科学 2020-08-27 Alejandro López-Cifuentes , Marcos Escudero-Viñolo , Jesús Bescós

Emotion Recognition in Conversation (ERC) has attracted widespread attention in the natural language processing field due to its enormous potential for practical applications. Existing ERC methods face challenges in achieving generalization…

计算与语言 · 计算机科学 2023-09-20 Shanglin Lei , Xiaoping Wang , Guanting Dong , Jiang Li , Yingjian Liu

Current facial emotion recognition systems are predominately trained to predict a fixed set of predefined categories or abstract dimensional values. This constrained form of supervision hinders generalization and applicability, as it…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Licai Sun , Xingxun Jiang , Haoyu Chen , Yante Li , Zheng Lian , Biu Liu , Yuan Zong , Wenming Zheng , Jukka M. Leppänen , Guoying Zhao

While exploring visual scenes, humans' scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Xianyu Chen , Ming Jiang , Qi Zhao

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

计算与语言 · 计算机科学 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

Human language processing relies on the brain's capacity for predictive inference. We present a machine learning framework for decoding neural (EEG) responses to dynamic visual language stimuli in Deaf signers. Using coherence between…

神经元与认知 · 定量生物学 2025-12-25 Sean C. Borneman , Julia Krebs , Ronnie B. Wilbur , Evie A. Malaia

In this paper, we study the problem of producing a comprehensive video summary following an unsupervised approach that relies on adversarial learning. We build on a popular method where a Generative Adversarial Network (GAN) is trained to…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Maria Nektaria Minaidi , Charilaos Papaioannou , Alexandros Potamianos

Understanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person's emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Matthias Jammot , Björn Braun , Paul Streli , Rafael Wampfler , Christian Holz

Significant inter-individual variability limits the generalization of EEG-based emotion recognition under cross-domain settings. We address two core challenges in multi-source adaptation: (1) dynamically modeling distributional…

机器学习 · 计算机科学 2025-10-21 Fo Hu , Can Wang , Qinxu Zheng , Xusheng Yang , Bin Zhou , Gang Li , Yu Sun , Wen-an Zhang

Speech encodes a wealth of information related to human behavior and has been used in a variety of automated behavior recognition tasks. However, extracting behavioral information from speech remains challenging including due to inadequate…

音频与语音处理 · 电气工程与系统科学 2021-04-09 Haoqi Li , Brian Baucom , Shrikanth Narayanan , Panayiotis Georgiou

Despite the remarkable performance of text-to-image diffusion models in image generation tasks, recent studies have raised the issue that generated images sometimes cannot capture the intended semantic contents of the text prompts, which…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Geon Yeong Park , Jeongsol Kim , Beomsu Kim , Sang Wan Lee , Jong Chul Ye

Learning in a multi-target environment without prior knowledge about the targets requires a large amount of samples and makes generalization difficult. To solve this problem, it is important to be able to discriminate targets through…

机器学习 · 计算机科学 2021-10-27 Kibeom Kim , Min Whoo Lee , Yoonsung Kim , Je-Hwan Ryu , Minsu Lee , Byoung-Tak Zhang

Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grained emotional modeling, existing systems remain limited to…

计算与语言 · 计算机科学 2026-04-30 Shuhao Xu , Yifan Hu , Jingjing Wu , Zhihao Du , Zheng Lian , Rui Liu

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video,…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Jilan Xu , Yifei Huang , Baoqi Pei , Junlin Hou , Qingqiu Li , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

In this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Zhaofeng Shi , Heqian Qiu , Lanxiao Wang , Fanman Meng , Qingbo Wu , Hongliang Li

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Vishakha Lall , Yisi Liu

Emotion Recognition in Conversations (ERC) aims to predict the emotional state of speakers in conversations, which is essentially a text classification task. Unlike the sentence-level text classification problem, the available supervised…

计算与语言 · 计算机科学 2020-10-07 Wenxiang Jiao , Michael R. Lyu , Irwin King

Learning an agent model that behaves like humans-capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective-is a fundamental challenge in computer vision. Existing methods…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Lu Chen , Yizhou Wang , Shixiang Tang , Qianhong Ma , Tong He , Wanli Ouyang , Xiaowei Zhou , Hujun Bao , Sida Peng

In this paper, we present an unsupervised learning approach for analyzing facial behavior based on a deep generative model combined with a convolutional neural network (CNN). We jointly train a variational auto-encoder (VAE) and a…

计算机视觉与模式识别 · 计算机科学 2018-05-14 Suman Saha , Rajitha Navarathna , Leonhard Helminger , Romann Weber

Training robust world models requires large-scale, precisely labeled multimodal datasets, a process historically bottlenecked by slow and expensive manual annotation. We present a production-tested GAZE pipeline that automates the…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Leela Krishna , Mengyang Zhao , Saicharithreddy Pasula , Harshit Rajgarhia , Abhishek Mukherji