中文
相关论文

相关论文: FERA: A Pose-Based Framework for Rule-Grounded Mul…

200 篇论文

Remote and webcam-based eye tracking in multi-line reading suffers from various noise factors and layout ambiguity, precisely where real-time reading support needs reliable, per-fixation line assignment. Prior work largely addresses this…

神经元与认知 · 定量生物学 2026-05-04 Franziska Kaltenberger , Wei-Ling Chen , Enkeleda Thaqi , Enkelejda Kasneci

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Representation-alignment objectives such as VideoREPA and MoAlign…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jiesong Lian , Zixiang Zhou , Ruizhe Zhong , Yuan Zhou , Qinglin Lu , Rui Wang , Long Hu , Yixue Hao , Baoru Huang

Modern diffusion models encounter a fundamental trade-off between training efficiency and generation quality. While existing representation alignment methods, such as REPA, accelerate convergence through patch-wise alignment, they often…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Hesen Chen , Junyan Wang , Zhiyu Tan , Hao Li

Traditional sheet metal forming relies on time-consuming and expensive Finite Element Analysis (FEA) for design validation, a process that significantly prolongs design cycles. While surrogate models offer faster iteration, current…

机器学习 · 计算机科学 2026-05-20 Jiajie Luo , Mohamed Mohamed , Osama Hassan , Haosu Zhou , Yingxue Zhao , Haoran Li , Xinrun Li , Zhutao Shao , Yang Long , Nan Li , Jichun Li

Current data analysis for the Canadian Olympic fencing team is primarily done manually by coaches and analysts. Due to the highly repetitive, yet dynamic and subtle movements in fencing, manual data analysis can be inefficient and…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Kevin Zhu , Alexander Wong , John McPhee

Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately label action segments. To address this issue, we consider…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Seth Z. Zhao , Reza Ghoddoosian , Isht Dwivedi , Nakul Agarwal , Behzad Dariush

Audiology entities are using Machine Learning (ML) models to guide their screening towards people at risk. Feature Engineering (FE) focuses on optimizing data for ML models, with evolutionary methods being effective in feature selection and…

机器学习 · 计算机科学 2025-02-14 Miguel Rabuge , Nuno Lourenço

Pose estimation is essential for many applications within computer vision and robotics. Despite its uses, few works provide rigorous uncertainty quantification for poses under dense or learned models. We derive a closed-form lower bound on…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Arun Muthukkumar

Predicting the future motion of observed vehicles is a crucial enabler for safe autonomous driving. The field of motion prediction has seen large progress recently with state-of-the-Art (sotA) models achieving impressive results on…

机器人学 · 计算机科学 2024-04-19 Marcel Hallgarten , Ismail Kisa , Martin Stoll , Andreas Zell

Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated strong performance on this task, enabling simpler…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Shijing Wang , Yaping Huang , Chaoqun Cui , David Wong , Yihua Cheng , Alexandros Neophytou , Hyung Jin Chang

For soft robots to work effectively in human-centered environments, they need to be able to estimate their state and external interactions based on (proprioceptive) sensors. Estimating disturbances allows a soft robot to perform desirable…

Generating safety-critical scenarios, which are essential yet difficult to collect at scale, offers an effective method to evaluate the robustness of autonomous vehicles (AVs). Existing methods focus on optimizing adversariality while…

机器人学 · 计算机科学 2024-10-14 Keyu Chen , Yuheng Lei , Hao Cheng , Haoran Wu , Wenchao Sun , Sifa Zheng

Algorithmic decision making driven by neural networks has become very prominent in applications that directly affect people's quality of life. In this paper, we study the problem of verifying, training, and guaranteeing individual fairness…

机器学习 · 计算机科学 2023-01-31 Kiarash Mohammadi , Aishwarya Sivaraman , Golnoosh Farnadi

Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Wei Suo , Mengyang Sun , Peng Wang , Qi Wu

In the realm of parameter-efficient fine-tuning (PEFT) methods, while options like LoRA are available, there is a persistent demand in the industry for a PEFT approach that excels in both efficiency and performance within the context of…

计算与语言 · 计算机科学 2025-02-04 Zequan Liu , Yi Zhao , Ming Tan , Wei Zhu , Aaron Xuxiang Tian

This study introduces the concept of finite element network analysis (FENA) which is a physics-informed, machine-learning-based, computational framework for the simulation of complex physical systems. The framework leverages the extreme…

计算物理 · 物理学 2021-02-24 Mehdi Jokar , Fabio Semperlotti

Human understanding of video dynamics relies on forming structured representations of entities, actions, and temporal relations before engaging in abstract reasoning. In contrast, existing Video-LLMs apply unstructured chain-of-thought…

计算与语言 · 计算机科学 2026-05-08 Zinuo Li , Yongxin Guo , Jun Liu , Jiawei Zhan , Xi Jiang , Chengjie Wang , Mohammed Bennamoun , Farid Boussaid , Feng Zheng , Qiuhong Ke

Inspired by human vision, we propose a new periphery-fovea multi-resolution driving model that predicts vehicle speed from dash camera videos. The peripheral vision module of the model processes the full video frames in low resolution. Its…

计算机视觉与模式识别 · 计算机科学 2019-03-26 Ye Xia , Jinkyu Kim , John Canny , Karl Zipser , David Whitney

The problem with existing camera-based Deep Reinforcement Learning approaches is twofold: they rarely integrate high-level scene context into the feature representation, and they rely on rigid, fixed reward functions. To address these…

机器人学 · 计算机科学 2026-02-06 Vinal Asodia , Iman Sharifi , Saber Fallah

Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrates three subtasks: emotion recognition, facial Action Unit (AU) recognition, and AU-based…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Jiulong Wu , Yucheng Shen , Lingyong Yan , Haixin Sun , Deguo Xia , Jizhou Huang , Min Cao