English
Related papers

Related papers: Automatic Identification of Non-Meaningful Body-Mo…

200 papers

Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker's comprehension of…

Computation and Language · Computer Science 2022-01-31 Kamini Sabu , Mithilesh Vaidya , Preeti Rao

Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames can be caused by controllable changes as well as exogenous…

Machine Learning · Computer Science 2025-11-13 Chuheng Zhang , Tim Pearce , Pushi Zhang , Kaixin Wang , Xiaoyu Chen , Wei Shen , Li Zhao , Jiang Bian

We present a system that demonstrates how the compositional structure of events, in concert with the compositional structure of language, can interplay with the underlying focusing mechanisms in video action recognition, thereby providing a…

Computer Vision and Pattern Recognition · Computer Science 2014-05-29 N. Siddharth , Andrei Barbu , Jeffrey Mark Siskind

Recent work has explored the use of personal information in the form of persona sentences or self-disclosures to improve modeling of individual characteristics and prediction of annotator labels for subjective tasks. The volume of personal…

Computation and Language · Computer Science 2026-01-27 Kieran Henderson , Kian Omoomi , Vasudha Varadarajan , Allison Lahnala , Charles Welch

We propose an automatic system for organizing the content of a collection of unstructured videos of an articulated object class (e.g. tiger, horse). By exploiting the recurring motion patterns of the class across videos, our system: 1)…

Computer Vision and Pattern Recognition · Computer Science 2016-08-12 Luca Del Pero , Susanna Ricco , Rahul Sukthankar , Vittorio Ferrari

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Arjun R. Akula , Song-Chun Zhu

With recent advancements in computer vision as well as machine learning (ML), video-based at-home exercise evaluation systems have become a popular topic of current research. However, performance depends heavily on the amount of available…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Sebastian Dill , Susi Zhihan , Maurice Rohr , Maziar Sharbafi , Christoph Hoog Antink

Speaker recognition performance in emotional talking environments is not as high as it is in neutral talking environments. This work focuses on proposing, implementing, and evaluating a new approach to enhance the performance in emotional…

Sound · Computer Science 2017-06-30 Ismail Shahin

The aim of this work is to detect and automatically generate high-level explanations of anomalous events in video. Understanding the cause of an anomalous event is crucial as the required response is dependant on its nature and severity.…

Computer Vision and Pattern Recognition · Computer Science 2021-12-13 Stanislaw Szymanowicz , James Charles , Roberto Cipolla

One of the key communicative competencies is the ability to maintain fluency in monologic speech and the ability to produce sophisticated language to argue a position convincingly. In this paper we aim to predict TED talk-style affective…

Computation and Language · Computer Science 2021-12-01 Yu Qiao , Sourabh Zanwar , Rishab Bhattacharyya , Daniel Wiechmann , Wei Zhou , Elma Kerz , Ralf Schlüter

In this study we collect and annotate human-human role-play dialogues in the domain of weight management. There are two roles in the conversation: the "seeker" who is looking for ways to lose weight and the "helper" who provides suggestions…

Computation and Language · Computer Science 2018-07-12 Ramesh Manuvinakurike , Sumanth Bharadwaj , Kallirroi Georgila

What makes a public talk resonate with large audiences? While prior research has emphasized speaker delivery or topic novelty, we reasoned that a core driver of engagement is linguistic clarity. This aligns with theories of processing…

Human-Computer Interaction · Computer Science 2026-04-07 Roni Segal , Matan Lary , Ralf Schmaelzle , Yossi Ben-Zion

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional…

Sound · Computer Science 2026-01-14 Runchuan Ye , Yixuan Zhou , Renjie Yu , Zijian Lin , Kehan Li , Xiang Li , Xin Liu , Guoyang Zeng , Zhiyong Wu

Behavioral annotation using signal processing and machine learning is highly dependent on training data and manual annotations of behavioral labels. Previous studies have shown that speech information encodes significant behavioral…

Machine Learning · Computer Science 2017-01-13 Haoqi Li , Brian Baucom , Panayiotis Georgiou

In this paper, we show that different body parts do not play equally important roles in recognizing a human action in video data. We investigate to what extent a body part plays a role in recognition of different actions and hence propose a…

Computer Vision and Pattern Recognition · Computer Science 2017-05-24 Yuping Shen , Hassan Foroosh

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Human pose forecasting is the task of predicting articulated human motion given past human motion. There exists a number of popular benchmarks that evaluate an array of different models performing human pose forecasting. These benchmarks do…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Maria Priisalu , Ted Kronvall , Cristian Sminchisescu

Accurately predicting heart activity and other biological signals is crucial for diagnosis and monitoring. Given that speech is an outcome of multiple physiological systems, a significant body of work studied the acoustic correlates of…

Sound · Computer Science 2024-06-11 Gasser Elbanna , Zohreh Mostaani , Mathew Magimai. -Doss

Objective criteria for universal semantic components that distinguish a humorous utterance from a non-humorous one are presently under debate. In this article, we give an in-depth observation of our system of self-paced reading for…

Computation and Language · Computer Science 2024-07-11 Elena Mikhalkova , Nadezhda Ganzherli , Julia Murzina

We consider the task of identifying human actions visible in online videos. We focus on the widely spread genre of lifestyle vlogs, which consist of videos of people performing actions while verbally describing them. Our goal is to identify…

Computation and Language · Computer Science 2021-09-10 Oana Ignat , Laura Burdick , Jia Deng , Rada Mihalcea
‹ Prev 1 4 5 6 7 8 10 Next ›