中文
相关论文

相关论文: Integrating Posture Control in Speech Motor Models…

200 篇论文

By analyzing the movements of quiet standing persons by means of wavelet statistics, we observe multiple scaling regions in the underlying body dynamics. The use of the wavelet-variance function opens the possibility to relate scaling…

生物物理 · 物理学 2009-11-06 S. Thurner , C. Mittermaier , R. Hanel , K. Ehrenberger

Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities. Although unified speech-text pre-training…

计算与语言 · 计算机科学 2024-10-15 Tengfei Yu , Xuebo Liu , Zhiyi Hou , Liang Ding , Dacheng Tao , Min Zhang

Motion simulators are widely employed in basic and applied research to study the neural mechanisms of perception and action under inertial stimulations. In these studies, uncontrolled simulator-introduced noise inevitably leads to a…

系统与控制 · 计算机科学 2015-06-17 Alessandro Nesti , Karl A Beykirch , Paul R MacNeilage , Michael Barnett-Cowan , Heinrich H Bülthoff

Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously difficult to train due to the lack of linguistic knowledge. It is…

音频与语音处理 · 电气工程与系统科学 2018-11-14 Hao Tang , James Glass

Why do human languages change at some times, and not others? We address this longstanding question from a computational perspective, focusing on the case of sound change. Sound change arises from the pronunciation variability ubiquitous in…

计算与语言 · 计算机科学 2015-07-17 James Kirby , Morgan Sonderegger

This paper explores the problem of 3D human pose estimation from only low-level acoustic signals. The existing active acoustic sensing-based approach for 3D human pose estimation implicitly assumes that the target user is positioned along a…

声音 · 计算机科学 2024-11-12 Yusuke Oumi , Yuto Shibata , Go Irie , Akisato Kimura , Yoshimitsu Aoki , Mariko Isogawa

In this review, we examine computational models that explore the role of neural oscillations in speech perception, spanning from early auditory processing to higher cognitive stages. We focus on models that use rhythmic brain activities,…

神经元与认知 · 定量生物学 2025-02-19 Olesia Dogonasheva , Denis Zakharov , Anne-Lise Giraud , Boris Gutkin

Large Language Models produce sequences learned as statistical patterns from large corpora. In order not to reproduce corpus biases, after initial training models must be aligned with human values, preferencing certain continuations over…

计算与语言 · 计算机科学 2024-01-02 Tsvetelina Hristova , Liam Magee , Karen Soldatic

Imitation learning models for robotic tasks typically rely on multi-modal inputs, such as RGB images, language, and proprioceptive states. While proprioception is intuitively important for decision-making and obstacle avoidance, simply…

机器人学 · 计算机科学 2025-07-02 Fuhang Kuang , Jiacheng You , Yingdong Hu , Tong Zhang , Chuan Wen , Yang Gao

Recognition of speech, and in particular the ability to generalize and learn from small sets of labelled examples like humans do, depends on an appropriate representation of the acoustic input. We formulate the problem of finding robust…

Robots that interact with humans in a physical space or application need to think about the person's posture, which typically comes from visual sensors like cameras and infra-red. Artificial intelligence and machine learning algorithms use…

How do language models "think"? This paper formulates a probabilistic cognitive model called the bounded pragmatic speaker, which can characterize the operation of different variations of language models. Specifically, we demonstrate that…

计算与语言 · 计算机科学 2024-01-03 Khanh Nguyen

Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on…

声音 · 计算机科学 2025-09-01 Junjie Cao

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

人工智能 · 计算机科学 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the human literature,…

计算与语言 · 计算机科学 2023-11-16 James A. Michaelov , Catherine Arnett , Tyler A. Chang , Benjamin K. Bergen

While large language models (LLMs) are generally considered proficient in generating language, how similar their language usage is to that of humans remains understudied. In this paper, we test whether models exhibit linguistic convergence,…

计算与语言 · 计算机科学 2026-02-13 Terra Blevins , Susanne Schmalwieser , Benjamin Roth

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Hongwei Yi , Hualin Liang , Yifei Liu , Qiong Cao , Yandong Wen , Timo Bolkart , Dacheng Tao , Michael J. Black

Background The development of a simulation model of full body reaching tasks that can predict endeffector trajectories and joint excursions consistent with experimental data is a non-trivial task. Because of the kinematic redundancy…

系统与控制 · 计算机科学 2011-08-10 Daohang Sha , James S Thomas

We present a cerebellar architecture with two main characteristics. The first one is that complex spikes respond to increases in sensory errors. The second one is that cerebellar modules associate particular contexts where errors have…

神经元与认知 · 定量生物学 2015-03-25 Sergio Verduzco-Flores , Randall C. O'Reilly

Pointing is a key mode of interaction with robots, yet most prior work has focused on recognition rather than generation. We present a motion capture dataset of human pointing gestures covering diverse styles, handedness, and spatial…

机器人学 · 计算机科学 2025-09-17 Anna Deichler , Siyang Wang , Simon Alexanderson , Jonas Beskow