中文
相关论文

相关论文: PoCaP Corpus: A Multimodal Dataset for Smart Opera…

200 篇论文

We introduce EchoXFlow, a clinical echocardiography dataset for learning from ultrasound in its native acquisition geometry rather than from scan-converted Cartesian videos. Existing public datasets offer limited opportunities to study…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Elias Stenhede , Joanna Sulkowska , Eivind Bjørkan Orstad , Henrik Schirmer , Arian Ranjbar

Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcity of emotion…

计算与语言 · 计算机科学 2023-08-30 Théo Deschamps-Berger , Lori Lamel , Laurence Devillers

Minimally invasive colorectal surgery is characterized by procedural variability, a difficult learning curve, and complications that impact quality and outcomes. Video-based assessment (VBA) offers an opportunity to generate data-driven…

Objectives: Analyze the types of studies and algorithms that are most applied, Identify the anatomical regions treated. Determine the application of parallel techniques used in studies carried out between 2010 and 2022 in research on noise…

图像与视频处理 · 电气工程与系统科学 2023-01-05 Sussana M. Florez-Aroni , Mijail A. Hancco-Condori , Fred Torres-Cruz

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

计算机视觉与模式识别 · 计算机科学 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

This paper briefly reports our ongoing attempt at the development of a multi-platform browser-based speech recording system. We designed the system toward a service of providing open service of building large-scale speech corpora at a…

人机交互 · 计算机科学 2019-12-20 Keita Ishizuka , Takashi Nose

Autonomy in robot-assisted minimally invasive surgery has the potential to reduce surgeon cognitive and task load, thereby increasing procedural efficiency. However, implementing accurate autonomous control can be difficult due to poor…

机器人学 · 计算机科学 2026-03-18 Shuyuan Yang , Zonghe Chua

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

计算与语言 · 计算机科学 2018-04-11 Pete Warden

In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations ($14,000$ hr), designed the task to represent the real word…

In order to simulate human language capacity, natural language processing systems must be able to reason about the dynamics of everyday situations, including their possible causes and effects. Moreover, they should be able to generalise the…

计算与语言 · 计算机科学 2020-10-28 Edoardo Maria Ponti , Goran Glavaš , Olga Majewska , Qianchu Liu , Ivan Vulić , Anna Korhonen

We present TiPToP, an extensible modular system that combines pretrained vision foundation models with an existing Task and Motion Planner (TAMP) to solve multi-step manipulation tasks directly from input RGB images and natural-language…

Computer-assisted multimodal training is an effective way of learning complex motor skills in various applications. In particular disciplines (eg. healthcare) incompetency in performing dexterous hands-on examinations (clinical palpation)…

人机交互 · 计算机科学 2020-01-17 A. Asadipour , K. Debattista , V. Patel , A. Chalmers

Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural…

计算与语言 · 计算机科学 2025-08-26 Bastien Le Guellec , Kokou Adambounou , Lisa C Adams , Thibault Agripnidis , Sung Soo Ahn , Radhia Ait Chalal , Tugba Akinci D Antonoli , Philippe Amouyel , Henrik Andersson , Raphael Bentegeac , Claudio Benzoni , Antonino Andrea Blandino , Felix Busch , Elif Can , Riccardo Cau , Armando Ugo Cavallo , Christelle Chavihot , Erwin Chiquete , Renato Cuocolo , Eugen Divjak , Gordana Ivanac , Barbara Dziadkowiec Macek , Armel Elogne , Salvatore Claudio Fanni , Carlos Ferrarotti , Claudia Fossataro , Federica Fossataro , Katarzyna Fulek , Michal Fulek , Pawel Gac , Martyna Gachowska , Ignacio Garcia Juarez , Marco Gatti , Natalia Gorelik , Alexia Maria Goulianou , Aghiles Hamroun , Nicolas Herinirina , Krzysztof Kraik , Dominik Krupka , Quentin Holay , Felipe Kitamura , Michail E Klontzas , Anna Kompanowska , Rafal Kompanowski , Alexandre Lefevre , Tristan Lemke , Maximilian Lindholz , Lukas Muller , Piotr Macek , Marcus Makowski , Luigi Mannacio , Aymen Meddeb , Antonio Natale , Beatrice Nguema Edzang , Adriana Ojeda , Yae Won Park , Federica Piccione , Andrea Ponsiglione , Malgorzata Poreba , Rafal Poreba , Philipp Prucker , Jean Pierre Pruvo , Rosa Alba Pugliesi , Feno Hasina Rabemanorintsoa , Vasileios Rafailidis , Katarzyna Resler , Jan Rotkegel , Luca Saba , Ezann Siebert , Arnaldo Stanzione , Ali Fuat Tekin , Liz Toapanta Yanchapaxi , Matthaios Triantafyllou , Ekaterini Tsaoulia , Evangelia Vassalou , Federica Vernuccio , Johan Wasselius , Weilang Wang , Szymon Urban , Adrian Wlodarczak , Szymon Wlodarczak , Andrzej Wysocki , Lina Xu , Tomasz Zatonski , Shuhang Zhang , Sebastian Ziegelmayer , Gregory Kuchcinski , Keno K Bressem

This study examined the use of voice recognition technology in perioperative services (Periop) to enable Periop staff to record workflow milestones using mobile technology. The use of mobile technology to improve patient flow and quality of…

音频与语音处理 · 电气工程与系统科学 2024-02-07 Majbah Uddin , Nathan Huynh , Jose M Vidal , Kevin M Taaffe , Lawrence D Fredendall , Joel S Greenstein

This paper introduces a novel Russian speech dataset called Golos, a large corpus suitable for speech research. The dataset mainly consists of recorded audio files manually annotated on the crowd-sourcing platform. The total duration of the…

音频与语音处理 · 电气工程与系统科学 2021-06-21 Nikolay Karpov , Alexander Denisenko , Fedor Minkin

We explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a…

This work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic information for a rich set of visual concepts with over…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Shuhuai Ren , Aston Zhang , Yi Zhu , Shuai Zhang , Shuai Zheng , Mu Li , Alex Smola , Xu Sun

Computer-Assisted Interventions enable clinicians to perform precise, minimally invasive procedures, often relying on advanced imaging methods. Cone-beam computed tomography (CBCT) can be used to facilitate computer-assisted interventions,…

图像与视频处理 · 电气工程与系统科学 2024-12-04 Maximilian E. Tschuchnig , Philipp Steininger , Michael Gadermayr

We introduce the Situated Corpus Of Understanding Transactions (SCOUT), a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration. The corpus was constructed from multiple Wizard-of-Oz experiments…

Phonocardiography has recently gained popularity in low-cost and remote monitoring, including passive fetal heart monitoring. Development for methods which analyse phonocardiographical data try to capitalize on this opportunity, and in…

信号处理 · 电气工程与系统科学 2024-08-26 Kristóf Müller , Janka Hatvani , Miklós Koller , Márton Áron Goda