中文
相关论文

相关论文: PoCaP Corpus: A Multimodal Dataset for Smart Opera…

200 篇论文

The use of autonomous robots for assistance tasks in hospitals has the potential to free up qualified staff and im-prove patient care. However, the ubiquity of deformable and transparent objects in hospital settings poses signif-icant…

机器人学 · 计算机科学 2023-12-22 Benjamin Alt , Minh Dang Nguyen , Andreas Hermann , Darko Katic , Rainer Jäkel , Rüdiger Dillmann , Eric Sax

Corpus phonetics has become an increasingly popular method of research in linguistic analysis. With advances in speech technology and computational power, large scale processing of speech data has become a viable technique. This tutorial…

计算与语言 · 计算机科学 2018-11-15 Eleanor Chodroff

When technical requirements are high, and patient outcomes are critical, opportunities for monitoring and improving surgical skills via objective motion analysis feedback may be particularly beneficial. This narrative review synthesises…

计算机视觉与模式识别 · 计算机科学 2024-02-20 Merryn D. Constable , Hubert P. H. Shum , Stephen Clark

Voice assistants have become an essential tool for people with various disabilities because they enable complex phone- or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these…

音频与语音处理 · 电气工程与系统科学 2022-02-17 Colin Lea , Zifang Huang , Dhruv Jain , Lauren Tooley , Zeinab Liaghat , Shrinath Thelapurath , Leah Findlater , Jeffrey P. Bigham

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

声音 · 计算机科学 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Autonomous surgical procedures, in particular minimal invasive surgeries, are the next frontier for Artificial Intelligence research. However, the existing challenges include precise identification of the human anatomy and the surgical…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Salman Maqbool , Aqsa Riaz , Hasan Sajid , Osman Hasan

Point-of-Care Ultrasound (POCUS) refers to clinician-performed and interpreted ultrasonography at the patient's bedside. Interpreting these images requires a high level of expertise, which may not be available during emergencies. In this…

Deep learning enables the development of efficient end-to-end speech processing applications while bypassing the need for expert linguistic and signal processing features. Yet, recent studies show that good quality speech resources and…

音频与语音处理 · 电气工程与系统科学 2020-09-16 Adriana Stan

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour…

计算与语言 · 计算机科学 2019-08-12 Michael Picheny , Zóltan Tüske , Brian Kingsbury , Kartik Audhkhasi , Xiaodong Cui , George Saon

In this manuscript, the topic of multi-corpus Speech Emotion Recognition (SER) is approached from a deep transfer learning perspective. A large corpus of emotional speech data, EmoSet, is assembled from a number of existing SER corpora. In…

声音 · 计算机科学 2021-03-16 Maurice Gerczuk , Shahin Amiriparian , Sandra Ottl , Björn Schuller

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

We developed a voice-driven artificial intelligence (AI) system that guides anyone - from paramedics to family members - through expert-level stroke evaluations using natural conversation, while also enabling smartphone video capture of key…

Automatic speech recognition (ASR) via call is essential for various applications, including AI for contact center (AICC) services. Despite the advancement of ASR, however, most publicly available call-based speech corpora such as…

Ultrasound scanning is a critical imaging technique for real-time, non-invasive diagnostics. However, variations in patient anatomy and complex human-in-the-loop interactions pose significant challenges for autonomous robotic scanning.…

机器人学 · 计算机科学 2025-11-21 Ruoqu Chen , Xiangjie Yan , Kangchen Lv , Gao Huang , Zheng Li , Xiang Li

Photoacoustic Microscopy (PAM) images integrating the advantages of optical contrast and acoustic resolution have been widely used in brain studies. However, there exists a trade-off between scanning speed and image resolution. Compared…

图像与视频处理 · 电气工程与系统科学 2025-03-06 Kai Pan , Linyang Li , Li Lin , Pujin Cheng , Junyan Lyu , Lei Xi , Xiaoyin Tang

The automatic analysis of the surgical process, from videos recorded during surgeries, could be very useful to surgeons, both for training and for acquiring new techniques. The training process could be optimized by automatically providing…

计算机视觉与模式识别 · 计算机科学 2016-10-19 Katia Charrière , Gwenolé Quellec , Mathieu Lamard , David Martiano , Guy Cazuguel , Gouenou Coatrieux , Béatrice Cochener

This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice transcribers produced…

计算与语言 · 计算机科学 2026-03-18 Martina Simonotti , Ludovica Pannitto , Eleonora Zucchini , Silvia Ballarè , Caterina Mauri

Early identification of stroke symptoms is essential for enabling timely intervention and improving patient outcomes, particularly in prehospital settings. This study presents a fast, non-invasive multimodal deep learning framework for…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Ngoc-Khai Hoang , Thi-Nhu-Mai Nguyen , Huy-Hieu Pham

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Vibravox is a dataset compliant with the General Data Protection Regulation (GDPR) containing audio recordings using five different body-conduction audio sensors: two in-ear microphones, two bone conduction vibration pickups, and a…

音频与语音处理 · 电气工程与系统科学 2025-03-28 Julien Hauret , Malo Olivier , Thomas Joubaud , Christophe Langrenne , Sarah Poirée , Véronique Zimpfer , Éric Bavu