中文
相关论文

相关论文: PoCaP Corpus: A Multimodal Dataset for Smart Opera…

200 篇论文

This paper introduces a new ROSbag-based multimodal affective dataset for emotional and cognitive states generated using Robot Operating System (ROS). We utilized images and sounds from the International Affective Pictures System (IAPS) and…

计算机与社会 · 计算机科学 2020-10-21 Wonse Jo , Shyam Sundar Kannan , Go-Eum Cha , Ahreum Lee , Byung-Cheol Min

The sophisticated sense of touch of the human hand significantly contributes to our ability to safely, efficiently, and dexterously manipulate arbitrary objects in our environment. Robotic and prosthetic devices lack refined, tactile…

机器人学 · 计算机科学 2021-08-02 Xiaying Wang , Fabian Geiger , Vlad Niculescu , Michele Magno , Luca Benini

The identification of intentionally delivered commands is a challenge in Brain Computer Interfaces (BCIs) based on Sensory-Motor Rhythms (SMR). It is of fundamental importance that BCI systems controlling a robotic device (i.e., upper limb…

人机交互 · 计算机科学 2019-05-27 Tortora Stefano , Beraldo Gloria , Tonin Luca , Menegatti Emanuele

Bimodal stimulation, combining cochlear implant (CI) and acoustic input from the opposite ear, typically enhances speech perception but varies due to factors like temporal mismatch. Previously, we used cortical auditory evoked potentials…

神经元与认知 · 定量生物学 2025-01-29 Hanna Dolhopiatenko , Waldo Nogueira

While existing critical care EHR datasets such as MIMIC and eICU have enabled significant advances in clinical AI research, the CRITICAL dataset opens new frontiers by providing extensive scale and diversity -- containing 1.95 billion…

机器学习 · 计算机科学 2025-09-24 Xiaolong Luo , Michael Lingzhi Li

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and…

The performance of speech and events recognition systems significantly improved recently thanks to deep learning methods. However, some of these tasks remain challenging when algorithms are deployed on robots due to the unseen mechanical…

音频与语音处理 · 电气工程与系统科学 2023-03-08 Pierre-Olivier Lagacé , François Ferland , François Grondin

We introduce RadioTalk, a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the fields of natural…

计算与语言 · 计算机科学 2019-09-18 Doug Beeferman , William Brannon , Deb Roy

For those experiencing severe-to-profound sensorineural hearing loss, the cochlear implant (CI) is the preferred treatment. Augmented reality (AR) aided surgery can potentially improve CI procedures and hearing outcomes. Typically, AR…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Yike Zhang , Eduardo Davalos , Dingjie Su , Ange Lou , Jack H. Noble

Substantial advances in multi-modal Artificial Intelligence (AI) facilitate the combination of diverse medical modalities to achieve holistic health assessments. We present COMPRER , a novel multi-modal, multi-objective pretraining…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Guy Lutsker , Hagai Rossman , Nastya Godiva , Eran Segal

Today, data collection has improved in various areas, and the medical domain is no exception. Auscultation, as an important diagnostic technique for physicians, due to the progress and availability of digital stethoscopes, lends itself well…

Ultrasound imaging reveals eye morphology and aids in diagnosing and treating eye diseases. However, interpreting diagnostic reports requires specialized physicians. We present a labeled ophthalmic dataset for the precise analysis and the…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Jing Wang , Junyan Fan , Meng Zhou , Yanzhu Zhang , Mingyu Shi

As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires…

This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic…

声音 · 计算机科学 2019-04-08 Heiga Zen , Viet Dang , Rob Clark , Yu Zhang , Ron J. Weiss , Ye Jia , Zhifeng Chen , Yonghui Wu

Modern knowledge workplaces increasingly strain human episodic memory as individuals navigate fragmented attention, overlapping meetings, and multimodal information streams. Existing workplace tools provide partial support through…

人机交互 · 计算机科学 2026-03-03 Lawrence Obiuwevwi , Krzysztof J. Rechowicz , Vikas Ashok , Sachin Shetty , Sampath Jayarathna

Automatic Speech Recognition (ASR) is greatly developed in recent years, which expedites many applications on other fields. For the ASR research, speech corpus is always an essential foundation, especially for the vertical industry, such as…

计算与语言 · 计算机科学 2021-02-17 Bo Yang , Xianlong Tan , Zhengmao Chen , Bing Wang , Dan Li , Zhongping Yang , Xiping Wu , Yi Lin

Humans and animals are constantly exposed to a continuous stream of sensory information from different modalities. At the same time, they form more compressed representations like concepts or symbols. In species that use language, this…

神经与进化计算 · 计算机科学 2017-06-09 Karla Stepanova , Matej Hoffmann , Zdenek Straka , Frederico B. Klein , Angelo Cangelosi , Michal Vavrecka

To date, endovascular surgeries are performed using the golden standard of Fluoroscopy, which uses ionising radiation to visualise catheters and vasculature. Prolonged Fluoroscopic exposure is harmful for the patient and the clinician, and…

图像与视频处理 · 电气工程与系统科学 2023-09-27 Alex Ranne , Yordanka Velikova , Nassir Navab , Ferdinando Rodriguez y Baena

Clinical finding summaries from an orthopantomogram, or a dental panoramic radiograph, have significant potential to improve patient communication and speed up clinical judgments. While orthopantomogram is a first-line tool for dental…

计算机视觉与模式识别 · 计算机科学 2021-07-07 Tzu-Ming Harry Hsu , Yin-Chih Chelsea Wang

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Philipp Klumpp , Tomás Arias-Vergara , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave
‹ 上一页 1 8 9 10 下一页 ›