English
Related papers

Related papers: PoCaP Corpus: A Multimodal Dataset for Smart Opera…

200 papers

This paper introduces a new ROSbag-based multimodal affective dataset for emotional and cognitive states generated using Robot Operating System (ROS). We utilized images and sounds from the International Affective Pictures System (IAPS) and…

Computers and Society · Computer Science 2020-10-21 Wonse Jo , Shyam Sundar Kannan , Go-Eum Cha , Ahreum Lee , Byung-Cheol Min

The sophisticated sense of touch of the human hand significantly contributes to our ability to safely, efficiently, and dexterously manipulate arbitrary objects in our environment. Robotic and prosthetic devices lack refined, tactile…

Robotics · Computer Science 2021-08-02 Xiaying Wang , Fabian Geiger , Vlad Niculescu , Michele Magno , Luca Benini

The identification of intentionally delivered commands is a challenge in Brain Computer Interfaces (BCIs) based on Sensory-Motor Rhythms (SMR). It is of fundamental importance that BCI systems controlling a robotic device (i.e., upper limb…

Human-Computer Interaction · Computer Science 2019-05-27 Tortora Stefano , Beraldo Gloria , Tonin Luca , Menegatti Emanuele

Bimodal stimulation, combining cochlear implant (CI) and acoustic input from the opposite ear, typically enhances speech perception but varies due to factors like temporal mismatch. Previously, we used cortical auditory evoked potentials…

Neurons and Cognition · Quantitative Biology 2025-01-29 Hanna Dolhopiatenko , Waldo Nogueira

While existing critical care EHR datasets such as MIMIC and eICU have enabled significant advances in clinical AI research, the CRITICAL dataset opens new frontiers by providing extensive scale and diversity -- containing 1.95 billion…

Machine Learning · Computer Science 2025-09-24 Xiaolong Luo , Michael Lingzhi Li

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and…

The performance of speech and events recognition systems significantly improved recently thanks to deep learning methods. However, some of these tasks remain challenging when algorithms are deployed on robots due to the unseen mechanical…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-08 Pierre-Olivier Lagacé , François Ferland , François Grondin

We introduce RadioTalk, a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the fields of natural…

Computation and Language · Computer Science 2019-09-18 Doug Beeferman , William Brannon , Deb Roy

For those experiencing severe-to-profound sensorineural hearing loss, the cochlear implant (CI) is the preferred treatment. Augmented reality (AR) aided surgery can potentially improve CI procedures and hearing outcomes. Typically, AR…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Yike Zhang , Eduardo Davalos , Dingjie Su , Ange Lou , Jack H. Noble

Substantial advances in multi-modal Artificial Intelligence (AI) facilitate the combination of diverse medical modalities to achieve holistic health assessments. We present COMPRER , a novel multi-modal, multi-objective pretraining…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Guy Lutsker , Hagai Rossman , Nastya Godiva , Eran Segal

Today, data collection has improved in various areas, and the medical domain is no exception. Auscultation, as an important diagnostic technique for physicians, due to the progress and availability of digital stethoscopes, lends itself well…

Ultrasound imaging reveals eye morphology and aids in diagnosing and treating eye diseases. However, interpreting diagnostic reports requires specialized physicians. We present a labeled ophthalmic dataset for the precise analysis and the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jing Wang , Junyan Fan , Meng Zhou , Yanzhu Zhang , Mingyu Shi

As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires…

This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic…

Sound · Computer Science 2019-04-08 Heiga Zen , Viet Dang , Rob Clark , Yu Zhang , Ron J. Weiss , Ye Jia , Zhifeng Chen , Yonghui Wu

Modern knowledge workplaces increasingly strain human episodic memory as individuals navigate fragmented attention, overlapping meetings, and multimodal information streams. Existing workplace tools provide partial support through…

Human-Computer Interaction · Computer Science 2026-03-03 Lawrence Obiuwevwi , Krzysztof J. Rechowicz , Vikas Ashok , Sachin Shetty , Sampath Jayarathna

Automatic Speech Recognition (ASR) is greatly developed in recent years, which expedites many applications on other fields. For the ASR research, speech corpus is always an essential foundation, especially for the vertical industry, such as…

Computation and Language · Computer Science 2021-02-17 Bo Yang , Xianlong Tan , Zhengmao Chen , Bing Wang , Dan Li , Zhongping Yang , Xiping Wu , Yi Lin

Humans and animals are constantly exposed to a continuous stream of sensory information from different modalities. At the same time, they form more compressed representations like concepts or symbols. In species that use language, this…

Neural and Evolutionary Computing · Computer Science 2017-06-09 Karla Stepanova , Matej Hoffmann , Zdenek Straka , Frederico B. Klein , Angelo Cangelosi , Michal Vavrecka

To date, endovascular surgeries are performed using the golden standard of Fluoroscopy, which uses ionising radiation to visualise catheters and vasculature. Prolonged Fluoroscopic exposure is harmful for the patient and the clinician, and…

Image and Video Processing · Electrical Eng. & Systems 2023-09-27 Alex Ranne , Yordanka Velikova , Nassir Navab , Ferdinando Rodriguez y Baena

Clinical finding summaries from an orthopantomogram, or a dental panoramic radiograph, have significant potential to improve patient communication and speed up clinical judgments. While orthopantomogram is a first-line tool for dental…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Tzu-Ming Harry Hsu , Yin-Chih Chelsea Wang

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-01 Philipp Klumpp , Tomás Arias-Vergara , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave
‹ Prev 1 8 9 10 Next ›