English
Related papers

Related papers: The Teacher-Student Chatroom Corpus

200 papers

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including about 68 hours of…

Computation and Language · Computer Science 2021-04-28 Ruiqing Zhang , Xiyang Wang , Chuanqiang Zhang , Zhongjun He , Hua Wu , Zhi Li , Haifeng Wang , Ying Chen , Qinfei Li

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

Classroom discourse is a core medium of instruction - analyzing it can provide a window into teaching and learning as well as driving the development of new tools for improving instruction. We introduce the largest dataset of mathematics…

Computation and Language · Computer Science 2023-05-26 Dorottya Demszky , Heather Hill

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge…

Computation and Language · Computer Science 2020-03-25 Meng Chen , Ruixue Liu , Lei Shen , Shaozu Yuan , Jingyan Zhou , Youzheng Wu , Xiaodong He , Bowen Zhou

This paper describes "TLT-school" a corpus of speech utterances collected in schools of northern Italy for assessing the performance of students learning both English and German. The corpus was recorded in the years 2017 and 2018 from…

Computation and Language · Computer Science 2020-01-23 Roberto Gretter , Marco Matassoni , Stefano Bannò , Daniele Falavigna

Tandem telecollaboration is a pedagogy used in second language learning where mixed groups of students meet online in videoconferencing sessions to practice their conversational skills in their target language. We have built and deployed a…

Human-Computer Interaction · Computer Science 2022-06-14 Aparajita Dey-Plissonneau , Hyowon Lee , Mingming Liu , Vyoma Patel , Michael Scriney , Alan F. Smeaton

Transitioning between topics is a natural component of human-human dialog. Although topic transition has been studied in dialogue for decades, only a handful of corpora based studies have been performed to investigate the subtleties of…

Computation and Language · Computer Science 2022-07-21 Mayank Soni , Brendan Spillane , Emer Gilmartin , Christian Saam , Benjamin R. Cowan , Vincent Wade

This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to…

Computation and Language · Computer Science 2021-06-17 Rohola Zandie , Mohammad H. Mahoor , Julia Madsen , Eshrat S. Emamian

We propose Teacher-Student Curriculum Learning (TSCL), a framework for automatic curriculum learning, where the Student tries to learn a complex task and the Teacher automatically chooses subtasks from a given set for the Student to train…

Machine Learning · Computer Science 2017-12-01 Tambet Matiisen , Avital Oliver , Taco Cohen , John Schulman

In Computer-Supported learning, monitoring and engaging a group of learners is a complex task for teachers, especially when learners are working collaboratively: Are my students motivated? What kind of progress are they making? Should I…

Computers and Society · Computer Science 2016-05-25 Eliana Scheihing , Matthieu Vernier , Javiera Born , Julio Guerra , Luis Carcamo

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

Computation and Language · Computer Science 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

Computation and Language · Computer Science 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

The scope of this paper was to find out how the students in Computer Science perceive different teaching styles and how the teaching style impacts the learning desire and interest in the course. To find out, we designed and implemented an…

Human-Computer Interaction · Computer Science 2023-07-11 Manuela Petrescu , Kuderna Bentasup

Multi-turn dialogues between a child and a caregiver are characterized by a property called contingency - that is, prompt, direct, and meaningful exchanges between interlocutors. We introduce ContingentChat, a teacher-student framework that…

Computation and Language · Computer Science 2025-10-24 Suchir Salhan , Hongyi Gu , Donya Rooein , Diana Galvan-Sosa , Gabrielle Gaudeau , Andrew Caines , Zheng Yuan , Paula Buttery

Evaluating teachers' skills is crucial for enhancing education quality and student outcomes. Teacher discourse, significantly influencing student performance, is a key component. However, coding this discourse can be laborious. This study…

Computation and Language · Computer Science 2024-12-20 Samuel Falcon , Jaime Leon

Modern text-to-speech (TTS) systems use deep learning to synthesize speech increasingly approaching human quality, but they require a database of high quality audio-text sentence pairs for training. Malayalam, the official language of the…

Sound · Computer Science 2022-11-24 Deepa P Gopinath , Thennal D K , Vrinda V Nair , Swaraj K S , Sachin G

This paper introduces a new multi-speaker English dataset for training text-to-speech models. The dataset is based on LibriVox audiobooks and Project Gutenberg texts, both in the public domain. The new dataset contains about 292 hours of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg , Yang Zhang

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

Computation and Language · Computer Science 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However,…

Computation and Language · Computer Science 2026-04-03 Wataru Nakata , Kentaro Seki , Hitomi Yanaka , Yuki Saito , Shinnosuke Takamichi , Hiroshi Saruwatari

We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large…

Computation and Language · Computer Science 2026-04-10 Matteo Rinaldi , Rossella Varvara , Viviana Patti