English
Related papers

Related papers: Using Weak Supervision and Data Augmentation in Qu…

200 papers

Time-series question answering (TSQA) tasks face significant challenges due to the lack of labeled data. Alternatively, with recent advancements in large-scale models, vision-language models (VLMs) have demonstrated the potential to analyze…

Machine Learning · Computer Science 2025-10-01 Takuya Fujimura , Kota Dohi , Natsuo Yamashita , Yohei Kawaguchi

The resurgence of self-supervised learning, whereby a deep learning model generates its own supervisory signal from the data, promises a scalable way to tackle the dramatically increasing size of real-world data sets without human…

Quantum Physics · Physics 2022-04-05 Ben Jaderberg , Lewis W. Anderson , Weidi Xie , Samuel Albanie , Martin Kiffner , Dieter Jaksch

We present semi-supervised models with data augmentation (SMDA), a semi-supervised text classification system to classify interactive affective responses. SMDA utilizes recent transformer-based models to encode each sentence and employs…

Computation and Language · Computer Science 2020-04-24 Jiaao Chen , Yuwei Wu , Diyi Yang

While medical image segmentation is an important task for computer aided diagnosis, the high expertise requirement for pixelwise manual annotations makes it a challenging and time consuming task. Since conventional data augmentations do not…

Image and Video Processing · Electrical Eng. & Systems 2021-06-21 Dwarikanath Mahapatra , Ankur Singh

Long-form question answering (LFQA) aims at generating in-depth answers to end-user questions, providing relevant information beyond the direct answer. However, existing retrievers are typically optimized towards information that directly…

Computation and Language · Computer Science 2024-10-14 Philipp Christmann , Svitlana Vakulenko , Ionut Teodor Sorodoc , Bill Byrne , Adrià de Gispert

Deep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions. Such reliance leads to performance degradation and poor generalization on unseen data. To…

Computation and Language · Computer Science 2025-11-21 Kyohoon Jin , Juhwan Choi , Jungmin Yun , Junho Lee , Soojin Jang , Youngbin Kim

The biggest challenge in the application of deep learning to the medical domain is the availability of training data. Data augmentation is a typical methodology used in machine learning when confronted with a limited data set. In a…

Image and Video Processing · Electrical Eng. & Systems 2024-03-21 Oleksandr Fedoruk , Konrad Klimaszewski , Aleksander Ogonowski , Rafał Możdżonek

Automatic relation extraction (RE) for types of interest is of great importance for interpreting massive text corpora in an efficient manner. Traditional RE models have heavily relied on human-annotated corpus for training, which can be…

Computation and Language · Computer Science 2017-11-27 Zeqiu Wu , Xiang Ren , Frank F. Xu , Ji Li , Jiawei Han

The reliance of text classifiers on spurious correlations can lead to poor generalization at deployment, raising concerns about their use in safety-critical domains such as healthcare. In this work, we propose to use counterfactual data…

Machine Learning · Computer Science 2024-01-10 Amir Feder , Yoav Wald , Claudia Shi , Suchi Saria , David Blei

Knowledge and language understanding of models evaluated through question answering (QA) has been usually studied on static snapshots of knowledge, like Wikipedia. However, our world is dynamic, evolves over time, and our models' knowledge…

This thesis work falls within the framework of question answering (QA) in the biomedical domain where several specific challenges are addressed, such as specialized lexicons and terminologies, the types of treated questions, and the…

Computation and Language · Computer Science 2023-07-26 Mourad Sarrouti

Being able to train Named Entity Recognition (NER) models for emerging topics is crucial for many real-world applications especially in the medical domain where new topics are continuously evolving out of the scope of existing models and…

Computation and Language · Computer Science 2022-10-11 Aleksander Ficek , Fangyu Liu , Nigel Collier

Large-scale language models like ChatGPT and GPT-4 have gained attention for their impressive conversational and generative capabilities. However, the creation of supervised paired question-answering data for instruction tuning presents…

Computation and Language · Computer Science 2023-05-23 Xuanyu Zhang , Qing Yang

In response to the need for rapid and accurate COVID-19 diagnosis during the global pandemic, we present a two-stage framework that leverages pseudo labels for domain adaptation to enhance the detection of COVID-19 from CT scans. By…

Image and Video Processing · Electrical Eng. & Systems 2024-03-19 Runtian Yuan , Qingqiu Li , Junlin Hou , Jilan Xu , Yuejie Zhang , Rui Feng , Hao Chen

Data augmentation is vital for deep learning neural networks. By providing massive training samples, it helps to improve the generalization ability of the model. Weakly supervised semantic segmentation (WSSS) is a challenging problem that…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Yukun Su , Ruizhou Sun , Guosheng Lin , Qingyao Wu

Dialogue understanding tasks often necessitate abundant annotated data to achieve good performance and that presents challenges in low-resource settings. To alleviate this barrier, we explore few-shot data augmentation for dialogue…

Computation and Language · Computer Science 2022-11-03 Maximillian Chen , Alexandros Papangelis , Chenyang Tao , Andy Rosenbaum , Seokhwan Kim , Yang Liu , Zhou Yu , Dilek Hakkani-Tur

The COVID-19 pandemic represents the most significant public health disaster since the 1918 influenza pandemic. During pandemics such as COVID-19, timely and reliable spatio-temporal forecasting of epidemic dynamics is crucial. Deep…

Machine Learning · Computer Science 2020-11-25 Lijing Wang , Aniruddha Adiga , Srinivasan Venkatramanan , Jiangzhuo Chen , Bryan Lewis , Madhav Marathe

Training deep neural networks requires many training samples, but in practice, training labels are expensive to obtain and may be of varying quality, as some may be from trusted expert labelers while others might be from heuristics or other…

Information Retrieval · Computer Science 2018-06-25 Mostafa Dehghani , Jaap Kamps

Safe and reliable natural language inference is critical for extracting insights from clinical trial reports but poses challenges due to biases in large pre-trained language models. This paper presents a novel data augmentation technique to…

Computation and Language · Computer Science 2024-04-16 Yuqi Wang , Zeqiang Wang , Wei Wang , Qi Chen , Kaizhu Huang , Anh Nguyen , Suparna De

Community-based Question Answering (CQA) sites play an important role in addressing health information needs. However, a significant number of posted questions remain unanswered. Automatically answering the posted questions can provide a…

Machine Learning · Statistics 2016-07-05 Papis Wongchaisuwat , Diego Klabjan , Siddhartha R. Jonnalagadda