English
Related papers

Related papers: ClueWeb22: 10 Billion Web Documents with Visual an…

200 papers

The rapid growth of scientific literature makes it challenging for researchers to identify novel and impactful ideas, especially across disciplines. Modern artificial intelligence (AI) systems offer new approaches, potentially inspiring…

Artificial Intelligence · Computer Science 2025-01-09 Xuemei Gu , Mario Krenn

While web data quality is crucial for large language models, most curation efforts focus on filtering and deduplication,treating HTML-to-text extraction as a fixed pre-processing step. Existing web corpora rely on heuristic-based extractors…

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language models (VLMs) remain limited. We introduce ChartNet, a…

English is the predominant language on the web, powering nearly half of the world's top ten million websites. Support for multilingual content is nevertheless growing, with many websites increasingly combining English with regional or…

Computation and Language · Computer Science 2025-08-27 Masudul Hasan Masud Bhuiyan , Matteo Varvello , Yasir Zaki , Cristian-Alexandru Staicu

Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when…

Information Retrieval · Computer Science 2025-02-12 Xiaochen Liu , Yanan Zhang

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

We present ACL OCL, a scholarly corpus derived from the ACL Anthology to assist Open scientific research in the Computational Linguistics domain. Integrating and enhancing the previous versions of the ACL Anthology, the ACL OCL contributes…

Computation and Language · Computer Science 2023-10-25 Shaurya Rohatgi , Yanxia Qin , Benjamin Aw , Niranjana Unnithan , Min-Yen Kan

The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduce FineVision, a meticulously collected, curated, and unified corpus of 24 million samples -…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Luis Wiedmann , Orr Zohar , Amir Mahla , Xiaohan Wang , Rui Li , Thibaud Frere , Leandro von Werra , Aritra Roy Gosthipaty , Andrés Marafioti

Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of…

Software Engineering · Computer Science 2020-03-31 Jordan Henkel , Christian Bird , Shuvendu K. Lahiri , Thomas Reps

This paper presents Lishu, a deployable web artifact for searching, monitoring, and interpreting literature from elite business and management journals. The system integrates the UTD-24 and Financial Times 50 (FT50) journal pools and…

Digital Libraries · Computer Science 2026-04-21 Chuang Zhao , Hongke Zhao , Yichen Li , Xiaoquan Zhi , Songyue Guo

Recently, neural models have been leveraged to significantly improve the performance of information extraction from semi-structured websites. However, a barrier for continued progress is the small number of datasets large enough to train…

Computation and Language · Computer Science 2023-06-16 Aidan San , Yuan Zhuang , Jan Bakus , Colin Lockard , David Ciemiewicz , Sandeep Atluri , Yangfeng Ji , Kevin Small , Heba Elfardy

Aligning coordinated text streams from multiple sources and multiple languages has opened many new research venues on cross-lingual knowledge discovery. In this paper we aim to advance state-of-the-art by: (1). extending coarse-grained…

Computation and Language · Computer Science 2016-09-28 Tao Ge , Qing Dou , Xiaoman Pan , Heng Ji , Lei Cui , Baobao Chang , Zhifang Sui , Ming Zhou

In recent years, we have witnessed the proliferation of large amounts of online content generated directly by users with virtually no form of external control, leading to the possible spread of misinformation. The search for effective…

Information Retrieval · Computer Science 2024-07-12 Rishabh Upadhyay , Gabriella Pasi , Marco Viviani

The automation of document processing is gaining recent attention due to the great potential to reduce manual work through improved methods and hardware. Neural networks have been successfully applied before - even though they have been…

Computation and Language · Computer Science 2021-06-15 Martin Holeček

Existing question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ``known unknowns'' with clear…

Computation and Language · Computer Science 2024-02-29 Corby Rosset , Ho-Lam Chung , Guanghui Qin , Ethan C. Chau , Zhuo Feng , Ahmed Awadallah , Jennifer Neville , Nikhil Rao

Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answer generation.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Kaixin zhang , Xiaohe Li , Jiahao Li , Haohua Wu , Xinyu Zhao , Zide Fan , Lei Wang

The success of deep learning has sparked interest in improving relational table tasks, like data preparation and search, with table representation models trained on large table corpora. Existing table corpora primarily contain tables…

Databases · Computer Science 2023-04-13 Madelon Hulsebos , Çağatay Demiralp , Paul Groth

We publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect…

Computation and Language · Computer Science 2017-06-13 Matthew Dunn , Levent Sagun , Mike Higgins , V. Ugur Guney , Volkan Cirik , Kyunghyun Cho

The Web is a potentially huge source of medical drug leads. But despite the significant amount of multi- dimensional information about drugs, currently commercial search engines accept only linear keyword strings as inputs. This work uses…

Information Retrieval · Computer Science 2014-04-15 Iaakov Exman

It is well-established that large, diverse datasets play a pivotal role in the performance of modern AI systems for text and image modalities. However, there are no datasets for tabular data of comparable size and diversity to those…

Computation and Language · Computer Science 2023-10-13 Gus Eggert , Kevin Huo , Mike Biven , Justin Waugh