English
Related papers

Related papers: Exploratory Analysis of a Terabyte Scale Web Corpu…

200 papers

Large Language Models (LLMs) demonstrate remarkable potential across various domains; however, they exhibit a significant performance gap in Information Extraction (IE). Note that high-quality instruction data is the vital key for enhancing…

Computation and Language · Computer Science 2024-05-28 Honghao Gui , Lin Yuan , Hongbin Ye , Ningyu Zhang , Mengshu Sun , Lei Liang , Huajun Chen

In this paper, we propose a fully automated system to extend knowledge graphs using external information from web-scale corpora. The designed system leverages a deep learning based technology for relation extraction that can be trained by a…

Computation and Language · Computer Science 2019-09-12 Sarthak Dash , Michael R. Glass , Alfio Gliozzo , Mustafa Canim

The tremendous increase in the amount of available research documents impels researchers to propose topic models to extract the latent semantic themes of a documents collection. However, how to extract the hidden topics of the documents…

Information Retrieval · Computer Science 2020-01-07 Mi Khine Oo , May Aye Khine

By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system…

Computation and Language · Computer Science 2026-05-19 Yueru Yan , Tuc Nguyen , Bo Su , Melissa Lieffers , Thai Le

Websites are capable of learning a wide range of information about the platform on which a browser is executing. One major source of such information is the set of standardised Application Programming Interfaces (APIs) provided within the…

Cryptography and Security · Computer Science 2019-10-17 Zhaoyi Fan

Many complex systems in nature and society can be described in terms of networks capturing the intricate web of connections among the units they are made of. A key question is how to interpret the global organization of such networks as the…

Physics and Society · Physics 2007-05-23 Gergely Palla , Imre Derenyi , Illes Farkas , Tamas Vicsek

With the increasing number of services in the internet, companies intranets, and home networks: service discovery becomes an integral part of modern networked system. This paper provides a comprehensive survey of major solutions for service…

Networking and Internet Architecture · Computer Science 2014-10-01 Bendaoud Karim Talal , Merzougui Rachid

Large language models (LLMs) have shown exceptional performance on a variety of natural language tasks. Yet, their capabilities for HTML understanding -- i.e., parsing the raw HTML of a webpage, with applications to automation of web-based…

While web data quality is crucial for large language models, most curation efforts focus on filtering and deduplication,treating HTML-to-text extraction as a fixed pre-processing step. Existing web corpora rely on heuristic-based extractors…

This study introduces a novel methodology for mapping scientific communities at scale, addressing challenges associated with network analysis in large bibliometric datasets. By leveraging enriched publication metadata from the French…

Digital Libraries · Computer Science 2025-01-20 Victor Barbier , Eric Jeangirard

Modeling how networks change under structural perturbations can yield foundational insights into network robustness, which is critical in many real-world applications. The largest connected component is a popular measure of network…

Physics and Society · Physics 2025-09-30 Jessica Jiang , Allison C. Zhuang , Petter Holme , Peter J. Mucha , Alice C. Schwarze

A large amount of data on the WWW remains inaccessible to crawlers of Web search engines because it can only be exposed on demand as users fill out and submit forms. The Hidden web refers to the collection of Web data which can be accessed…

Information Retrieval · Computer Science 2014-07-23 Sonali Gupta , Komal Kumar Bhatia

Modeling relations between languages can offer understanding of language characteristics and uncover similarities and differences between languages. Automated methods applied to large textual corpora can be seen as opportunities for novel…

Computation and Language · Computer Science 2019-12-24 Blaž Škrlj , Senja Pollak

Biclustering is a two way clustering approach involving simultaneous clustering along two dimensions of the data matrix. Finding biclusters of web objects (i.e. web users and web pages) is an emerging topic in the context of web usage…

Neural and Evolutionary Computing · Computer Science 2011-06-14 R. Rathipriya , Dr. K. Thangavel , J. Bagyamani

Although several datasets annotated for anaphoric reference/coreference exist, even the largest such datasets have limitations in terms of size, range of domains, coverage of anaphoric phenomena, and size of documents included. Yet, the…

Computation and Language · Computer Science 2022-10-12 Juntao Yu , Silviu Paun , Maris Camilleri , Paloma Carretero Garcia , Jon Chamberlain , Udo Kruschwitz , Massimo Poesio

Percolation processes on random networks have been the subject of intense research activity over the last decades: the overall phenomenology of standard percolation on uncorrelated and unclustered topologies is well known. Still some…

Statistical Mechanics · Physics 2024-12-06 Lorenzo Cirigliano , Gábor Timár , Claudio Castellano

Pretraining large language models (LLMs) on high-quality, structured data such as mathematics and code substantially enhances reasoning capabilities. However, existing math-focused datasets built from Common Crawl suffer from degraded…

Computation and Language · Computer Science 2025-08-22 Rabeeh Karimi Mahabadi , Sanjeev Satheesh , Shrimai Prabhumoye , Mostofa Patwary , Mohammad Shoeybi , Bryan Catanzaro

The size of web has increased exponentially over the past few years with thousands of documents related to a subject available to the user. With this much amount of information available, it is not possible to take the full advantage of the…

Information Retrieval · Computer Science 2012-11-07 R. K. Roul , S. K. Sahay

Clearly, no one likes webpages with poor quality of experience (QoE). Being perceived as slow or fast is a key element in the overall perceived QoE of web applications. While extensive effort has been put into optimizing web applications…

Networking and Internet Architecture · Computer Science 2017-04-06 Qingzhu Gao , Prasenjit Dey , Parvez Ahammad

Recent success in large multimodal models (LMMs) has sparked promising applications of agents capable of autonomously completing complex web tasks. While open-source LMM agents have made significant advances in offline evaluation…

Artificial Intelligence · Computer Science 2025-06-02 Vardaan Pahuja , Yadong Lu , Corby Rosset , Boyu Gou , Arindam Mitra , Spencer Whitehead , Yu Su , Ahmed Awadallah