English
Related papers

Related papers: Stories From the Past Web

200 papers

Designing keyword-based access paths is a common practice in digital libraries. They are easy to use and accepted by users and come with moderate costs for content providers. However, users usually have to break down the search into pieces…

Digital Libraries · Computer Science 2022-08-23 Hermann Kroll , Niklas Mainzer , Wolf-Tilo Balke

Networks have become the de facto diagram of the Big Data age (try searching Google Images for [big data AND visualisation] and see). The concept of networks has become central to many fields of human inquiry and is said to revolutionise…

Social and Information Networks · Computer Science 2018-06-22 Liliana Bounegru , Tommaso Venturini , Jonathan Gray , Mathieu Jacomy

Web Data Extraction is an important problem that has been studied by means of different scientific tools and in a broad range of applications. Many approaches to extracting data from the Web have been designed to solve specific problems and…

Information Retrieval · Computer Science 2017-03-07 Emilio Ferrara , Pasquale De Meo , Giacomo Fiumara , Robert Baumgartner

A growing body of work shows that many problems in fairness, accountability, transparency, and ethics in machine learning systems are rooted in decisions surrounding the data collection and annotation process. In spite of its fundamental…

Machine Learning · Computer Science 2019-12-24 Eun Seo Jo , Timnit Gebru

Although the Internet Archive's Wayback Machine is the largest and most well-known web archive, there have been a number of public web archives that have emerged in the last several years. With varying resources, audiences and collection…

Digital Libraries · Computer Science 2013-01-08 Scott G. Ainsworth , Ahmed AlSum , Hany SalahEldeen , Michele C. Weigle , Michael L. Nelson

Providing effective access paths to content is a key task in digital libraries. Oftentimes, such access paths are realized through advanced query languages, which, on the one hand, users may find challenging to learn or use, and on the…

Digital Libraries · Computer Science 2023-05-01 Hermann Kroll , Christin Katharina Kreutz , Pascal Sackhoff , Wolf-Tilo Balke

New ways of documenting and describing language via electronic media coupled with new ways of distributing the results via the World-Wide Web offer a degree of access to language resources that is unparalleled in history. At the same time,…

Computation and Language · Computer Science 2007-05-23 Gary Simons , Steven Bird

The progressive digitization of historical archives provides new, often domain specific, textual resources that report on facts and events which have happened in the past; among these, memoirs are a very common type of primary source. In…

Computation and Language · Computer Science 2021-02-25 Marco Rovera , Federico Nanni , Simone Paolo Ponzetto

Retrieving information from social networks is the first and primordial step many data analysis fields such as Natural Language Processing, Sentiment Analysis and Machine Learning. Important data science tasks relay on historical data…

Information Retrieval · Computer Science 2018-03-28 A. Hernandez-Suarez , G. Sanchez-Perez , K. Toscano-Medina , V. Martinez-Hernandez , V. Sanchez , H. Perez-Meana

Web usage mining is a process of extracting useful information from server logs i.e. users history. Web usage mining is a process of finding out what users are looking for on the internet. Some users might be looking at only textual data,…

Information Retrieval · Computer Science 2013-10-25 P YesuRaju , P KiranSree

Nowadays, the Internet is indispensable when it comes to information dissemination. People rely on the Internet to inform themselves on current news events, as well as to verify facts. We, as a community, are quickly approaching an…

Digital Libraries · Computer Science 2018-09-18 Waqar Detho

We explore story generation: creative systems that can build coherent and fluent passages of text about a topic. We collect a large dataset of 300K human-written stories paired with writing prompts from an online forum. Our dataset enables…

Computation and Language · Computer Science 2018-05-15 Angela Fan , Mike Lewis , Yann Dauphin

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

We document strategies and lessons learned from sampling the web by collecting 27.3 million URLs with 3.8 billion archived pages spanning 26 years (1996-2021) from the Internet Archive's (IA) Wayback Machine. Our goal is to revisit…

Digital Libraries · Computer Science 2025-07-22 Kritika Garg , Sawood Alam , Dietrich Ayala , Mark Graham , Michele C. Weigle , Michael L. Nelson

The Archives Unleashed project aims to improve scholarly access to web archives through a multi-pronged strategy involving tool creation, process modeling, and community building - all proceeding concurrently in mutually-reinforcing…

Digital Libraries · Computer Science 2020-01-16 Nick Ruest , Jimmy Lin , Ian Milligan , Samantha Fritz

The proliferation of social media has the potential for changing the structure and organization of the web. In the past, scientists have looked at the web as a large connected component to understand how the topology of hyperlinks…

Social and Information Networks · Computer Science 2013-08-27 Tommy Nguyen , Boleslaw K. Szymanski

Archiving the web is socially and culturally critical, but presents problems of scale. The Internet Archive's Wayback Machine can replay captured web pages as they existed at a certain point in time, but it has limited ability to provide…

Information Retrieval · Computer Science 2013-06-12 Ahmed AlSum , Michael L. Nelson

The social Web is transforming the way information is created and distributed. Blog authoring tools enable users to publish content, while sites such as Digg and Del.icio.us are used to distribute content to a wider audience. With content…

Computers and Society · Computer Science 2008-12-18 Kristina Lerman , Aram Galstyan

The arXiv is the most popular preprint repository in the world. Since its inception in 1991, the arXiv has allowed researchers to freely share publication-ready articles prior to formal peer review. The growth and the popularity of the…

Digital Libraries · Computer Science 2017-09-22 Alberto Pepe , Matteo Cantiello , Josh Nicholson

Retrieval and content management are assumed to be mutually exclusive. In this paper we suggest that they need not be so. In the usual information retrieval scenario, some information about queries leading to a website (due to `hits' or…

Information Retrieval · Computer Science 2019-08-29 C Ravindranath Chowdary , Anil Kumar Singh , Anil Nelakanti
‹ Prev 1 3 4 5 6 7 10 Next ›