English
Related papers

Related papers: Fundus: A Simple-to-Use News Scraper Optimized for…

200 papers

Despite the advancements in search engine features, ranking methods, technologies, and the availability of programmable APIs, current-day open-access digital libraries still rely on crawl-based approaches for acquiring their underlying…

Information Retrieval · Computer Science 2016-04-19 Sujatha Das Gollapalli , Krutarth Patel , Cornelia Caragea

Crawling parallel texts -- texts that are mutual translations -- from the Internet is usually done following a brute-force approach: documents are massively downloaded in an unguided process, and only a fraction of them end up leading to…

Computation and Language · Computer Science 2026-04-22 Cristian García-Romero , Miquel Esplà-Gomis , Felipe Sánchez-Martínez

In the last decade we have observed a mass increase of information, in particular information that is shared through smartphones. Consequently, the amount of information that is available does not allow the average user to be aware of all…

Information Retrieval · Computer Science 2017-07-04 Akshay Kumar Chaturvedi , Filipa Peleja , Ana Freire

Wikipedia is the largest web repository of free knowledge. Volunteer editors devote time and effort to creating and expanding articles in more than 300 language editions. As content quality varies from article to article, editors also spend…

Computers and Society · Computer Science 2024-04-16 Paramita Das , Isaac Johnson , Diego Saez-Trumper , Pablo Aragón

With the continuous advancement of artificial intelligence, natural language processing technology has become widely utilized in various fields. At the same time, there are many challenges in creating Chinese news summaries. First of all,…

Computation and Language · Computer Science 2024-06-27 Yiming Chen , Haobin Chen , Simin Liu , Yunyun Liu , Fanhao Zhou , Bing Wei

Accessing daily news content still remains a big challenge for people with print-impairment including blind and low-vision due to opacity of printed content and hindrance from online sources. In this paper, we present our approach for…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Vishal Agarwal , Tanuja Ganu , Saikat Guha

Social media is becoming an increasingly important data source for learning about breaking news and for following the latest developments of ongoing news. This is in part possible thanks to the existence of mobile devices, which allows…

Computation and Language · Computer Science 2019-09-12 Arkaitz Zubiaga

Optimising deep learning inference across edge devices and optimisation targets such as inference time, memory footprint and power consumption is a key challenge due to the ubiquity of neural networks. Today, production deep learning…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-08-05 Perry Gibson , José Cano

We present a generic framework to make wrapper induction algorithms tolerant to noise in the training data. This enables us to learn wrappers in a completely unsupervised manner from automatically and cheaply obtained noisy training data,…

Databases · Computer Science 2011-03-15 Nilesh Dalvi , Ravi Kumar , Mohamed Soliman

Easier access to the internet and social media has made disseminating information through online sources very easy. Sources like Facebook, Twitter, online news sites and personal blogs of self-proclaimed journalists have become significant…

Computation and Language · Computer Science 2021-09-28 Shaily Bhatt , Sakshi Kalra , Naman Goenka , Yashvardhan Sharma

This research introduces ScoreRAG, an approach to enhance the quality of automated news generation. Despite advancements in Natural Language Processing and large language models, current news generation methods often struggle with…

Computation and Language · Computer Science 2025-06-05 Pei-Yun Lin , Yen-lung Tsai

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

Digital Libraries · Computer Science 2023-06-21 Muhammad Zohaib

Blog is becoming an increasingly popular media for information publishing. Besides the main content, most of blog pages nowadays also contain noisy information such as advertisements etc. Removing these unrelated elements can improves user…

Information Retrieval · Computer Science 2017-08-29 Kui Zhao , Yi Wang , Xia Hu , Can Wang

A key challenge of online news recommendation is to help users find articles they are interested in. Traditional news recommendation methods usually use single news information, which is insufficient to encode news and user representation.…

Information Retrieval · Computer Science 2021-12-20 Songqiao Han , Hailiang Huang , Jiangwei Liu

Social media has become a popular means for people to consume news. Meanwhile, it also enables the wide dissemination of fake news, i.e., news with intentionally false information, which brings significant negative effects to the society.…

Social and Information Networks · Computer Science 2019-03-28 Kai Shu , Deepak Mahudeswaran , Suhang Wang , Dongwon Lee , Huan Liu

There is an overwhelming number of news articles published every day around the globe. Following the evolution of a news-story is a difficult task given that there is no such mechanism available to track back in time to study the diffusion…

Information Retrieval · Computer Science 2017-12-22 Roberto Camacho Barranco , Arnold P. Boedihardjo , M. Shahriar Hossain

Document extraction is an important step before retrieval-augmented generation (RAG), knowledge bases, and downstream generative AI can work. It turns unstructured documents like PDFs and scans into structured text and layout-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Aman Ulla

The traditional offline approaches are no longer sufficient for building modern recommender systems in domains such as online news services, mainly due to the high dynamics of environment changes and necessity to operate on a large scale…

Information Retrieval · Computer Science 2019-11-26 Joanna Misztal-Radecka , Dominik Rusiecki , Michał Żmuda , Artur Bujak

This paper presents a novel two-stage framework to extract opinionated sentences from a given news article. In the first stage, Naive Bayes classifier by utilizing the local features assigns a score to each sentence - the score signifies…

Computation and Language · Computer Science 2021-01-26 Rajkumar Pujari , Swara Desai , Niloy Ganguly , Pawan Goyal

The explosion in the amount of news and journalistic content being generated across the globe, coupled with extended and instantaneous access to information through online media, makes it difficult and time-consuming to monitor news…

Computation and Language · Computer Science 2018-08-06 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona