English
Related papers

Related papers: CSMD: Curated Multimodal Dataset for Chinese Stock…

200 papers

Most existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introduce a new…

Computer Vision and Pattern Recognition · Computer Science 2020-02-14 Yipeng Sun , Jiaming Liu , Wei Liu , Junyu Han , Errui Ding , Jingtuo Liu

Nowadays, foundation models become one of fundamental infrastructures in artificial intelligence, paving ways to the general intelligence. However, the reality presents two urgent challenges: existing foundation models are dominated by the…

This paper develops and empirically evaluates a Sharpe-driven stock selection and liquidity-constrained portfolio optimization framework designed for the Chinese equity market. The proposed methodology integrates three sequential stages:…

Operating Systems · Computer Science 2025-11-18 Thanh Nguyen

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the…

Artificial Intelligence · Computer Science 2024-12-24 Priyaranjan Pattnayak , Hitesh Laxmichand Patel , Bhargava Kumar , Amit Agarwal , Ishan Banerjee , Srikant Panda , Tejaswini Kumar

Recently, many studies incorporate external knowledge into character-level feature based models to improve the performance of Chinese relation extraction. However, these methods tend to ignore the internal information of the Chinese…

Computation and Language · Computer Science 2023-03-10 Jing Yang , Bin Ji , Shasha Li , Jun Ma , Long Peng , Jie Yu

With the profound development of large language models(LLMs), their safety concerns have garnered increasing attention. However, there is a scarcity of Chinese safety benchmarks for LLMs, and the existing safety taxonomies are inadequate,…

Computation and Language · Computer Science 2024-09-04 Wenjing Zhang , Xuejiao Lei , Zhaoxiang Liu , Meijuan An , Bikun Yang , KaiKai Zhao , Kai Wang , Shiguo Lian

Accurate forecasting in financial markets requires integrating diverse data sources, from historical prices to macroeconomic indicators and financial news. However, existing models often fail to align these modalities effectively, limiting…

Machine Learning · Computer Science 2025-11-04 Yunhua Pei , John Cartlidge , Anandadeep Mandal , Daniel Gold , Enrique Marcilio , Riccardo Mazzon

Large language models (LLMs) are increasingly being applied across various specialized fields, leveraging their extensive knowledge to empower a multitude of scenarios within these domains. However, each field encompasses a variety of…

Computation and Language · Computer Science 2024-04-09 Yuhang Zhou , Zeping Li , Siyu Tian , Yuchen Ni , Sen Liu , Guangnan Ye , Hongfeng Chai

We propose MCGrad, a novel and scalable multicalibration algorithm. Multicalibration - calibration in subgroups of the data - is an important property for the performance of machine learning-based systems. Existing multicalibration methods…

Understanding the relationship between textual news and time-series evolution is a critical yet under-explored challenge in applied data science. While multimodal learning has gained traction, existing multimodal time-series datasets fall…

Computation and Language · Computer Science 2026-02-12 Jialin Chen , Aosong Feng , Ziyu Zhao , Juan Garza , Gaukhar Nurbek , Cheng Qin , Ali Maatouk , Leandros Tassiulas , Yifeng Gao , Rex Ying

Financial organizations collect a huge amount of temporal (sequential) data about clients, which is typically collected from multiple sources (modalities). Despite the urgent practical need, developing deep learning techniques suitable to…

Machine Learning · Computer Science 2025-06-03 Dzhambulat Mollaev , Alexander Kostin , Maria Postnova , Ivan Karpukhin , Ivan Kireev , Gleb Gusev , Andrey Savchenko

In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS…

Computation and Language · Computer Science 2024-05-24 Yong Hu , Fandong Meng , Jie Zhou

To predict the future movements of stock markets, numerous studies concentrate on daily data and employ various machine learning (ML) models as benchmarks that often vary and lack standardization across different research works. This paper…

Computational Finance · Quantitative Finance 2024-07-16 Han Gui

Discovering a meaningful symbolic expression that explains experimental data is a fundamental challenge in many scientific fields. We present a novel, open-source computational framework called Scientist-Machine Equation Detector (SciMED),…

Machine Learning · Computer Science 2023-03-02 Liron Simon Keren , Alex Liberzon , Teddy Lazebnik

Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, especially in multimodal contexts, remains underexplored. To…

Computation and Language · Computer Science 2026-04-27 Rui Zhao , Xuewen Zhong , Xiaoyun Zheng , Jinsong Su , Yidong Chen

This article applies natural language processing (NLP) to extract and quantify textual information to predict stock performance. Using an extensive dataset of Chinese analyst reports and employing a customized BERT deep learning model for…

Computation and Language · Computer Science 2025-03-19 Rui Liu , Jiayou Liang , Haolong Chen , Yujia Hu

Multimodal learning, which aims to understand and analyze information from multiple modalities, has achieved substantial progress in the supervised regime in recent years. However, the heavy dependence on data paired with expensive human…

Machine Learning · Computer Science 2024-08-19 Yongshuo Zong , Oisin Mac Aodha , Timothy Hospedales

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about a…

Computation and Language · Computer Science 2024-10-15 Shengxiang Gao , Fang nan , Yongbing Zhang , Yuxin Huang , Kaiwen Tan , Zhengtao Yu

Financial market predictions utilize historical data to anticipate future stock prices and market trends. Traditionally, these predictions have focused on the statistical analysis of quantitative factors, such as stock prices, trading…

Statistical Finance · Quantitative Finance 2024-02-13 Zihan Dong , Xinyu Fan , Zhiyuan Peng

Chinese Spelling Correction (CSC) is a critical task in natural language processing, aimed at detecting and correcting spelling errors in Chinese text. This survey provides a comprehensive overview of CSC, tracing its evolution from…

Computation and Language · Computer Science 2025-02-18 Changchun Liu , Kai Zhang , Junzhe Jiang , Zixiao Kong , Qi Liu , Enhong Chen