中文
相关论文

相关论文: VN-MTEB: Vietnamese Massive Text Embedding Benchma…

200 篇论文

Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack…

计算与语言 · 计算机科学 2026-05-07 Minjie Qiang , Mingming Zhang , Xiaoyi Bao , Xing Fu , Yu Cheng , Weiqiang Wang , Zhongqing Wang , Ningtao Wang

This paper addresses the gap between general-purpose text embeddings and the specific demands of item retrieval tasks. We demonstrate the shortcomings of existing models in capturing the nuances necessary for zero-shot performance on item…

信息检索 · 计算机科学 2024-03-01 Yuxuan Lei , Jianxun Lian , Jing Yao , Mingqi Wu , Defu Lian , Xing Xie

We introduce VietSuperSpeech, a large-scale Vietnamese automatic speech recognition (ASR) dataset of 52,023 audio-text pairs totaling 267.39 hours, with a distinctive focus on casual conversational speech. Unlike existing Vietnamese ASR…

声音 · 计算机科学 2026-03-03 Loan Do , Thanh Ngoc Nguyen , Thanh Pham , Vinh Do , Hien Nguyen , Charlotte Nguyen

Recent advancements in hate speech detection (HSD) in Vietnamese have made significant progress, primarily attributed to the emergence of transformer-based pre-trained language models, particularly those built on the BERT architecture.…

计算与语言 · 计算机科学 2024-06-05 Luan Thanh Nguyen

Embedders play a central role in machine learning, projecting any object into numerical representations that can, in turn, be leveraged to perform various downstream tasks. The evaluation of embedding models typically depends on…

机器学习 · 计算机科学 2024-11-19 Maxime Darrin , Philippe Formont , Ismail Ben Ayed , Jackie CK Cheung , Pablo Piantanida

Machine translation (MT) systems universally degrade when faced with code-mixed text. This problem is more acute for low-resource languages that lack dedicated parallel corpora. This work directly addresses this gap for Vietnamese-English,…

计算与语言 · 计算机科学 2026-01-12 Hieu Tran , Phuong-Anh Nguyen-Le , Huy Nghiem , Quang-Nhan Nguyen , Wei Ai , Marine Carpuat

In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 million sentences describing over 40 million images crawled…

机器学习 · 计算机科学 2016-11-28 Junhua Mao , Jiajing Xu , Yushi Jing , Alan Yuille

This paper presents a zero-shot system for fact-checked claim retrieval. We employed several state-of-the-art large language models to obtain text embeddings. The models were then combined to obtain the best possible result. Our approach…

计算与语言 · 计算机科学 2025-08-14 Ladislav Lenc , Daniel Cífka , Jiří Martínek , Jakub Šmíd , Pavel Král

Vision-Language Translation (VLT) is a challenging task that requires accurately recognizing multilingual text embedded in images and translating it into the target language with the support of visual context. While recent Large…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Xintong Wang , Jingheng Pan , Yixiao Liu , Xiaohu Zhao , Chenyang Lyu , Minghao Wu , Chris Biemann , Longyue Wang , Linlong Xu , Weihua Luo , Kaifu Zhang

Machine comprehension(MC) style question answering is a representative problem in natural language processing. Previous methods rarely spend time on the improvement of encoding layer, especially the embedding of syntactic information and…

人工智能 · 计算机科学 2017-07-31 Boyuan Pan , Hao Li , Zhou Zhao , Bin Cao , Deng Cai , Xiaofei He

Due to the lack of large-scale datasets, the prevailing approach in visual sentiment analysis is to leverage models trained for object classification in large datasets like ImageNet. However, objects are sentiment neutral which hinders the…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Ziad Al-Halah , Andrew Aitken , Wenzhe Shi , Jose Caballero

This paper describes RETVec, an efficient, resilient, and multilingual text vectorizer designed for neural-based text processing. RETVec combines a novel character encoding with an optional small embedding model to embed words into a…

计算与语言 · 计算机科学 2024-04-24 Elie Bursztein , Marina Zhang , Owen Vallis , Xinyu Jia , Alexey Kurakin

Accurate mapping of the built asset information to established data classification systems and taxonomies is crucial for effective asset management, whether for compliance at project handover or ad-hoc data integration scenarios. Due to the…

计算与语言 · 计算机科学 2024-11-20 Mehrzad Shahinmoghadam , Ali Motamedi

A recent trend in speech processing is the use of embeddings created through machine learning models trained on a specific task with large datasets. By leveraging the knowledge already acquired, these models can be reused in new tasks where…

声音 · 计算机科学 2023-06-27 Andrés Carofilis , Laura Fernández-Robles , Enrique Alegre , Eduardo Fidalgo

We present ViT5, a pretrained Transformer-based encoder-decoder model for the Vietnamese language. With T5-style self-supervised pretraining, ViT5 is trained on a large corpus of high-quality and diverse Vietnamese texts. We benchmark ViT5…

计算与语言 · 计算机科学 2022-05-27 Long Phan , Hieu Tran , Hieu Nguyen , Trieu H. Trinh

Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our results show that this assumption does not hold in practice.…

Hateful memes are widespread in social media and convey negative information. The main challenge of hateful memes detection is that the expressive meaning can not be well recognized by a single modality. In order to further integrate modal…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Weibo Zhang , Guihua Liu , Zhuohua Li , Fuqing Zhu

Due to privacy restrictions, there's a shortage of publicly available speech recognition datasets in the medical domain. In this work, we present VietMed - a Vietnamese speech recognition dataset in the medical domain comprising 16h of…

计算与语言 · 计算机科学 2025-04-07 Khai Le-Duc

Fact-checking is essential due to the explosion of misinformation in the media ecosystem. Although false information exists in every language and country, most research to solve the problem mainly concentrated on huge communities like…

计算与语言 · 计算机科学 2026-03-17 Hung Tuan Le , Long Truong To , Manh Trong Nguyen , Kiet Van Nguyen

As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing specialized VTON models. Meanwhile, universal multi-reference image editing models have…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Xiaoye Liang , Zhiyuan Qu , Mingye Zou , Jiaxin Liu , Lai Jiang , Mai Xu , Yiheng Zhu
‹ 上一页 1 8 9 10 下一页 ›