中文
相关论文

相关论文: InDEX: Indonesian Idiom and Expression Dataset for…

200 篇论文

While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set…

计算与语言 · 计算机科学 2017-01-11 Omer Levy , Anders Søgaard , Yoav Goldberg

The goal of this paper is to learn more about how idiomatic information is structurally encoded in embeddings, using a structural probing method. We repurpose an existing English verbal multi-word expression (MWE) dataset to suit the…

计算与语言 · 计算机科学 2023-04-28 Filip Klubička , Vasudevan Nedumpozhimana , John D. Kelleher

Prompt-based methods leverage the knowledge of pre-trained language models (PLMs) trained with a masked language modeling (MLM) objective; however, these methods are sensitive to template, verbalizer, and few-shot instance selection,…

计算与语言 · 计算机科学 2025-07-02 Mohna Chakraborty , Adithya Kulkarni , Qi Li

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like…

计算与语言 · 计算机科学 2026-01-15 Fadhil Muhammad , Alwin Djuliansah , Adrian Aryaputra Hamzah , Kurniawati Azizah

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks:…

计算与语言 · 计算机科学 2025-08-19 Alham Fikri Aji , Trevor Cohn

Large Language Models (LLMs) have gained significant traction across critical domains owing to their impressive contextual understanding and generative capabilities. However, their increasing deployment in high stakes applications…

计算与语言 · 计算机科学 2025-10-06 Santhosh G S , Akshay Govind S , Gokul S Krishnan , Balaraman Ravindran , Sriraam Natarajan

While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance under complex interleaved instructions. This limitation stems…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yabo Zhang , Kunchang Li , Dewei Zhou , Xinyu Huang , Xun Wang

Currently, text-to-image synthesis uses text encoder and image generator architecture. Research on this topic is challenging. This is because of the domain gap between natural language and vision. Nowadays, most research on this topic only…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Made Raharja Surya Mahadi , Nugraha Priya Utama

Detecting gender-based hate speech in Indonesian social media remains challenging due to limited labeled datasets. While binary hate speech classification has advanced, a more granular category like gender-targeted hate speech is…

计算与语言 · 计算机科学 2025-03-07 Muhammad Amien Ibrahim , Faisal , Tora Sangputra Yopie Winarto , Zefanya Delvin Sulistiya

Sentence fusion is the task of joining several independent sentences into a single coherent text. Current datasets for sentence fusion are small and insufficient for training modern neural models. In this paper, we propose a method for…

计算与语言 · 计算机科学 2019-03-19 Mor Geva , Eric Malmi , Idan Szpektor , Jonathan Berant

Standard informativeness measures used to evaluate Automatic Text Summarization mostly rely on n-gram overlapping between the automatic summary and the reference summaries. These measures differ from the metric they use (cosine, ROUGE,…

信息检索 · 计算机科学 2020-04-16 Carlos-Emiliano González-Gallardo , Eric SanJuan , Juan-Manuel Torres-Moreno

Large-scale Vision-Language models have achieved remarkable results in various domains, such as image captioning and conditioned image generation. Nevertheless, these models still encounter difficulties in achieving human-like compositional…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Jiahao Liu , Senhao Cao

In this paper, we propose a new word embedding based corpus consisting of more than 61 million words crawled from multiple web resources. We design a preprocessing pipeline for the filtration of unwanted text from crawled data. Afterwards,…

计算与语言 · 计算机科学 2024-08-29 Wazir Ali , Saifullah Tumrani , Jay Kumar , Tariq Rahim Soomro

The Yunshan Cup 2020 track focused on creating a framework for evaluating different methods of part-of-speech (POS). There were two tasks for this track: (1) POS tagging for the Indonesian language, and (2) POS tagging for the Lao tagging.…

计算与语言 · 计算机科学 2022-04-07 Yingwen Fu , Jinyi Chen , Nankai Lin , Xixuan Huang , Xinying Qiu , Shengyi Jiang

This study presents a domain adaptation approach for speaker diarization targeting conversational Indonesian audio. We address the challenge of adapting an English-centric diarization pipeline to a low-resource language by employing…

This paper describes a preliminary study for producing and distributing a large-scale database of embeddings from the Portuguese Twitter stream. We start by experimenting with a relatively small sample and focusing on three challenges:…

计算与语言 · 计算机科学 2017-09-05 Pedro Saleiro , Luís Sarmento , Eduarda Mendes Rodrigues , Carlos Soares , Eugénio Oliveira

This paper proposes the task of automatic assessment of Sentence Translation Exercises (STEs), that have been used in the early stage of L2 language learning. We formalize the task as grading student responses for each rubric criterion…

计算与语言 · 计算机科学 2024-03-07 Naoki Miura , Hiroaki Funayama , Seiya Kikuchi , Yuichiroh Matsubayashi , Yuya Iwase , Kentaro Inui

MICE is a corpus of emotion words in four languages which is currently working progress. There are two sections to this study, Part I: Emotion word corpus and Part II: Emotion word survey. In Part 1, the method of how the emotion data is…

计算与语言 · 计算机科学 2021-06-10 Ng Bee Chin , Yosephine Susanto , Erik Cambria

Clickbait spoiling aims to generate a short text to satisfy the curiosity induced by a clickbait post. As it is a newly introduced task, the dataset is only available in English so far. Our contributions include the construction of manually…

计算与语言 · 计算机科学 2023-10-13 Ni Putu Intan Maharani , Ayu Purwarianti , Alham Fikri Aji

Significant progress has been made on Indonesian NLP. Nevertheless, exploration of the code-mixing phenomenon in Indonesian is limited, despite many languages being frequently mixed with Indonesian in daily conversation. In this work, we…

计算与语言 · 计算机科学 2023-11-22 Muhammad Farid Adilazuarda , Samuel Cahyawijaya , Genta Indra Winata , Pascale Fung , Ayu Purwarianti