中文
相关论文

相关论文: Crowdsource, Crawl, or Generate? Creating SEA-VL, …

200 篇论文

Multimodal Large Language Models excel in high-resource settings, but often misinterpret long-tail cultural entities and underperform in low-resource languages. To address this gap, we propose a data-centric approach that directly grounds…

计算与语言 · 计算机科学 2025-08-13 Jean de Dieu Nyandwi , Yueqi Song , Simran Khanuja , Graham Neubig

As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Bryan Chen Zhengyu Tan , Zheng Weihua , Zhengyuan Liu , Nancy F. Chen , Hwaran Lee , Kenny Tsu Wei Choo , Roy Ka-Wei Lee

Current multimodal approaches predominantly treat visual generation as an external process, relying on pixel rendering or code execution, thereby overlooking the native visual representation capabilities latent within Large Language Models…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yiren Zheng , Shibo Li , Jiaming Liu , Haofan Wang , Yiren Song

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approaches typically rely on human annotators, making the process…

计算与语言 · 计算机科学 2025-06-02 ChaeHun Park , Yujin Baek , Jaeseok Kim , Yu-Jung Heo , Du-Seong Chang , Jaegul Choo

Verifying the authenticity of AI-generated images presents a growing challenge on social media platforms these days. While vision-language models (VLMs) like CLIP outdo in multimodal representation, their capacity for AI-generated image…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Ziyang Ou

Document retrieval is an important task for search and Retrieval-Augmented Generation (RAG) applications. Large Language Models (LLMs) have contributed to improving the accuracy of text-based document retrieval. However, documents with…

Cross-modal retrieval (CMR) is a fundamental task in multimedia research, focused on retrieving semantically relevant targets across different modalities. While traditional CMR methods match text and image via embedding-based similarity…

信息检索 · 计算机科学 2025-04-18 Haoxuan Li , Yi Bin , Yunshan Ma , Guoqing Wang , Yang Yang , See-Kiong Ng , Tat-Seng Chua

This paper describes Georeference Contrastive Learning of visual Representation (GeoCLR) for efficient training of deep-learning Convolutional Neural Networks (CNNs). The method leverages georeference information by generating a similar…

计算机视觉与模式识别 · 计算机科学 2022-06-28 Takaki Yamada , Adam Prügel-Bennett , Stefan B. Williams , Oscar Pizarro , Blair Thornton

Successful applications of complex vision-based behaviours underwater have lagged behind progress in terrestrial and aerial domains. This is largely due to the degraded image quality resulting from the physical phenomena involved in…

计算机视觉与模式识别 · 计算机科学 2023-03-08 Stewart Jamieson , Jonathan P. How , Yogesh Girdhar

Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. However, prior work has noted that such generated data lacks diversity. In this paper, we propose…

计算与语言 · 计算机科学 2026-04-29 Avinash Amballa , Yashas Malur Saidutta , Chi-Heng Lin , Vivek Kulkarni , Srinivas Chappidi

As an emergent process, creativity relies on explorations via sampling and prototyping for problem construction. These activities compile knowledge, provide a context enveloping the solution, and answer questions. With Generative AI,…

人机交互 · 计算机科学 2026-01-12 Alicia Guo , David Ledo , George Fitzmaurice , Fraser Anderson

The availability of accurate localization is critical for multi-robot exploration strategies; noisy or inconsistent localization causes failure in meeting exploration objectives. We aim to achieve high localization accuracy with…

机器人学 · 计算机科学 2023-06-23 Ehsan Latif , Ramviyas Parasuraman

Recent advances in generative AI have sparked renewed interest and expanded possibilities for music generation. However, the performance and versatility of these systems across musical genres are heavily influenced by the availability of…

声音 · 计算机科学 2025-08-26 Atharva Mehta , Shivam Chauhan , Monojit Choudhury

The vision of an inclusive World Wide Web is impeded by a severe linguistic divide, particularly for communities in low-resource regions of Southeast Asia. While large language models (LLMs) offer a potential solution for translation, their…

计算与语言 · 计算机科学 2026-03-23 Zhixiang Lu , Chong Zhang , Yulong Li , Angelos Stefanidis , Anh Nguyen , Imran Razzak , Jionglong Su , Zhengyong Jiang

Representation learning approaches require a massive amount of discriminative training data, which is unavailable in many scenarios, such as healthcare, smart city, education, etc. In practice, people refer to crowdsourcing to get annotated…

机器学习 · 计算机科学 2021-12-17 Yang Hao , Wenbiao Ding , Zitao Liu

Large language models (LLMs) and vision-language models (VLMs) have demonstrated remarkable performance across a wide range of tasks and domains. Despite this promise, spatial understanding and reasoning -- a fundamental component of human…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Jiayu Wang , Yifei Ming , Zhenmei Shi , Vibhav Vineet , Xin Wang , Yixuan Li , Neel Joshi

This paper examines how deep learning (DL) representations, in contrast to traditional engineered features, can support semantic interaction (SI) in visual analytics. SI attempts to model user's cognitive reasoning via their interaction…

人机交互 · 计算机科学 2020-08-03 Yali Bian , John Wenskovitch , Chris North

Diffusion models are a strong backbone for visual generation, but their inherently sequential denoising process leads to slow inference. Previous methods accelerate sampling by caching and reusing intermediate outputs based on feature…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jiwoo Chung , Sangeek Hyun , MinKyu Lee , Byeongju Han , Geonho Cha , Dongyoon Wee , Youngjun Hong , Jae-Pil Heo

As text-to-image models become increasingly prevalent, ensuring their equitable performance across diverse cultural contexts is critical. Efforts to mitigate cross-cultural biases have been hampered by trade-offs, including a loss in…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Arnav Yayavaram , Siddharth Yayavaram , Simran Khanuja , Michael Saxon , Graham Neubig

Visual metaphors are powerful rhetorical devices used to persuade or communicate creative ideas through images. Similar to linguistic metaphors, they convey meaning implicitly through symbolism and juxtaposition of the symbols. We propose a…