中文
相关论文

相关论文: How Does India Cook Biryani?

200 篇论文

Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, a versatile food…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Yuehao Yin , Huiyan Qi , Bin Zhu , Jingjing Chen , Yu-Gang Jiang , Chong-Wah Ngo

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a…

Vision-language models (VLMs) have shown impressive performance in substantial downstream multi-modal tasks. However, only comparing the fine-tuned performance on downstream tasks leads to the poor interpretability of VLMs, which is adverse…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Zheng Ma , Mianzhi Pan , Wenhan Wu , Kanzhi Cheng , Jianbing Zhang , Shujian Huang , Jiajun Chen

The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of benchmarks and evaluation methodologies specifically designed…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Yogesh Kumar

Large language models (LLMs) are used worldwide, yet exhibit Western cultural tendencies. Many countries are now building ``regional'' or ``sovereign'' LLMs, but it remains unclear whether they reflect local values and practices or merely…

计算与语言 · 计算机科学 2026-01-26 Dhruv Agarwal , Anya Shukla , Sunayana Sitaram , Aditya Vashistha

Any national cuisine is a sum total of its variety of regional cuisines, which are the cultural and historical identifiers of their respective regions. India is home to a number of regional cuisines that showcase its culinary diversity.…

物理与社会 · 物理学 2016-02-17 Anupam Jain , Rakhi N K , Ganesh Bagler

In this paper, we introduce Bangla-Bayanno, an open-ended Visual Question Answering (VQA) Dataset in Bangla, a widely used, low-resource language in multimodal AI research. The majority of existing datasets are either manually annotated…

计算与语言 · 计算机科学 2025-08-28 Mohammed Rakibul Hasan , Rafi Majid , Ahanaf Tahmid

Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It…

计算与语言 · 计算机科学 2018-09-05 Semih Yagcioglu , Aykut Erdem , Erkut Erdem , Nazli Ikizler-Cinbis

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper,…

人工智能 · 计算机科学 2025-01-03 Shudong Liu , Yiqiao Jin , Cheng Li , Derek F. Wong , Qingsong Wen , Lichao Sun , Haipeng Chen , Xing Xie , Jindong Wang

Multi-language recipe personalisation and recommendation is an under-explored field of information retrieval in academic and production systems. The existing gaps in our current understanding are numerous, even on fundamental questions such…

信息检索 · 计算机科学 2020-08-19 Niall Twomey , Mikhail Fain , Andrey Ponikar , Nadine Sarraf

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Munan Ning , Bin Zhu , Yujia Xie , Bin Lin , Jiaxi Cui , Lu Yuan , Dongdong Chen , Li Yuan

News videos are carefully edited multimodal narratives that combine narration, visuals, and external quotations into coherent storylines. In recent years, there have been significant advances in evaluating multimodal large language models…

机器学习 · 计算机科学 2026-01-08 Zibo Liu , Muyang Li , Zhe Jiang , Shigang Chen

Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly "understand" real-world dynamics, as accurate numerical deduction necessitates a profound grasp of temporal events, object…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Shaoyang Cui , Lingbei Meng

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to…

Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness…

Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models synthesize images and…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Bernini Team , Chenchen Liu , Junyi Chen , Lei Li , Lu Chi , Mingzhen Sun , Zhuoying Li , Yi Fu , Ruoyu Guo , Yiheng Wu , Ge Bai , Zehuan Yuan

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken languages but lack performance on low-resource languages and…

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an…

Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference.…