English
Related papers

Related papers: Pythia v0.1: the Winning Entry to the VQA Challeng…

200 papers

Visual Question Answering (VQA) has emerged as a promising area of research to develop AI-based systems for enabling interactive and immersive learning. Numerous VQA datasets have been introduced to facilitate various tasks, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Ngoc Dung Huynh , Mohamed Reda Bouadjenek , Sunil Aryal , Imran Razzak , Hakim Hacid

We describe our two-stage system for the Multilingual Information Access (MIA) 2022 Shared Task on Cross-Lingual Open-Retrieval Question Answering. The first stage consists of multilingual passage retrieval with a hybrid dense and sparse…

Computation and Language · Computer Science 2022-07-19 Zhucheng Tu , Sarguna Janani Padmanabhan

Reinforcement fine-tuning (RFT) is a proliferating paradigm for LMM training. Analogous to high-level reasoning tasks, RFT is similarly applicable to low-level vision domains, including image quality assessment (IQA). Existing RFT-based IQA…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Ziheng Jia , Jiaying Qian , Zicheng Zhang , Zijian Chen , Xiongkuo Min

Question-answering for domain-specific applications has recently attracted much interest due to the latest advancements in large language models (LLMs). However, accurately assessing the performance of these applications remains a…

This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question -- an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Chenhui Gou , Zilong Chen , Zeyu Wang , Feng Li , Deyao Zhu , Zicheng Duan , Kunchang Li , Chaorui Deng , Hongyi Yuan , Haoqi Fan , Cihang Xie , Jianfei Cai , Hamid Rezatofighi

In visual question answering (VQA), an algorithm must answer text-based questions about images. While multiple datasets for VQA have been created since late 2014, they all have flaws in both their content and the way algorithms are…

Computer Vision and Pattern Recognition · Computer Science 2017-09-15 Kushal Kafle , Christopher Kanan

Perceptual organization remains one of the very few established theories on the human visual system. It underpinned many pre-deep seminal works on segmentation and detection, yet research has seen a rapid decline since the preferential…

Computer Vision and Pattern Recognition · Computer Science 2021-04-09 Yonggang Qi , Kai Zhang , Aneeshan Sain , Yi-Zhe Song

This study systematically evaluates 27 frontier Large Language Models on eight biology benchmarks spanning molecular biology, genetics, cloning, virology, and biosecurity. Models from major AI developers released between November 2022 and…

Machine Learning · Computer Science 2025-05-23 Lennart Justen

Answering visual questions need acquire daily common knowledge and model the semantic connection among different parts in images, which is too difficult for VQA systems to learn from images with the only supervision from answers. Meanwhile,…

Computation and Language · Computer Science 2018-05-23 Jialin Wu , Zeyuan Hu , Raymond J. Mooney

Current work on Visual Question Answering (VQA) explore deterministic approaches conditioned on various types of image and question features. We posit that, in addition to image and question pairs, other modalities are useful for teaching…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Zixu Wang , Yishu Miao , Lucia Specia

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce…

Computation and Language · Computer Science 2022-03-29 Yingshan Chang , Mridu Narang , Hisami Suzuki , Guihong Cao , Jianfeng Gao , Yonatan Bisk

Machine learning has advanced dramatically, narrowing the accuracy gap to humans in multimodal tasks like visual question answering (VQA). However, while humans can say "I don't know" when they are uncertain (i.e., abstain from answering a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Spencer Whitehead , Suzanne Petryk , Vedaad Shakib , Joseph Gonzalez , Trevor Darrell , Anna Rohrbach , Marcus Rohrbach

Prior studies in privacy policies frame the question answering (QA) task as identifying the most relevant text segment or a list of sentences from a policy document given a user query. Existing labeled datasets are heavily imbalanced (only…

Computation and Language · Computer Science 2023-04-25 Md Rizwan Parvez , Jianfeng Chi , Wasi Uddin Ahmad , Yuan Tian , Kai-Wei Chang

Building a deep learning model for a Question-Answering (QA) task requires a lot of human effort, it may need several months to carefully tune various model architectures and find a best one. It's even harder to find different excellent…

Computation and Language · Computer Science 2022-01-27 Sinan Tan , Hui Xue , Qiyu Ren , Huaping Liu , Jing Bai

Information comes in diverse modalities. Multimodal native AI models are essential to integrate real-world information and deliver comprehensive understanding. While proprietary multimodal native models exist, their lack of openness imposes…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Dongxu Li , Yudong Liu , Haoning Wu , Yue Wang , Zhiqi Shen , Bowen Qu , Xinyao Niu , Fan Zhou , Chengen Huang , Yanpeng Li , Chongyan Zhu , Xiaoyi Ren , Chao Li , Yifan Ye , Peng Liu , Lihuan Zhang , Hanshu Yan , Guoyin Wang , Bei Chen , Junnan Li

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We…

Computation and Language · Computer Science 2025-12-03 DeepSeek-AI , Aixin Liu , Aoxue Mei , Bangcai Lin , Bing Xue , Bingxuan Wang , Bingzheng Xu , Bochao Wu , Bowei Zhang , Chaofan Lin , Chen Dong , Chengda Lu , Chenggang Zhao , Chengqi Deng , Chenhao Xu , Chong Ruan , Damai Dai , Daya Guo , Dejian Yang , Deli Chen , Erhang Li , Fangqi Zhou , Fangyun Lin , Fucong Dai , Guangbo Hao , Guanting Chen , Guowei Li , H. Zhang , Hanwei Xu , Hao Li , Haofen Liang , Haoran Wei , Haowei Zhang , Haowen Luo , Haozhe Ji , Honghui Ding , Hongxuan Tang , Huanqi Cao , Huazuo Gao , Hui Qu , Hui Zeng , Jialiang Huang , Jiashi Li , Jiaxin Xu , Jiewen Hu , Jingchang Chen , Jingting Xiang , Jingyang Yuan , Jingyuan Cheng , Jinhua Zhu , Jun Ran , Junguang Jiang , Junjie Qiu , Junlong Li , Junxiao Song , Kai Dong , Kaige Gao , Kang Guan , Kexin Huang , Kexing Zhou , Kezhao Huang , Kuai Yu , Lean Wang , Lecong Zhang , Lei Wang , Liang Zhao , Liangsheng Yin , Lihua Guo , Lingxiao Luo , Linwang Ma , Litong Wang , Liyue Zhang , M. S. Di , M. Y Xu , Mingchuan Zhang , Minghua Zhang , Minghui Tang , Mingxu Zhou , Panpan Huang , Peixin Cong , Peiyi Wang , Qiancheng Wang , Qihao Zhu , Qingyang Li , Qinyu Chen , Qiushi Du , Ruiling Xu , Ruiqi Ge , Ruisong Zhang , Ruizhe Pan , Runji Wang , Runqiu Yin , Runxin Xu , Ruomeng Shen , Ruoyu Zhang , S. H. Liu , Shanghao Lu , Shangyan Zhou , Shanhuang Chen , Shaofei Cai , Shaoyuan Chen , Shengding Hu , Shengyu Liu , Shiqiang Hu , Shirong Ma , Shiyu Wang , Shuiping Yu , Shunfeng Zhou , Shuting Pan , Songyang Zhou , Tao Ni , Tao Yun , Tian Pei , Tian Ye , Tianyuan Yue , Wangding Zeng , Wen Liu , Wenfeng Liang , Wenjie Pang , Wenjing Luo , Wenjun Gao , Wentao Zhang , Xi Gao , Xiangwen Wang , Xiao Bi , Xiaodong Liu , Xiaohan Wang , Xiaokang Chen , Xiaokang Zhang , Xiaotao Nie , Xin Cheng , Xin Liu , Xin Xie , Xingchao Liu , Xingkai Yu , Xingyou Li , Xinyu Yang , Xinyuan Li , Xu Chen , Xuecheng Su , Xuehai Pan , Xuheng Lin , Xuwei Fu , Y. Q. Wang , Yang Zhang , Yanhong Xu , Yanru Ma , Yao Li , Yao Li , Yao Zhao , Yaofeng Sun , Yaohui Wang , Yi Qian , Yi Yu , Yichao Zhang , Yifan Ding , Yifan Shi , Yiliang Xiong , Ying He , Ying Zhou , Yinmin Zhong , Yishi Piao , Yisong Wang , Yixiao Chen , Yixuan Tan , Yixuan Wei , Yiyang Ma , Yiyuan Liu , Yonglun Yang , Yongqiang Guo , Yongtong Wu , Yu Wu , Yuan Cheng , Yuan Ou , Yuanfan Xu , Yuduan Wang , Yue Gong , Yuhan Wu , Yuheng Zou , Yukun Li , Yunfan Xiong , Yuxiang Luo , Yuxiang You , Yuxuan Liu , Yuyang Zhou , Z. F. Wu , Z. Z. Ren , Zehua Zhao , Zehui Ren , Zhangli Sha , Zhe Fu , Zhean Xu , Zhenda Xie , Zhengyan Zhang , Zhewen Hao , Zhibin Gou , Zhicheng Ma , Zhigang Yan , Zhihong Shao , Zhixian Huang , Zhiyu Wu , Zhuoshu Li , Zhuping Zhang , Zian Xu , Zihao Wang , Zihui Gu , Zijia Zhu , Zilin Li , Zipeng Zhang , Ziwei Xie , Ziyi Gao , Zizheng Pan , Zongqing Yao , Bei Feng , Hui Li , J. L. Cai , Jiaqi Ni , Lei Xu , Meng Li , Ning Tian , R. J. Chen , R. L. Jin , S. S. Li , Shuang Zhou , Tianyu Sun , X. Q. Li , Xiangyue Jin , Xiaojin Shen , Xiaosha Chen , Xinnan Song , Xinyi Zhou , Y. X. Zhu , Yanping Huang , Yaohui Li , Yi Zheng , Yuchen Zhu , Yunxian Ma , Zhen Huang , Zhipeng Xu , Zhongyu Zhang , Dongjie Ji , Jian Liang , Jianzhong Guo , Jin Chen , Leyi Xia , Miaojun Wang , Mingming Li , Peng Zhang , Ruyi Chen , Shangmian Sun , Shaoqing Wu , Shengfeng Ye , T. Wang , W. L. Xiao , Wei An , Xianzu Wang , Xiaowen Sun , Xiaoxiang Wang , Ying Tang , Yukun Zha , Zekai Zhang , Zhe Ju , Zhen Zhang , Zihua Qu

The advent of Artificial Intelligence (AI) tools, such as Large Language Models, has introduced new possibilities for Qualitative Data Analysis (QDA), offering both opportunities and challenges. To help navigate the responsible integration…

Computers and Society · Computer Science 2026-03-02 Elisabeth Kirsten , Annalina Buckmann , Leona Lassak , Nele Borgert , Abraham Mhaidli , Steffen Becker

We propose the task of free-form and open-ended Visual Question Answering (VQA). Given an image and a natural language question about the image, the task is to provide an accurate natural language answer. Mirroring real-world scenarios,…

Computation and Language · Computer Science 2016-10-28 Aishwarya Agrawal , Jiasen Lu , Stanislaw Antol , Margaret Mitchell , C. Lawrence Zitnick , Dhruv Batra , Devi Parikh

Existing VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, properties, and spatial relationships of entities, a significant…

Computer Vision and Pattern Recognition · Computer Science 2020-06-17 Ramprasaath R. Selvaraju , Purva Tendulkar , Devi Parikh , Eric Horvitz , Marco Ribeiro , Besmira Nushi , Ece Kamar

This paper describes our system for the SciVQA 2025 Shared Task on Scientific Visual Question Answering. Our system employs an ensemble of two Multimodal Large Language Models and various few-shot example retrieval strategies. The model and…

Computation and Language · Computer Science 2025-07-04 Christian Jaumann , Annemarie Friedrich , Rainer Lienhart
‹ Prev 1 4 5 6 7 8 10 Next ›