中文
相关论文

相关论文: Multi-Lingual DALL-E Storytime

200 篇论文

Pre-trained vision and language models such as CLIP have witnessed remarkable success in connecting images and texts with a primary focus on English texts. Despite recent efforts to extend CLIP to support other languages, disparities in…

计算与语言 · 计算机科学 2023-10-31 Zhen Zhang , Jialu Wang , Xin Eric Wang

Preserving ancient languages is essential for understanding humanity's cultural and linguistic heritage, yet Old English remains critically under-resourced, limiting its accessibility to modern natural language processing (NLP) techniques.…

计算与语言 · 计算机科学 2025-07-29 Rodrigo Gabriel Salazar Alva , Matías Nuñez , Cristian López , Javier Martín Arista

The diversity of human language, shaped by social, cultural, and regional influences, presents significant challenges for natural language processing (NLP) systems. Existing benchmarks often overlook intra-language variations, leaving…

计算与语言 · 计算机科学 2025-04-11 Abhay Gupta , Jacob Cheung , Philip Meng , Shayan Sayyed , Austen Liao , Kevin Zhu , Sean O'Brien

We present TALE, a novel training-free framework harnessing the generative capabilities of text-to-image diffusion models to address the cross-domain image composition task that focuses on flawlessly incorporating user-specified objects…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Kien T. Pham , Jingye Chen , Qifeng Chen

Text-to-image generation advancements have been predominantly English-centric, creating barriers for non-English speakers and perpetuating digital inequities. While existing systems rely on translation pipelines, these introduce semantic…

计算与语言 · 计算机科学 2025-07-09 Mohammad Mahdi Derakhshani , Dheeraj Varghese , Marzieh Fadaee , Cees G. M. Snoek

The proliferation of deepfake technologies poses urgent challenges and serious risks to digital integrity, particularly within critical sectors such as forensics, journalism, and the legal system. While existing detection systems have made…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Shahroz Tariq , Simon S. Woo , Priyanka Singh , Irena Irmalasari , Saakshi Gupta , Dev Gupta

In this paper, we examine how generative machine learning systems produce a new politics of visual culture. We focus on DALL-E 2 and related models as an emergent approach to image-making that operates through the cultural techniques of…

计算机与社会 · 计算机科学 2022-11-14 Fabian Offert , Thao Phan

Text-to-Image artificial intelligence (AI) recently saw a major breakthrough with the release of Dall-E and its open-source counterpart, Stable Diffusion. These programs allow anyone to create original visual art pieces by simply providing…

人机交互 · 计算机科学 2023-01-06 Nassim Dehouche , Kullathida Dehouche

While mechanistic interpretability has developed powerful tools to analyze the internal workings of Large Language Models (LLMs), their complexity has created an accessibility gap, limiting their use to specialists. We address this…

计算与语言 · 计算机科学 2026-02-23 Aaron Louis Eidt , Nils Feldhus

Recent image generation models excel at creating high-quality images from brief captions. However, they fail to maintain consistency of multiple instances across images when encountering lengthy contexts. This inconsistency is largely due…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Zilyu Ye , Jinxiu Liu , Ruotian Peng , Jinjin Cao , Zhiyang Chen , Yiyang Zhang , Ziwei Xuan , Mingyuan Zhou , Xiaoqian Shen , Mohamed Elhoseiny , Qi Liu , Guo-Jun Qi

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models…

计算与语言 · 计算机科学 2026-02-10 Bonaventure F. P. Dossou , Henri Aïdasso

In recent years, large language models (e.g., Open AI's GPT-4, Meta's LLaMa, Google's PaLM) have become the dominant approach for building AI systems to analyze and generate language online. However, the automated systems that increasingly…

计算与语言 · 计算机科学 2023-06-14 Gabriel Nicholas , Aliya Bhatia

We present RALL-E, a robust language modeling method for text-to-speech (TTS) synthesis. While previous work based on large language models (LLMs) shows impressive performance on zero-shot TTS, such methods often suffer from poor…

音频与语音处理 · 电气工程与系统科学 2024-05-21 Detai Xin , Xu Tan , Kai Shen , Zeqian Ju , Dongchao Yang , Yuancheng Wang , Shinnosuke Takamichi , Hiroshi Saruwatari , Shujie Liu , Jinyu Li , Sheng Zhao

Recent advances in text-to-image synthesis have led to large pretrained transformers with excellent capabilities to generate visualizations from a given text. However, these models are ill-suited for specialized tasks like story…

计算机视觉与模式识别 · 计算机科学 2022-09-14 Adyasha Maharana , Darryl Hannan , Mohit Bansal

Currently, natural language processing (NLP) models proliferate language discrimination leading to potentially harmful societal impacts as a result of biased outcomes. For example, part-of-speech taggers trained on Mainstream American…

计算与语言 · 计算机科学 2022-06-22 Jamell Dacon

We introduce KALL-E, a novel autoregressive (AR) language model for text-to-speech (TTS) synthesis that operates by predicting the next distribution of continuous speech frames. Unlike existing methods, KALL-E directly models the continuous…

音频与语音处理 · 电气工程与系统科学 2025-09-18 Kangxiang Xia , Xinfa Zhu , Jixun Yao , Wenjie Tian , Wenhao Li , Lei Xie

Generative AI models like DALL-E 2 can interpret textual prompts and generate high-quality images exhibiting human creativity. Though public enthusiasm is booming, systematic auditing of potential gender biases in AI-generated images…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Luhang Sun , Mian Wei , Yibing Sun , Yoo Ji Suh , Liwei Shen , Sijia Yang

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence…

计算与语言 · 计算机科学 2025-08-21 Admitos Passadakis , Yingjin Song , Albert Gatt

The advancement of computational psychology requires AI tools capable of deeply understanding counseling dialogues. Existing audio language models (AudioLLMs) often rely on single speech encoders pre-trained on general data, struggling to…

音频与语音处理 · 电气工程与系统科学 2025-10-06 Yongqi Kang , Yong Zhao