English
Related papers

Related papers: LPM 1.0: Video-based Character Performance Model

200 papers

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xingpei Ma , Jiaran Cai , Yuansheng Guan , Shenneng Huang , Qiang Zhang , Shunsi Zhang

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

It has been established in recent work that Large Language Models (LLMs) can be prompted to "self-play" conversational games that probe certain capabilities (general instruction following, strategic goal orientation, language understanding…

Computation and Language · Computer Science 2024-06-03 Anne Beyer , Kranti Chalamalasetti , Sherzod Hakimov , Brielen Madureira , Philipp Sadler , David Schlangen

Large Language Model (LLM) personas with explicit specifications of attributes, background, and behavioural tendencies are increasingly used to simulate human conversations for tasks such as user modeling, social reasoning, and behavioural…

Computation and Language · Computer Science 2026-03-04 Eliseo Bao , Anxo Perez , Xi Wang , Javier Parapar

Though Large Vision-Language Models (LVLMs) are being actively explored in medicine, their ability to conduct complex real-world telemedicine consultations combining accurate diagnosis with professional dialogue remains underexplored. This…

Human-Computer Interaction · Computer Science 2025-11-12 Ivan Sviridov , Amina Miftakhova , Artemiy Tereshchenko , Galina Zubkova , Pavel Blinov , Andrey Savchenko

Evaluating the abilities of large models and manifesting their gaps are challenging. Current benchmarks adopt either ground-truth-based score-form evaluation on static datasets or indistinct textual chatbot-style human preferences…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Zijian Chen , Lirong Deng , Zhengyu Chen , Kaiwei Zhang , Qi Jia , Yuan Tian , Yucheng Zhu , Guangtao Zhai

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

Multimodal Large Language Models (MLLMs) are distinguished by their multimodal comprehensive ability and widely used in many real-world applications including GPT-4o, autonomous driving and robotics. Despite their impressive performance,…

Machine Learning · Computer Science 2024-09-17 Zhenyu Ning , Jieru Zhao , Qihao Jin , Wenchao Ding , Minyi Guo

Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xingpei Ma , Shenneng Huang , Jiaran Cai , Yuansheng Guan , Shen Zheng , Hanfeng Zhao , Qiang Zhang , Shunsi Zhang

Long-context understanding poses significant challenges in natural language processing, particularly for real-world dialogues characterized by speech-based elements, high redundancy, and uneven information density. Although large language…

Computation and Language · Computer Science 2025-04-25 Yongxuan Wu , Runyu Chen , Peiyu Liu , Hongjin Qian

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

People now regularly interface with Large Language Models (LLMs) via speech and text (e.g., Bard) interfaces. However, little is known about the relationship between how users anthropomorphize an LLM system (i.e., ascribe human-like…

Human-Computer Interaction · Computer Science 2024-05-13 Michelle Cohn , Mahima Pushkarna , Gbolahan O. Olanubi , Joseph M. Moran , Daniel Padgett , Zion Mengesha , Courtney Heldreth

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

Artificial Intelligence · Computer Science 2024-11-06 Zhifei Xie , Changqiao Wu

Optimization is as much about modeling the right problem as solving it. Identifying the right objectives, constraints, and trade-offs demands extensive interaction between researchers and stakeholders. Large language models can empower…

Artificial Intelligence · Computer Science 2026-04-06 Joshua Drossman , Alexandre Jacquillat , Sébastien Martin

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Luchuan Song , Pinxin Liu , Haiyang Liu , Zhenchao Jin , Yolo Yunlong Tang , Zichong Xu , Susan Liang , Jing Bi , Jason J Corso , Chenliang Xu

Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are unclear, impatient, or reluctant to share information. However, collecting real interaction…

Artificial Intelligence · Computer Science 2026-05-14 Harshita Chopra , Kshitish Ghate , Aylin Caliskan , Tadayoshi Kohno , Chirag Shah , Natasha Jaques

We introduce EgoMem, the first lifelong memory agent tailored for full-duplex models that process real-time omnimodal streams. EgoMem enables real-time models to recognize multiple users directly from raw audiovisual streams, to provide…

Artificial Intelligence · Computer Science 2026-02-02 Yiqun Yao , Naitong Yu , Xiang Li , Xin Jiang , Xuezhi Fang , Wenjia Ma , Xuying Meng , Jing Li , Aixin Sun , Yequan Wang

Recent Large Language Models (LLMs) have shown remarkable capabilities in mimicking fictional characters or real humans in conversational settings. However, the realism and consistency of these responses can be further enhanced by providing…

Computation and Language · Computer Science 2023-12-29 Seokhoon Jeong , Assentay Makhmud

Doctor-patient consultations require multi-turn, context-aware communication tailored to diverse patient personas. Training or evaluating doctor LLMs in such settings requires realistic patient interaction systems. However, existing…

Artificial Intelligence · Computer Science 2025-10-30 Daeun Kyung , Hyunseung Chung , Seongsu Bae , Jiho Kim , Jae Ho Sohn , Taerim Kim , Soo Kyung Kim , Edward Choi

Customizable role-playing in large language models (LLMs), also known as character generalization, is gaining increasing attention for its versatility and cost-efficiency in developing and deploying role-playing dialogue agents. This study…

Computation and Language · Computer Science 2025-02-19 Xiaoyang Wang , Hongming Zhang , Tao Ge , Wenhao Yu , Dian Yu , Dong Yu
‹ Prev 1 4 5 6 7 8 10 Next ›