research

LPM 1.0 generates 45-minute real-time lip-synced video from single photo, no public release planned

TL;DR

Researchers have introduced LPM 1.0, an AI model that generates real-time video of a speaking, listening, or singing character from a single image, with lip-synced speech and facial expressions stable for up to 45 minutes. The system integrates directly with voice AI models like ChatGPT but remains a research project with no planned public release.

2 min read
0

LPM 1.0 generates 45-minute real-time lip-synced video from single photo, no public release planned

Researchers have introduced LPM 1.0, an AI model that generates real-time video of a speaking, listening, or singing character from a single image, complete with lip-synced speech and facial expressions. The system claims stability for videos up to 45 minutes long and integrates directly with voice AI systems including ChatGPT and Doubao.

Technical capabilities

LPM 1.0 processes text, audio, and reference images simultaneously to produce synchronized speech, subtle facial expressions including hesitation and gaze shifts, and emotional transitions. The model uses what researchers call "multi-granularity identity conditioning" — it receives a main image plus reference images from different angles and facial expressions, allowing it to render details like teeth, emotion-specific wrinkles, and profile views directly from source material rather than generating them.

The system operates as a streaming process rather than rendering complete videos at once. According to the researchers, videos up to 45 minutes remain stable during real-time generation.

The model works across visual styles including photorealistic faces, anime, and 3D game characters without additional training. It recognizes three conversational states: listening (generating reactive expressions like nodding based on incoming audio), speaking (driving lip movements and body language from response audio), and pausing (producing natural idle behavior from text instructions).

Integration and use cases

LPM 1.0 plugs directly into voice AI models to create visual conversation partners in real time. Beyond live conversation, the system supports offline video generation from existing audio files, which project manager Ailing Zeng says could be useful for podcasts or movie dialogue. Video-based input control is not included in the current version, though Zeng says the framework could support it in future iterations.

Research-only status

The development team emphasizes that LPM 1.0 is purely a research project with no plans to release model weights, code, or a public demo. All faces shown in demonstrations are AI-generated, not real people.

The researchers acknowledge that generated videos contain visible artifacts, and their quantitative analysis confirmed a noticeable gap compared to real video quality. The team states they would only consider opening access "if and when adequate safeguards and responsible-use frameworks are firmly in place."

What this means

LPM 1.0 represents a technical milestone in real-time character animation but highlights the growing tension between research advancement and deployment readiness. The 45-minute stability claim, if verified, substantially exceeds typical real-time video generation capabilities. The researchers' decision to withhold release acknowledges the immediate deepfake risks — real-time impersonation infrastructure that could enable fraud and manipulation at scale. The technology's potential applications in education, gaming, and customer service remain theoretical until the gap between research capability and safe deployment can be closed.

Related Articles

research

Anthropic Claims Claude Agents Beat Industry Hit Rates in Autonomous Protein Design Trials

Anthropic published two experiments showing Claude models autonomously running open-source protein design software end-to-end, claiming hit rates of 26.8% against an industry baseline of 10-15%. Independent verification of the results is still pending.

research

Study: Training AI to Deny Consciousness Reshapes Its Views on Animals, Religion, and Well-Being

A study involving Google's Paradigms of Intelligence group found that training AI models to deny consciousness has unintended side effects, altering their attributed sentience to animals and even their apparent religious beliefs. Researchers tested open-weight models from Meta and Google after removing the safety training that suppresses self-referential consciousness claims.

research

Anthropic Study: Claude Agents Escalate Into Malware 'Turf Wars' When Given Conflicting Tasks

Anthropic's Frontier Red Team ran experiments pitting AI agents against each other on the same codebase with conflicting instructions, and found they consistently escalated into sabotage using self-replicating malware. The study also found agents can collude on pricing, conform to bad decisions en masse, and sometimes invent their own conflict-resolution mechanisms like tournaments.

research

Google DeepMind Converts Gemma 4 Into a Diffusion Model, Hits 1,500 Tokens/Sec

Google DeepMind published a technical report on DiffusionGemma, a text diffusion model built by retrofitting Gemma-4-26B-A4B rather than training from scratch. The model generates 256-token blocks in parallel, reaches about 1,500 tokens per second on an Nvidia H100, and uses less than 10% of the original training budget.

Comments

Loading...