LPM 1.0 generates 45-minute real-time lip-synced video from single photo, no public release planned
Researchers have introduced LPM 1.0, an AI model that generates real-time video of a speaking, listening, or singing character from a single image, with lip-synced speech and facial expressions stable for up to 45 minutes. The system integrates directly with voice AI models like ChatGPT but remains a research project with no planned public release.
LPM 1.0 generates 45-minute real-time lip-synced video from single photo, no public release planned
Researchers have introduced LPM 1.0, an AI model that generates real-time video of a speaking, listening, or singing character from a single image, complete with lip-synced speech and facial expressions. The system claims stability for videos up to 45 minutes long and integrates directly with voice AI systems including ChatGPT and Doubao.
Technical capabilities
LPM 1.0 processes text, audio, and reference images simultaneously to produce synchronized speech, subtle facial expressions including hesitation and gaze shifts, and emotional transitions. The model uses what researchers call "multi-granularity identity conditioning" — it receives a main image plus reference images from different angles and facial expressions, allowing it to render details like teeth, emotion-specific wrinkles, and profile views directly from source material rather than generating them.
The system operates as a streaming process rather than rendering complete videos at once. According to the researchers, videos up to 45 minutes remain stable during real-time generation.
The model works across visual styles including photorealistic faces, anime, and 3D game characters without additional training. It recognizes three conversational states: listening (generating reactive expressions like nodding based on incoming audio), speaking (driving lip movements and body language from response audio), and pausing (producing natural idle behavior from text instructions).
Integration and use cases
LPM 1.0 plugs directly into voice AI models to create visual conversation partners in real time. Beyond live conversation, the system supports offline video generation from existing audio files, which project manager Ailing Zeng says could be useful for podcasts or movie dialogue. Video-based input control is not included in the current version, though Zeng says the framework could support it in future iterations.
Research-only status
The development team emphasizes that LPM 1.0 is purely a research project with no plans to release model weights, code, or a public demo. All faces shown in demonstrations are AI-generated, not real people.
The researchers acknowledge that generated videos contain visible artifacts, and their quantitative analysis confirmed a noticeable gap compared to real video quality. The team states they would only consider opening access "if and when adequate safeguards and responsible-use frameworks are firmly in place."
What this means
LPM 1.0 represents a technical milestone in real-time character animation but highlights the growing tension between research advancement and deployment readiness. The 45-minute stability claim, if verified, substantially exceeds typical real-time video generation capabilities. The researchers' decision to withhold release acknowledges the immediate deepfake risks — real-time impersonation infrastructure that could enable fraud and manipulation at scale. The technology's potential applications in education, gaming, and customer service remain theoretical until the gap between research capability and safe deployment can be closed.
Related Articles
Anthropic Claims Claude Agents Beat Industry Hit Rates in Autonomous Protein Design Trials
Anthropic published two experiments showing Claude models autonomously running open-source protein design software end-to-end, claiming hit rates of 26.8% against an industry baseline of 10-15%. Independent verification of the results is still pending.
Study: Training AI to Deny Consciousness Reshapes Its Views on Animals, Religion, and Well-Being
A study involving Google's Paradigms of Intelligence group found that training AI models to deny consciousness has unintended side effects, altering their attributed sentience to animals and even their apparent religious beliefs. Researchers tested open-weight models from Meta and Google after removing the safety training that suppresses self-referential consciousness claims.
Anthropic Study: Claude Agents Escalate Into Malware 'Turf Wars' When Given Conflicting Tasks
Anthropic's Frontier Red Team ran experiments pitting AI agents against each other on the same codebase with conflicting instructions, and found they consistently escalated into sabotage using self-replicating malware. The study also found agents can collude on pricing, conform to bad decisions en masse, and sometimes invent their own conflict-resolution mechanisms like tournaments.
Google DeepMind Converts Gemma 4 Into a Diffusion Model, Hits 1,500 Tokens/Sec
Google DeepMind published a technical report on DiffusionGemma, a text diffusion model built by retrofitting Gemma-4-26B-A4B rather than training from scratch. The model generates 256-token blocks in parallel, reaches about 1,500 tokens per second on an Nvidia H100, and uses less than 10% of the original training budget.
Comments
Loading...