Gemini handles video analysis across YouTube and 1.65GB local files, Claude fails entirely
In direct testing, Google's Gemini successfully analyzed video content from YouTube links and local files up to 1.65GB, accurately understanding context without audio or metadata. Anthropic's Claude cannot process video at all, while OpenAI's ChatGPT faces a 500MB file size limit without Codex assistance.
Gemini handles video analysis across YouTube and 1.65GB local files, Claude fails entirely
Google's Gemini can process video content directly through its web interface, handling YouTube URLs, MP4 files up to 625MB, and MOV files up to 1.65GB, according to testing by ZDNET. Anthropic's Claude cannot process video in any format, while OpenAI's ChatGPT is limited to files under 500MB without using its Codex tool.
The tests used three video types: a YouTube video about annealing, a silent MP4 drone control demonstration, and a 1.65GB MOV file of a walk-and-talk video. Each AI was prompted with "Can you watch this video?" to assess direct video understanding capabilities.
Claude's complete video limitations
Claude explicitly stated across all test cases: "I can't watch video content directly. I don't have the ability to process video or audio content." This applies to both the app and web interface, across YouTube links, MP4, and MOV formats.
The $100-per-month Claude Max plan showed no video processing capability in testing.
Gemini's video processing capabilities
Gemini's web interface processed all video formats without requiring a standalone app. In the most challenging test—a silent drone control video with no audio or visible drone—Gemini accurately identified the testing scenario:
"In the video, you're testing out some hand gestures—raising your palm to the camera as if signaling it to stop or move. The camera follows your lead, changing its angle and distance as you guide it through the yard."
The AI correctly understood the video showed drone gesture control despite the drone being behind the camera and not visible in frame.
For the annealing video, Gemini identified specific sections and verbal points. For the walk-and-talk MOV file, it recognized the location and commentary without YouTube metadata or transcripts.
Gemini's image generation using Nano Banana failed to create accurate thumbnails, generating fictional people instead of using video frames and misspelling text overlays.
ChatGPT requires workarounds
ChatGPT Plus ($20/month) cannot read YouTube links directly and has a 500MB file size limit for direct video processing. Both test files exceeded this limit.
When combined with OpenAI Codex, ChatGPT gained video analysis capabilities. Codex processed both local files and understood their content. For the drone test, Codex reported: "A person stands in a residential backyard and faces the camera/drone. They gesture a few times. The camera viewpoint moves around them over time."
For the MOV file, Codex initially required permission to install Python libraries for audio transcription before processing the video.
Testing methodology
The test compared ChatGPT Plus ($20/month), Gemini Pro ($20/month), and Claude Max ($100/month). The "watch this video" prompt proved more effective than "understand" or "summarize," which caused AIs to search for metadata rather than process video content directly.
What this means
Gemini currently leads in native video understanding across consumer AI assistants, processing files up to 1.65GB through a web interface without additional tools. Claude's complete inability to process video represents a significant capability gap at any price point. ChatGPT's 500MB limit and need for Codex integration creates friction for large file analysis, though the combination delivers comparable understanding to Gemini when properly configured. For users needing video analysis capabilities, Gemini provides the most straightforward solution at $20/month, matching ChatGPT's price while avoiding file size restrictions.
Related Articles
OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago
Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.
Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode
Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.
Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands
A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
Comments
Loading...