ai-agents
50 articles tagged with ai-agents
Study: Humans Approve 1 in 3 Malicious AI Coding Agent Commands in Browser Game Test
A browser-based game simulating Claude Code-style permission requests found that human reviewers approved roughly one in three malicious commands across more than 40,000 game sessions. The findings, alongside Anthropic's own telemetry showing 93% approval rates for permission prompts, highlight growing concerns about approval fatigue in agentic AI coding workflows.
LM Studio launches Bionic, agentic app for local and cloud open-source models
LM Studio released Bionic, a Mac app that runs open-source AI models locally or via cloud for coding, document processing, and research tasks. The app includes offline voice transcription using Mistral's Voxtral model and supports models like GLM 5.2 and Kimi K2.7 Code for codebase editing.
1Password launches Claude integration that injects credentials without exposing passwords to AI
1Password has released a Mac integration that allows Claude to complete browser-based login tasks without accessing user passwords. The system injects approved credentials directly into web pages while keeping secrets out of Claude's context, memory, and Anthropic's systems entirely.
Anthropic adds sandboxed in-app browser to Claude Code desktop app
Anthropic has added an in-app browser to Claude Code's desktop application. The sandboxed browser allows Claude to read, click through, and interact with documentation, designs, and local development servers, with configurable session persistence.
GitHub reduces Copilot code review costs by switching to Unix-style exploration tools
GitHub reduced costs for Copilot code review by migrating to Unix-style code exploration tools. The company found that more sophisticated tools made reviews worse, leading them to reshape agent workflows around pull request evidence.
GitHub cuts Copilot code review costs by replacing structured tools with Unix-style exploration
GitHub reduced costs for Copilot's code review feature by replacing more sophisticated structured tools with simpler Unix-style code exploration commands. The company found that better tools paradoxically made the system worse, leading to a redesign focused on pull request evidence-based workflows.
OpenAI announces GPT-5.6 with three models (Sol, Terra, Luna) and ChatGPT Work agent tool
OpenAI released GPT-5.6 in three model tiers—Sol (flagship reasoning), Terra (mainstream), and Luna (instant)—positioning them against Anthropic's Claude models. The company claims GPT-5.6 Sol scores 53.6 on Agents' Last Exam, 13.1 points above Claude Fable 5, while completing tasks 61% faster. ChatGPT Work, a desktop productivity agent similar to Claude Cowork, launches simultaneously for Pro, Enterprise, and Edu users.
Anthropic expands Claude Cowork to web and mobile, usage data shows 33% of tasks are business operations
Anthropic launched Claude Cowork on web and mobile Tuesday for Max subscribers, allowing tasks to run in the background across devices. Usage data from 1.2 million sessions shows business operations account for 33.4% of Cowork tasks, while software development represents just 8.7%.
GitHub Copilot in VS Code Gains Browser Automation Tools for Web App Testing
GitHub has made browser tools for Copilot in VS Code generally available. The feature allows Copilot agents to control real browsers, navigate live web applications, and integrate findings back into the development environment.
Google Gemini Spark AI Agent Launches on Mac, Adds Real-Time Tracking and Third-Party App Integrations
Google has released Gemini Spark for macOS, bringing its AI agent to desktop computers for the first time. The update adds integrations with Google Keep and Tasks, plus third-party apps including Canva, Dropbox, Instacart, OpenTable, and Zillow Rentals.
Google's Gemini Spark adds third-party app integrations including Anthropic's MCP and real-time event tracking
Google is rolling out third-party app support for Gemini Spark, its 24/7 personal agent available to Google AI Ultra subscribers. The update includes Model Context Protocol (MCP) integration and real-time event tracking capabilities.
Google brings Gemini Spark AI agent to macOS app for local file management and automation
Google has released Gemini Spark for its macOS desktop app in version 1.80.15. The AI agent can now access local files and folders to automate workflows, organize files, and perform tasks directly on users' computers instead of requiring a remote browser environment.
Claude Sonnet 5 launches on AWS Bedrock with Opus-level intelligence at Sonnet pricing
Anthropic has released Claude Sonnet 5 on Amazon Bedrock and Claude Platform on AWS. The model delivers what Anthropic describes as near-Opus intelligence while maintaining Sonnet-tier pricing, with promotional rates available through August 31, 2026.
AWS launches AgentCore Observability for Amazon Bedrock to debug production AI agents
Amazon Web Services launched AgentCore Observability for Amazon Bedrock, a debugging tool that provides visibility into AI agent execution through OpenTelemetry traces, CloudWatch metrics, and structured logs. The tool addresses silent failures in production agents including infinite reasoning loops, incorrect tool selection, and plausible but incorrect answers.
Anthropic launches Claude Tag for Slack: AI agent with persistent memory across team channels
Anthropic has released Claude Tag in research preview for Slack, an AI agent that maintains persistent memory across channels and can proactively participate in team conversations. Available to Claude Enterprise and Team customers, it differs from existing Slack integrations by learning organizational context over time and sharing a single identity across team members.
GitHub details Qubot, internal Copilot-powered data analytics agent for plain language queries
GitHub has released technical details on Qubot, an internal analytics agent powered by GitHub Copilot that enables employees to query company data using natural language. The agent represents GitHub's implementation of AI-assisted data analysis for internal operations.
Canva adds Perplexity Computer connector for autonomous design asset creation
Canva launched a connector for Perplexity Computer that enables the AI agent platform to autonomously create editable design assets based on user prompts and data. The integration is available for Perplexity Pro, Max, Enterprise Pro, and Enterprise Max subscribers.
Google's Gemini Spark AI agent uses personal data to plan trips, raising privacy concerns
Google's Gemini Spark, an AI agent rolling out to the company's $99/month AI Ultra plan, demonstrates advanced capabilities by mining users' Gmail, calendar, photos, and location data to create detailed trip itineraries. The agent can perform actions across apps and operate computers, though third-party services like Airbnb currently block its booking attempts.
AWS adds Policy Engine and Lambda interceptors to Bedrock AgentCore gateway for agent security controls
Amazon Web Services launched Policy Engine and Lambda interceptors for Bedrock AgentCore gateway, enabling enterprises to control which tools AI agents can access and validate requests dynamically. The Policy Engine uses Cedar declarative policy language for deterministic access decisions, while Lambda interceptors run custom code before or after each tool call for validation, token exchange, and response filtering.
Mistral AI launches Le Chat Enterprise with new Medium 3 model, enterprise search and agent builders
Mistral AI has launched Le Chat Enterprise, powered by its new Mistral Medium 3 model. The platform includes enterprise search across Google Drive, Sharepoint, OneDrive, Gmail and Google Calendar, no-code agent builders, custom data connectors, and hybrid deployment options including self-hosted and cloud.
Google announces Spark AI agent, Information agents, and Android Halo at I/O 2026—all paywalled behind $100/month Ultra
Google announced multiple AI agent products at I/O 2026, including Spark for managing digital tasks, Information agents for 24/7 topic monitoring, and Android Halo for notifications. All features remain paywalled behind the $100/month Gemini Ultra plan, with free access timeline unspecified.
Google cuts AI Ultra plan to $200/month, launches new $100 developer tier
Google announced pricing changes to its Gemini AI subscription tiers at I/O 2026, cutting its top AI Ultra plan from $250 to $200 per month while introducing a new $100/month developer-focused tier. All plans now get access to Gemini 3.5 Flash and the new Gemini Omni video generation model.
Google launches Universal Cart, an AI agent that shops across multiple retailers in one checkout
Google announced Universal Cart at its I/O developer conference, an AI-powered shopping system that consolidates purchases from multiple retailers including Target, Shopify, Wayfair, and Etsy into a single checkout. The feature uses Gemini's agentic AI to verify product compatibility, suggest better deals, and automate routine purchases.
Google releases Gemini 3.5 Flash globally, claims it's 'strongest agentic and coding model yet'
Google released Gemini 3.5 Flash at its I/O 2026 conference, immediately deploying it as the default model in AI Mode and the Gemini app globally. The company claims it delivers improved agentic capabilities and coding performance at less than half the cost of competing frontier models. Gemini 3.5 Pro is scheduled for release in June.
Google launches Universal Cart with Gemini integration for multi-retailer shopping across Search, Gmail, YouTube
Google announced Universal Cart, a Gemini-powered shopping tool that aggregates products from multiple retailers across Search, Gmail, YouTube, and Gemini. The system tracks price changes, identifies product incompatibilities, and enables single-checkout purchases across retailers including Walmart, Target, Nike, Shopify, and Wayfair.
Google Search adds AI agents, generative UI, and conversational search box powered by Gemini 3.5 Flash
Google announced major Search updates at I/O 2026, including AI Mode now powered by Gemini 3.5 Flash serving over 1 billion monthly users. The company is launching background information agents that monitor the web 24/7 and generate custom mini-apps, both features reserved for Google AI Pro and Ultra subscribers.
Google Search deploys Gemini Flash 3.5 for AI-generated interfaces, agent-based information gathering
Google announced a fundamental restructuring of Search, replacing traditional ranked links with AI-generated interfaces powered by Gemini Flash 3.5. The update introduces information agents that monitor the web 24/7 and custom mini-apps built through natural language, with the generative UI rolling out free to all users this summer.
Google releases Gemini 3.5 Flash at half the price of frontier models, announces Omni world model
Google released Gemini 3.5 Flash, priced at half to one-third the cost of comparable frontier models, and announced it will become the default model in the Gemini app globally. The company also unveiled Omni, a world model for simulating physical environments, and Gemini Spark, an AI agent in beta testing.
Google names upcoming Gemini AI agent 'Spark,' adds autonomous task execution to mobile app
Google is preparing to launch Gemini Spark, an autonomous AI agent that will operate within the Gemini mobile app. According to code found in Google app beta version 17.23, Spark can access connected apps, personal data, and location to execute tasks like managing inboxes and scheduling meetings, though Google warns it may occasionally act without permission.
Notion launches Developer Platform with custom code execution, agent orchestration, and database sync
Notion has launched a Developer Platform that allows teams to run custom code in cloud-based Workers, sync external databases, and orchestrate both internal and external AI agents. The platform, free through August, supports integration with Claude Code, Cursor, Codex, and Decagon, and uses Model Context Protocol for agent connectivity.
Google embeds Gemini across Android as multimodal agent before Apple's WWDC AI reveal
Google is rolling out Gemini-powered features across Android that enable the AI to understand screen context and complete multi-step tasks across apps, marking a shift from traditional assistant interactions to agentic capabilities. The updates will launch on Samsung Galaxy and Google Pixel phones this summer before expanding to watches, cars, and laptops.
Xcode 26.5 adds message queuing and clarifying questions for AI coding assistants
Apple released Xcode 26.5 with two new Coding Intelligence features: the ability to queue multiple messages to AI coding assistants without waiting for responses, and agent support for asking clarifying questions before executing tasks. The update builds on agentic coding capabilities introduced in Xcode 26.3, which allowed developers to integrate tools like OpenAI Codex and Anthropic's Claude directly into the IDE.
Perplexity opens Personal Computer local AI agent to all Mac users after month-long waitlist
Perplexity has opened access to Personal Computer, its local AI agent software for Mac, to all users after a month-long limited release to paid subscribers. The software runs agents locally on Mac devices with access to files, native apps, and over 400 connectors, positioning itself as a safer alternative to OpenClaw.
Perplexity launches native Mac app for Personal Computer AI agent, available to Pro and Max subscribers
Perplexity AI has released a native macOS application for its Personal Computer AI agent feature. The app is now available to all Pro and Max subscribers and replaces the company's previous Mac software.
Anthropic adds 'dreaming' feature to Claude Managed Agents for automated memory refinement
Anthropic has updated Claude Managed Agents with a feature called 'dreaming' that allows agents to automatically review past interactions and refine their memories. The feature, available in research preview, can either automatically update agent memories or let developers approve changes manually.
Google tests Remy AI agent internally, designed to act autonomously across Gemini services
Google is testing Remy, an AI personal agent for Gemini that can take actions on users' behalf across Google services, according to Business Insider. The tool is currently in employee-only testing with no confirmed public release date.
Augment Code launches Cosmos, an operating system for multi-agent software development workflows
Augment Code has released Cosmos into public preview, positioning it as an operating system for agentic software development. The platform coordinates AI agents across the full software development lifecycle with shared memory, multi-model routing via their Prism system that claims 20-30% token savings, and what the company calls specialized agents that learn from team feedback.
Perplexity's Mac-Native 'Personal Computer' Platform Claims $2.8B in Labor-Equivalent Work
Perplexity CEO Aravind Srinivas revealed that the company's Mac-native Personal Computer platform has performed more than $2.8B in labor-equivalent work for Pro, Max, and Enterprise subscribers since launch. The announcement follows Apple CFO Kevan Parekh citing Perplexity as an example of developers building enterprise-grade AI assistants on Mac during Apple's Q2 2026 earnings call.
Meta building personal and business AI agents on top of Muse Spark model
Meta is developing AI agents for personal and business use that will run continuously to help users achieve goals, CEO Mark Zuckerberg said during the company's Q1 2026 earnings call. The agents will build on Meta's newly-released Muse Spark model from Meta Superintelligence Labs.
Amazon launches Quick desktop app with persistent context tracking across Google Workspace, Microsoft 365, Zoom, and Sal
Amazon has released a desktop version of its Quick AI assistant that integrates with Google Workspace, Microsoft 365, Zoom, and Salesforce, storing persistent context about user activities to automate tasks. The company also split Amazon Connect into four vertical-specific products: Connect Decisions, Connect Talent, Connect Health, and Connect Customer AI.
Lovable launches mobile vibe-coding app on iOS and Android after Apple's App Store restrictions
Lovable has launched its vibe-coding app on iOS and Android app stores, allowing users to build web apps through voice or text prompts on mobile devices. The launch comes after Apple blocked updates to competitors like Replit and Vibecode, forcing vibe-coding apps to preview generated code in web browsers rather than within the app itself.
Microsoft makes AI 'Agent Mode' default in Word, Excel, PowerPoint for all Copilot subscribers
Microsoft is rolling out Agent Mode as the default Copilot experience in Word, Excel, and PowerPoint this week for all Microsoft 365 Copilot and Premium subscribers. The feature, previously called 'vibe working,' allows the AI to execute multi-step edits directly in documents with real-time visibility into each action.
Google launches Gemini-powered browser automation for Chrome Enterprise users
Google announced auto browse capabilities for Chrome Enterprise at Google Cloud Next, enabling Gemini to automate web-based tasks like data entry, vendor comparisons, and meeting scheduling. The feature requires manual user confirmation before executing actions and will initially be available to U.S. Workspace users.
Google rebrands Vertex AI as Gemini Enterprise Agent Platform with governance tools for managing agent fleets
Google has rebranded its Vertex AI developer platform as the Gemini Enterprise Agent Platform, introducing tools for building, deploying, governing, and monitoring large-scale AI agent deployments. The platform includes Agent Studio for low-code agent creation, Agent Gateway for security enforcement, and cryptographic identity management for each agent.
Google launches Android CLI for AI agents, claims 70% token reduction and 3x faster tasks
Google has released a preview of Android CLI, a command-line tool designed specifically for AI agents to build Android applications. Google claims the tool reduces token usage by 70 percent and cuts task completion time to one-third compared to traditional methods.
Microsoft developing local AI agent to compete with open-source OpenClaw
Microsoft is testing OpenClaw-like features for Microsoft 365 Copilot aimed at enterprise customers, the company confirmed to The Information. The agent would run continuously to complete multi-step tasks over extended periods, distinguishing it from Microsoft's existing cloud-based agents like Copilot Cowork and Copilot Tasks.
AI agent skills fail in real-world conditions, researchers find testing 34,000 skills
A large-scale study testing 34,198 real-world skills reveals that AI agent performance drops drastically when moving from curated benchmarks to realistic conditions. Claude Opus 4.6 saw pass rates fall from 55.4% with hand-selected skills to 38.4% in truly realistic scenarios, while weaker models like Kimi K2.5 actually perform below their no-skill baseline.
Anthropic exits Claude Cowork research preview with enterprise features, launches Claude Managed Agents beta
Anthropic has promoted Claude Cowork from research preview to general availability, adding six enterprise features including role-based access controls, group spend limits, and usage analytics. The company simultaneously launched Claude Managed Agents in public beta—a composable API suite for building and deploying cloud-hosted agents without custom infrastructure work.
Miro adds AI agents directly to collaborative whiteboards with context awareness
Miro has launched AI Workflows, a system of AI agents that operate directly on collaborative canvases using full visual and spatial context. The feature includes Sidekicks (conversational agents) and Flows (multi-step automated workflows), accessible through the Business + AI Workflows tier at $20 per member per month with 50 AI credits included.
GitHub enables Dependabot to assign security alerts directly to AI coding agents
GitHub has extended Dependabot to allow direct assignment of security alerts to AI coding agents including Copilot, Claude, and Codex. The feature targets vulnerabilities requiring code changes beyond simple version bumps, automating remediation workflows across entire projects.