model releaseStability AI

Stability AI releases Stable Audio Open Small for on-device audio generation with Arm

TL;DR

Stability AI has open-sourced Stable Audio Open Small in partnership with Arm, a smaller and faster variant of its text-to-audio model designed for on-device deployment. The model maintains output quality and prompt adherence while reducing computational requirements for real-world edge deployment on devices powered by Arm's technology, which runs on 99% of smartphones globally.

2 min read
0

Stability AI Releases Stable Audio Open Small for On-Device Deployment

Stability AI has open-sourced Stable Audio Open Small, a compact audio generation model developed in partnership with Arm, targeting real-world deployment on edge devices and smartphones.

Key Details

Stable Audio Open Small represents a size-optimized variant of Stability AI's existing Stable Audio Open text-to-audio model. The partnership leverages Arm's processor architecture, which according to Stability AI powers 99% of smartphones globally.

The model maintains the core capabilities of the full Stable Audio Open while reducing model size and inference latency. Stability AI claims the variant preserves output quality and maintains adherence to text prompts despite the architectural compression.

No specific model parameters, file size, inference speed benchmarks, or inference cost metrics have been disclosed at launch. The pricing structure for commercial use remains undisclosed.

Technical Approach

The release focuses on addressing a key constraint in audio generation: computational efficiency on resource-limited devices. By optimizing Stable Audio Open for Arm processors, the model can theoretically run directly on smartphones and edge devices without requiring cloud inference.

The open-source release suggests this is not a proprietary commercial variant. Developers can integrate the model into applications that require on-device audio generation without relying on external API calls.

Market Context

Audio generation remains an underdeveloped area compared to text and image generation. Most commercial audio models (including OpenAI's text-to-speech API and ElevenLabs) require cloud-based inference. On-device audio generation has seen limited adoption outside specialized applications.

Stability AI's focus on open-sourcing the model aligns with its broader strategy of releasing foundation models freely to developers, contrasting with the closed commercial approaches of competitors like OpenAI and Anthropic.

What This Means

Stable Audio Open Small enables developers to build audio generation features directly into mobile applications without external dependencies. This reduces latency, improves privacy (processing stays on-device), and removes per-inference costs. However, the practical impact depends on undisclosed details: actual model size, latency benchmarks, audio quality metrics, and real-world performance on typical smartphone hardware. The 99% Arm smartphone penetration claim suggests broad potential reach, but device-level variable performance may limit deployment across older or lower-end devices. Whether this model achieves production-grade audio quality comparable to cloud-based systems remains unconfirmed.

Related Articles

model release

Xiaomi Releases MiMo-V2.6-Flash-RL, a 309B-Parameter MoE Model with 1M-Token Context and Native Omnimodal Support

Xiaomi's MiMo team released MiMo-V2.6-Flash-RL, an efficiency-tier checkpoint in the MiMo-V2.6 series featuring a 309B-parameter (15B active) Mixture-of-Experts architecture, 1M-token context, and native support for text, image, video, and audio. The model uses a single mixed reinforcement learning run across coding, agentic, visual, and cybersecurity tasks rather than domain-specific training.

model release

Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens

Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.

model release

Yandex Releases AliceAI-Foundation-80B-A3B-Base, an 80B-Parameter MoE Model with 262K Context

Yandex has released AliceAI-Foundation-80B-A3B-Base, an 80-billion-parameter hybrid MoE base model with 3 billion active parameters per token and a 262,144-token context window. The model was trained fully from scratch and, according to Yandex, outperforms larger open-source models on Russian-language factual and educational benchmarks.

model release

Alibaba Releases Qwen-Image-2.1, a 7B Unified Text-to-Image and Editing Model

Alibaba's Qwen team has open-sourced Qwen-Image-2.1, a 7B parameter unified model for text-to-image generation and image editing. The release adds native transparent (RGBA) image support and editing with up to 10 reference images.

Comments

Loading...