model release

OpenRouter Releases Elephant Alpha: 100B-Parameter Model with 256K Context Window and Free Pricing

TL;DR

OpenRouter has released Elephant Alpha, a 100B-parameter text model with a 256K context window and 32K output token limit. The model is available at no cost through OpenRouter's platform, supporting function calling, structured output, and prompt caching.

2 min read
1

OpenRouter Releases Elephant Alpha: 100B-Parameter Model with 256K Context Window and Free Pricing

OpenRouter has released Elephant Alpha, a 100B-parameter text model designed for "intelligence efficiency" with a 256K context window and support for up to 32K output tokens. The model is available at $0 per million tokens for both input and output through OpenRouter's routing platform.

Technical Specifications

Elephant Alpha features:

  • 100 billion parameters
  • 262,144 (256K) token context window
  • 32,768 (32K) maximum output tokens
  • Function calling support
  • Structured output capabilities
  • Prompt caching
  • Released April 13, 2025

According to OpenRouter, the model focuses on "delivering strong reasoning performance while minimizing token usage," though specific benchmark scores have not been disclosed.

Target Use Cases

OpenRouter positions Elephant Alpha for three primary applications:

  • Code completion and debugging
  • Rapid document processing
  • Lightweight agent interactions

The model is available through OpenRouter's unified API, which routes requests across multiple providers with automatic fallbacks. OpenRouter notes that prompts and completions may be logged by the provider and used for model improvement.

Pricing and Access

The model is currently available at zero cost through OpenRouter's platform, with no charges for input or output tokens. This pricing is managed through OpenRouter's routing system, which normalizes requests and responses across providers.

The model supports OpenAI-compatible API calls and can be accessed through the OpenAI SDK as well as various third-party SDKs and frameworks.

What This Means

Elephant Alpha enters a crowded field of large language models with a distinctive positioning around "intelligence efficiency" and a notably large context window at 256K tokens. The free pricing through OpenRouter makes it accessible for experimentation, though the lack of published benchmarks makes it difficult to assess performance claims against established models. The 32K output token limit is substantially higher than many competing models, which could be useful for document generation tasks. However, the data logging policy and absence of performance metrics warrant careful evaluation for production deployments.

Related Articles

model release

Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents

Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.

model release

Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context

Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.

model release

Z.ai Confirmed as Creator of Chart-Topping 'Ox Alpha' Model, Weights Coming Wednesday

Z.ai, maker of the GLM model series, has confirmed it is behind Ox Alpha, the mysterious open-weight model that appeared anonymously on OpenRouter and topped benchmark leaderboards. The company will release the model's weights on Wednesday.

model release

Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture

Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.

Comments

Loading...