← Back to Articles Directory
AI Models • July 30, 2026 •Updated September 27, 2026 • 4 min read

Agnes 2.5 Pro Alpha: API Status and Migration

Agnes has deprecated the hosted Alpha endpoint but retains its Apache 2.0 weights; learn how to migrate to the stable API or self-host.

Co-Founder & Lead Programmer of AcceleratedLogic AI

Agnes AI has deprecated the hosted agnes-2.5-pro-alpha API. Its current documentation directs new API integrations to the stable agnes-2.5-pro model. The Alpha model weights remain available under Apache 2.0, so the hosted service and the downloadable checkpoint have different status. This guide separates those options and shows what to verify before migrating.
Checked September 27, 2026. Model names, access, and prices can change. The linked Agnes AI documentation and account dashboard are the source of truth for a live deployment.

Current status at a glance

Component What Agnes AI currently documents What that means for developers
Hosted agnes-2.5-pro-alpha Deprecated; Agnes says to migrate Avoid starting a new integration against the Alpha model ID.
Hosted agnes-2.5-pro Commercial stable API model Use the stable ID and confirm account access and current billing.
Alpha checkpoint Open weights under Apache 2.0 Self-hosting remains an option, subject to the license and substantial hardware needs.
The Alpha API notice states that deprecation applies to its hosted API and does not affect the open-source weights. The stable Agnes 2.5 Pro documentation identifies agnes-2.5-pro as the commercial model for new integrations. Treat an API model ID, a hosted service, and a downloadable checkpoint as separate things when planning an upgrade.

What the Alpha checkpoint is

Agnes AI's official model card describes Agnes 2.5 Pro Alpha as a post-trained derivative of Qwen3.5-397B-A17B and publishes the weights under Apache 2.0. The card lists 397 billion parameters in BF16, text and image input, text output, a 1,048,576-token context window, and up to 65,536 output tokens. These are publisher-listed model specifications; they do not guarantee the same limits or performance from every runtime.
The model card's local-serving guidance is a useful reality check: its SGLang example uses eight-way tensor parallelism and recommends eight NVIDIA H200 GPUs or equivalent, plus about 1 TB of free fast storage for the model files and cache. A permissive weight license does not make a 397B BF16 checkpoint a practical single-GPU download. Review the card's current requirements before budgeting a self-hosted deployment.

Move a hosted integration to the stable model

The stable model documentation lists the same OpenAI-compatible Chat Completions route, https://apihub.agnes-ai.com/v1/chat/completions, with agnes-2.5-pro as the model name. A minimal request looks like this:
curl https://apihub.agnes-ai.com/v1/chat/completions \
  -H "Authorization: Bearer $AGNES_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-2.5-pro",
    "messages": [{"role": "user", "content": "Summarize this incident report."}],
    "max_tokens": 1200
  }'
This follows the official stable-model request reference. Keep the key in a server-side environment variable or secret store. Do not paste it into a public client bundle, repository, or shared prompt. Confirm that the account has access to the stable model before switching production traffic.
The stable documentation describes text and image-URL input, streaming, and tool calling. A model-name change can still affect output quality, token use, and tool behavior. Before rollout, send a representative test set through both paths if the old endpoint remains available to your account; otherwise, record a baseline from your existing outputs. Include image requests, long inputs, streamed responses, and tool-call handling if your application uses them. Check for incomplete responses and errors rather than judging only a successful short text prompt.

Recheck pricing before estimating cost

On the date checked above, Agnes AI's official pricing page listed stable agnes-2.5-pro at $0.45 per million input tokens, $0.90 per million output tokens, and $0.045 per million cached-input tokens. The page says the cached price applies only to input the service confirms as a cache hit. These are current published API prices for the stable model; they should not be presented as a continuing price for the deprecated Alpha API. Check the pricing page and account billing before forecasting spend, because rates and access can change.
Estimate a workload from measured input and output token counts, then track confirmed cache hits separately. Include retries, long reasoning outputs, and peak traffic in the estimate. A low per-token rate alone does not tell you whether a model meets your latency, reliability, or quality requirements.

Choosing between the API and open weights

Use the stable hosted API when you want Agnes AI to operate the serving infrastructure and your account, data-handling requirements, budget, and region permit that arrangement. Choose the Alpha weights only if you need to manage the deployment yourself and can meet the model card's hardware requirements. For either route, pin the model identifier in configuration, test the exact endpoint or checkpoint you will ship, and keep a rollback path. The Alpha benchmark table is a historical publisher-hosted reference; it is not a current comparison of the stable Pro API, so this guide does not use it to rank present-day models.

Official sources