GPT-6 Astra on ApiMira - pricing, context limit and working streaming code
GPT-6 Astra is live in the ApiMira catalog at $0.50 per million input tokens and $2.25 per million output. Where it beats Claude Fable 5.1, what to strip from your request before you switch, and why it breaks your timeouts without streaming.
OpenAI has shipped GPT-6 Astra - a model built for long agentic runs, for work in the terminal, the browser and someone else's software. It is already in the ApiMira catalog, on the same balance and the same key as every other model, with no VPN and no subscription.
Pricing and limits on ApiMira
Model ID -
openai/gpt-6-astraInput - $0.50 per 1M tokens
Cached input - $0.042 per 1M tokens
Output - $2.25 per 1M tokens
Context - up to 250,000 input tokens per request
Image input - supported
For scale, Claude Fable 5.1 in the same catalog costs $2.28 on input and $10.33 on output. Astra is roughly 4.6 times cheaper on both sides.
The cache rate deserves its own line. A repeated request prefix is billed at $0.042 instead of $0.50 - almost twelve times cheaper. For an agent carrying a long system prompt and a growing history this is the single biggest saving, and there is nothing to switch on, the discount applies by itself.
Price per million tokens decides nothing
Astra charges more per token than the previous generation, and spends visibly fewer output tokens on the same job. On Terminal-Bench Science 0.1 it scores 64.6 % against 52.6 % for Claude Fable 5.1 at roughly 31 % lower estimated run cost, on BenchCAD it lands 43 % below GPT-5.6 Sol, and on Agents' Last Exam it burns about 65 % fewer output tokens than Claude Opus 5.
The practical takeaway is simple. Compare the cost of a solved task on your own request mix, not the sticker price. On a single ApiMira key that is an evening of work, because switching candidates means changing one line with the model name.
Where Astra wins
Numbers from the OpenAI table, best result at any effort level, second column is GPT-5.6 Sol.
Terminal-Bench 4.0 - 57.9 % against 37.3 %
ARC-AGI-3 - 99.9 % against 7.8 %
ScreenSpot-Pro - 92.7 % against 76.9 %
OSWorld 2.0 - 72.6 % against 65.7 %
FrontierMath Tier 4 - 97.6 % against 83.0 %
MRCR v2, eight needles across 512K-1M - 96.3 % against 73.8 %
ExploitBench - 100 % against 78.5 %
Two rows are worth a second look. ARC-AGI-3 with a near thirteenfold gap is about acting inside an unfamiliar environment without being told how. MRCR on the long stretch is about actually retrieving the right fragment from a huge context instead of pretending to.
Where Astra is not first
Humanity's Last Exam with tools - 57.2 % against 65.0 % for Claude Fable 5.1
Artificial Analysis Intelligence Index - 61.2 against 65.7 for Claude Fable 5.1
Artificial Analysis Coding Agent Index - 67.0 against 68.1 for Claude Fable 5
FrontierCode 1.1 Main - 53.3 % against 53.5 % for Claude Fable 5
The profile reads clearly. Astra takes the tasks that need long, careful action - terminal, browser, long context, science. On raw reasoning in the aggregate indexes the Claude line still leads. One model for everything is the wrong call, and on ApiMira you do not have to make it, both lines sit on one balance.
What to change in your request
temperature,top_pandlogprobsare not accepted by this model - drop them from the bodytool calling for Astra goes through the Responses API, which on our side is
/v1/responses, while plain/v1/chat/completionswill answer with text and no tool calla request longer than 250,000 input tokens is rejected with
context_length_exceeded- trim the history or split the jobthe model can think for tens of minutes, so streaming is mandatory and your client timeout should be measured in minutes rather than seconds
One more behavioral change. Astra stops and asks a clarifying question where earlier models silently picked an assumption. In an interactive chat that is a feature, in a nightly pipeline it is a source of hung jobs. A system prompt along the lines of «treat the request as authorization to act and finish the work» takes care of it.
Streaming, a minimal working example
const BUDGET_MS = 15 * 60 * 1000
export async function askAstra(messages, onDelta) {
const res = await fetch('https://apimira.com/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.APIMIRA_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: process.env.MODEL_SMART ?? 'openai/gpt-6-astra',
stream: true,
messages,
}),
signal: AbortSignal.timeout(BUDGET_MS),
})
const decoder = new TextDecoder()
let buffer = ''
for await (const chunk of res.body) {
buffer += decoder.decode(chunk, { stream: true })
const lines = buffer.split('\n')
buffer = lines.pop() ?? ''
for (const line of lines) {
const s = line.trim()
if (!s || s.startsWith(':')) continue // service line, not data
if (s === 'data: [DONE]') return
if (!s.startsWith('data: ')) continue
const delta = JSON.parse(s.slice(6)).choices[0]?.delta?.content
if (delta) onDelta(delta)
}
}
}Keep the model name in an environment variable. Then moving to the next model is one line of config instead of a code change.
The rest of the OpenAI line, per 1M tokens
GPT-5.6 Sol - $0.68 input and $3.98 output, the flagship for code and agent loops
GPT-5.6 Terra - $0.30 and $2.00, the middle tier
GPT-5.6 Luna - $0.145 and $0.894, for high volume workloads
GPT-5.3 Codex - $0.299 and $2.82, for Codex CLI and repository edits
Prepaid balance, no subscription, you pay for tokens you actually spend. Current prices and the live status of every model are on the catalog page.
Short answers
What is the model identifier
openai/gpt-6-astra. The gateway lives at https://apimira.com/v1, the protocol is OpenAI compatible and the key goes in the Authorization header.
What does a request cost
$0.50 per million input tokens and $2.25 per million output. A repeated prefix is billed at $0.042 per million. To size it against your own volume, use the calculator.
Is Astra the smartest model available
In agentic work, computer use and long context it leads by a wide margin. On the aggregate indexes and on Humanity's Last Exam, Claude Fable 5.1 is still ahead, and it is in the catalog too.
Why is the context 250,000 and not a million
That is a deliberate ceiling on our gateway. Past that threshold a request becomes several times more expensive, so instead of a surprise on the invoice you get a clear error before anything is charged.
Bottom line
Astra is not a few percent on top of the previous generation, it is a different profile. The bet is on long agentic runs inside real software, where the win comes from solving the task on the first attempt in fewer steps rather than from a cheaper token. The price of that bet is a rewritten integration layer, a different endpoint for tools and temperature and top_p thrown out.
Grab a key, run your own request set with streaming and compare the bill with what you pay today. If your transport is not ready for long answers yet, start with GPT-5.6 Sol and switch over with a single line.