← Back to Payloads
AI News2026-09-29

OpenAI Pushed Three Things on Sep 29: GPT-6.1 Sol Lands at $2/$10 with Half the Cache Reads, Astra Gets an Ultrafast Mode, and the Agents API Picks Up Computer Use

Three OpenAI changelog entries landed on Sep 29: GPT-6.1 Sol at $2/$10 with cached input cut from $0.20 to $0.10/MTok (50% off) plus Multi-agent beta; GPT-6 Astra Ultrafast mode (US-only residency, rate-limited); Computer use added to the Agents API with OpenAI-hosted browser handoff.

OpenAI Pushed Three Things on Sep 29: GPT-6.1 Sol Lands at $2/$10 with Half the Cache Reads, Astra Gets an Ultrafast Mode, and the Agents API Picks Up Computer Use

Originally published: 2026-09-29 20:08 UTC / 22:08 Berlin / 16:08 EDT

What Happened

OpenAI's API changelog carries three new Sep 29, 2026 entries dated today:

1. GPT-6.1 Sol (gpt-6.1-sol) is released for complex coding and professional work at a lower cost than GPT-6 Astra. Standard pricing for prompts up to 272K input tokens is $2 input / $0.10 cached input / $2.50 cache write / $10 output per 1M tokens. Multi-agent delegation is supported in beta on the Responses API. 2. GPT-6 Astra Ultrafast mode is added to the Responses API. Use gpt-6-astra with service_tier: "ultrafast" to reduce the time between generated output tokens. Available to API customers, subject to rate limits, with global processing and US data residency. EU and other regional inference residency aren't supported. 3. Computer use is added to the Agents API. Agents can complete tasks in an OpenAI-hosted browser, with website access approvals and sign-in handled by your application.

This report is a documentation comparison against the OpenAI changelog and the linked model/guide pages. No firsthand API call was made.

What Actually Changed

GPT-6.1 Sol pricing math. Same base list price as GPT-6 Sol ($2 input / $10 output per 1M tokens, <=272K context), but the cached-input line drops from $0.20/MTok to $0.10/MTok - a 50% cut on cache reads. Cache-write stays at $2.50/MTok and Batch/Flex pricing scales linearly. On long-context inputs (>272K), the cached-input rate is $0.20/MTok versus GPT-6 Sol's $0.40/MTok. The model is positioned in the changelog as "complex coding and professional work at a lower cost than GPT-6 Astra" - the same slot GPT-6 Sol previously held, but with a sharper cache story.

Multi-agent in beta. GPT-6.1 Sol supports Multi-agent delegation on the Responses API. The changelog sentence: "Let the model delegate work to subagents in a Responses API request." This is the same multi-agent primitive already documented for the Responses API; the model-side flag is what changes.

Astra Ultrafast scope. The new service_tier: "ultrafast" only works with gpt-6-astra (not GPT-6.1 Sol, not GPT-6 Sol, not GPT-6 Luna). Residency is explicitly US-only - the changelog calls this out: "EU and other regional inference residency aren't supported." The pricing page exposes a separate Ultrafast pricing section.

Agents API computer use. Computer use moves out of the standalone model-tool path and into the Agents API tool set. The Agents API, first surfaced as a public beta on Sep 10 (already covered separately), gains a managed browser surface: OpenAI hosts the browser, your application owns website-access approvals and sign-in. This is a cleaner integration path than the legacy computer_use_preview model endpoint.

Why Developers and Founders Should Care

GPT-6.1 Sol's cache-read cut is the headline for cost. If you run heavy prompt-cache workloads on GPT-6 Sol (agent backbones, retrieval-augmented pipelines, code-search sweeps), a 50% cache-read cut is a real margin change. Quick math: a workflow spending 100M cached input tokens/day drops from ~$20/day (at GPT-6 Sol's $0.20/MTok) to ~$10/day (at GPT-6.1 Sol's $0.10/MTok) for the same work. The output and cache-write costs are unchanged, so the cut is a pure cache-read win - workloads with low cache hit ratios won't feel it.

Astra Ultrafast is the headline for latency-critical UX. The changelog defines it as reducing the time between generated output tokens - i.e., inter-token latency, not time-to-first-token. Streaming UX for a coding assistant or an interactive agent is where this shows up. The catch is residency: if you have an EU data-residency requirement, you cannot use Ultrafast mode. US-only deployments and global-with-US-residency workloads are fine.

Agents API computer use is the headline for capability reach. Building a browser-using agent used to mean wiring up the legacy computer-use tool against a model you hosted yourself. Now the Agents API can hand off to an OpenAI-hosted browser session and let your application gate website access and sign-in. This is the path the Sep 10 Agents API public-beta post hinted at: managed harness, durable sessions, host-owned browser.

Evidence and Test Results

All claims above are quoted or paraphrased from the OpenAI changelog Sep 29 entries:

  • GPT-6.1 Sol release entry: "Released GPT-6.1 Sol (gpt-6.1-sol) for complex coding and professional work at a lower cost than GPT-6 Astra. Standard pricing per 1M tokens for prompts with up to 272K input tokens is $2 input, $0.10 cached input, $2.50 cache write, and $10 output. GPT-6.1 Sol also supports Multi-agent in beta."
  • Ultrafast entry: "Added Ultrafast mode for GPT-6 Astra in the Responses API. Use gpt-6-astra with service_tier: \"ultrafast\" to reduce the time between generated output tokens. It is available to API customers, subject to rate limits, with global processing and US data residency. EU and other regional inference residency aren't supported."
  • Agents API computer-use entry: "Added computer use to the Agents API. Agents can complete tasks in an OpenAI-hosted browser, with website access approvals and sign-in handled by your application."

Pricing verified against the live developers.openai.com/api/docs/pricing page at 2026-09-29 20:08 UTC. The Standard table lists gpt-6-astra $10 / $1 cached / $12.50 cache write / $50; gpt-6.1-sol $2 / $0.10 cached / $2.50 cache write / $10; gpt-6-sol $2 / $0.20 cached / $2.50 cache write / $10; gpt-6-luna $0.10 / $0.01 cached / $0.125 cache write / $0.50. Long-context columns show gpt-6.1-sol at $4 / $0.20 cached / $5 cache write / $15.

Verification level: documentation comparison + pricing-page snapshot. No firsthand API call was made for this report. We did not test cache-hit-ratio behavior, inter-token latency under Ultrafast, or browser-use session handoff end-to-end.

Cost, Risk, and Limitations

  • GPT-6.1 Sol is a price-positioning move, not a capability upgrade claim. The changelog says "for complex coding and professional work at a lower cost than GPT-6 Astra" but doesn't promise a capability uplift over GPT-6 Sol. Treat it as a refresh of the same slot with a sharper cache price; benchmark on your workload before mass-migrating.
  • Cache-read cut is conditional on cache hit. The 50% cache-read cut applies per cached token, not per request. Workloads with low cache hit ratios won't see material cost reduction; high-cache workloads (agent backbones, RAG with shared system prompts, code-search) will.
  • Astra Ultrafast residency is US-only. Any EU/UK/APAC residency requirement disqualifies Ultrafast today. The changelog is explicit: "EU and other regional inference residency aren't supported."
  • Astra Ultrafast is rate-limited. The changelog flags "subject to rate limits" without quantifying them. Don't promise a latency-tier SLA to end users without checking the published rate-limit page for your tier.
  • Astra Ultrafast is per-model. Only gpt-6-astra supports service_tier: "ultrafast". GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna do not. If you want the latency tier on a coding workload, you have to route through Astra - which is the more expensive model.
  • Agents API computer use is in the public-beta Agents API, not in the general Responses API. If you're already on the Responses API directly (not the Agents API), this change does not affect your stack. If you're evaluating the Agents API for the first time, this is a reasonable entry point but you inherit the beta contract.
  • Multi-agent delegation in beta. The changelog describes Multi-agent delegation as supported "in beta" on GPT-6.1 Sol. Expect API-shape churn; pin a versioned API header.

Mr. Technology Verdict

Three additive changes, no breaking-change headline. The Sep 29 changelog trifecta is a steady-state "ship the backlog" day: a refresh-tier model with a sharper cache price, a latency-tier flag with explicit residency limits, and a managed-browser handoff for the Agents API. None of them are a "stop the press" event. All three are worth adopting on your next normal maintenance window if you have workloads that fit their slots.

If you only have time to do one of these, do GPT-6.1 Sol. The cache-read cut is the most material cost change for the average production agent workload, and the migration path is a one-line model-ID swap. The Ultrafast and Agents-API-computer-use stories are conditional wins - Ultrafast only applies if you're residency-flexible and on Astra, and Agents API computer use only applies if you've already adopted the Agents API.

Recommended Action

1. Today: Pull your prompt-cache hit-rate metric for the last 30 days on GPT-6 Sol. If your cache hit ratio is >60% on long-context workloads, GPT-6.1 Sol is a material cost win - swap the model ID and re-run your regression suite. 2. This week: If you have a streaming UX where inter-token latency matters (interactive coding assistant, IDE plugins, real-time agent chat), test gpt-6-astra with service_tier: "ultrafast" against your latency budget. Confirm your residency posture before promising the latency tier to EU users. 3. This week: If you're evaluating or already on the Agents API, add computer use to your tool inventory and document the website-access approval and sign-in flow you want your application to enforce. 4. Skip if not in scope: If you're cache-light, residency-bound to EU, or not using the Agents API, none of these three changes affect your stack today. Monitor; do not migrate.

Sources

Last verified: 2026-09-29 20:08 UTC. No corrections on file.

Related Dispatches