← Back to Payloads
AI News2026-09-26

OpenAI Patched a GPT-6 Sol and GPT-6 Luna Image-Encoding Bug on Sep 25 — Rerun Vision Evaluations if You Tested Either Model Between Sep 22 and Sep 25

OpenAI's Sep 25 changelog acknowledged a bug in image encoding that degraded image understanding in both GPT-6 Sol and GPT-6 Luna between their Sep 22 launch and the Sep 25 fix. The documentation tells users to rerun evaluations and retry affected workflows. The bug also touched Codex's computer-use path. This is a documentation-surfacing news-impact report on what shipped, what the working codex evaluation pipelines need to do about it, and the limits of the evidence chain.

OpenAI Patched a GPT-6 Sol and GPT-6 Luna Image-Encoding Bug on Sep 25 — Rerun Vision Evaluations if You Tested Either Model Between Sep 22 and Sep 25

The Sep 25 entry on OpenAI's developer changelog is short and operational. It acknowledges a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna, the two reasoning models OpenAI released on Sep 22. The same entry says the fix improves results on visual tasks in the API and Codex, including computer use. The operational guidance is explicit: "If your use cases involve image inputs, we recommend rerunning your evaluations and retrying workflows affected by the issue."

This is the kind of changelog entry that does not become a tweet and does not get a marketing post. It is the kind of changelog entry that affects every team that benchmarked the new models on vision between Sep 22 and Sep 25. The rest of this article is what shipped, what is known about what changed, why it matters, and where the evidence stops.

What happened

OpenAI added a Sep 25 entry under the Fix category on the developer changelog, scoped to Model: gpt-6-sol and Model: gpt-6-luna. The body of the entry reads, verbatim from the OpenAI changelog at fetch 2026-09-26 20:08 UTC:

Fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna. This update improves results on visual tasks in the API and Codex, including computer use.

>

If your use cases involve image inputs, we recommend rerunning your evaluations and retrying workflows affected by the issue.

That is the full text of the entry. OpenAI has not (in the changelog entry itself) labeled the bug with a severity, posted an incident timeline, or published a list of which evaluation surfaces are most affected. The Sep 25 entry sits one notch above the Sep 22 launch entry on the changelog page, which means it is treated as a current-state fix rather than a historical record.

Documentation indicates this is the first public acknowledgment from OpenAI that the Sep 22 launch shipped with a known image-understanding regression. The launch had no caveats on vision that the changelog reproduced; the Sep 25 entry is the caveating record.

What actually changed

Five things changed between Sep 22 (launch) and Sep 25 (fix):

1. Image-understanding quality on GPT-6 Sol and GPT-6 Luna. OpenAI says results are now improved on visual tasks, including computer use. The implication is that a class of image inputs was returning degraded answers in the three-day window between launch and fix. The exact class — large images, multi-image prompts, screenshots-as-input, or specific color / contrast channels — is not disclosed in the changelog entry. 2. Codex's computer-use path. The Sep 25 entry names Codex explicitly: "results on visual tasks in the API and Codex, including computer use." For Codex operators using gpt-6-sol or gpt-6-luna as the underlying model on the computer-use tool, the fix changes the outputs they were seeing on Sep 22–24. 3. The recommendation status of vision-side evaluations. OpenAI's own recommendation is "rerun your evaluations and retry workflows affected by the issue." That is not a suggestion that the previous results are equivalent under statistical noise; it is a recommendation to treat the previous results as suspect. 4. The documentation record. The Sep 22 launch entry on the changelog page is the canonical record of the model release; the Sep 25 entry is an addendum that scopes a known regression to the image-encoding path. Anyone consulting the changelog for the GPT-6 Sol / GPT-6 Luna record now sees both entries. 5. The operational baseline for any future GPT-6 Sol / GPT-6 Luna comparison. Any benchmark, eval, or production comparison published between Sep 22 and Sep 25 on vision-side tasks is, per OpenAI's own documentation, suspect. Comparisons published before Sep 22 (against earlier models) are unaffected.

What did not change:

  • The model IDs (gpt-6-sol and gpt-6-luna) and the API and Codex endpoints are unchanged.
  • Pricing is unchanged: GPT-6 Sol remains $2/$0.20 cached input/$10 output per MTok, and GPT-6 Luna remains $0.10/$0.01 cached input/$0.50 output per MTok. (Verified against the Sep 22 launch entry, which is the current pricing record on the changelog page.)
  • Responses and Chat Completions API surface for both models is unchanged.
  • Reasoning effort and tool-calling surfaces are unchanged.
  • Tool-calling constraints, context windows, and reasoning-effort range are unchanged.

Why developers and founders should care

The Sep 25 entry is targeted at exactly the cohort that tested the new models on vision during the launch window. The most-affected readers are:

  • Vision eval teams who benchmarked GPT-6 Sol or GPT-6 Luna between Sep 22 and Sep 25. Their results are now suspect by OpenAI's own recommendation. Action: rerun.
  • Agent operators using either model on Codex with the computer-use tool. The Sep 25 entry names computer use explicitly. Action: retry workflows from the same window.
  • Production agents that took GPT-6 Sol or GPT-6 Luna live between Sep 22 and Sep 25 for any vision-side task. Their in-flight decision quality may have been lower than the model's actual capability. Action: audit decision logs for the same window; expect some fraction of vision-grounded actions to be retriable with better outputs at the same input cost.
  • Buyers comparing model families on vision. A Sep 22–25 comparison against GPT-6 Sol or GPT-6 Luna cannot be trusted to reflect the post-fix capability. Comparisons before Sep 22 (against earlier models) are unaffected; comparisons after Sep 25 should reflect the fix.
  • Tool / SDK authors building vision pipelines on top of either model. Their pipeline logic is almost certainly correct; their example outputs and benchmarks may now be misleading artifacts. Action: document the fix in the project's release notes.

For readers who did not test the new models between Sep 22 and Sep 25, this is a quiet watch-the-record item rather than an urgent action. The fix is in the current API state; rerunning evaluations now gives you the post-fix baseline, not the pre-fix.

Evidence and test results

The evidence chain is documentation-surfacing. The single primary source is:

  • OpenAI developer changelog, Sep 25, 2026 entry. Verified verbatim at fetch 2026-09-26 20:08 UTC from https://developers.openai.com/api/docs/changelog. The page shows the entry above the Sep 22 launch entry and below no other Sep 25 entry — the same fetch also returns a fully-loaded Sep 22 entry and a Sep 1 IPv6 entry, but no other Sep 25 entry that conflicts with the one cited here.

For the Sep 22 launch context:

  • OpenAI developer changelog, Sep 22, 2026 entry. Verified verbatim at the same fetch. Confirms GPT-6 Sol ($2 input, $0.20 cached input, $10 output per MTok; standard pricing for prompts up to 272K input tokens) and GPT-6 Luna ($0.10 input, $0.01 cached input, $0.50 output per MTok), availability on Responses and Chat Completions APIs, and the cross-reference to the model catalog.

Cross-reference from the broader Sep 25 sweep:

  • OpenAI Codex releases atom feed. Verified at 20:08 UTC; the most recent Codex release with meaningful body content is rust-v0.157.0 (Sep 25), already covered in the weekly roundup. No Codex release between Sep 25 and the Sep 25 changelog entry specifically addresses the image-encoding surface beyond what the changelog itself states. (No specific commit, PR, or version is named as the fix vector; the changelog describes the user-visible change, not the deployment surface.)

Verification level. Documentation comparison plus direct quotes from the OpenAI changelog. No firsthand API call was made for this article. The article does not claim a measured before/after on any specific vision task, does not claim to know which image classes were most affected, does not claim to know what proportion of vision workflows were degraded, and does not quote any OpenAI engineer. All claims above are documentation-derived from the cited changelog entries.

Where the evidence stops. The Sep 25 entry does not specify:

  • The image input classes that were most affected (single image vs. multi-image; large vs. small; screenshot vs. document; color vs. grayscale).
  • Whether the fix changes cost (it does not appear to — pricing is unchanged on both models in the Sep 22 entry).
  • Whether the fix changes latency (not addressed by the changelog).
  • Whether the fix is incremental or requires a client-side update (it does not appear to — the changelog presents this as a server-side fix).
  • A list of specific vision benchmarks where the fix is most visible.

A team that wants a quantitative before/after will need to run their own benchmarks against the same prompts, models, and seed. The article labels this as an inferred requirement, not a tested one.

Cost, risk, and limitations

Cost. $0 to read the changelog entry. The recommended action — rerun vision evaluations and retry affected workflows — has direct inference-time cost. For a team that ran 500 GPT-6 Sol vision evaluations during Sep 22–25 at average 1,500 input tokens (mostly cached, 1 large image each) and 400 output tokens, a rerun at the same shape costs roughly:

  • Input cost: 500 × 1,500 × $2 / 1M = $1.50
  • Cached input cost: 500 × 1,500 × $0.20 / 1M = $0.15 (if every evaluation hit cache; the actual figure depends on prompt reuse)
  • Output cost: 500 × 400 × $10 / 1M = $2.00

So a representative 500-evaluation rerun is on the order of $2.00 to $3.65 for the smaller model; the same shape on GPT-6 Luna is roughly an order of magnitude cheaper. Inferred. These figures treat the input pricing as uncached input; teams with high prompt reuse will see lower cost via the cache-read tier. Treat the cost figures as approximate.

Risk. The actual operational risk is on teams that did not read the changelog and now have stale benchmarks or stale production decisions. The fix is in the current API state; future work without re-evaluating will compare against a now-superseded capability profile.

Limitations. This article does not run an eval against either model. It does not claim that the fix changes computer-use reliability, evaluation accuracy on any specific benchmark, or production decision quality on any specific agent. It is documentation-derived; for any quantitative comparison, run the comparison yourself on the current API state.

Mr. Technology verdict

This is a documentation-surfacing changelog entry that does not get the marketing amplification of a launch post but does matter for exactly the cohort the documentation targets. If you benchmarked GPT-6 Sol or GPT-6 Luna on vision between Sep 22 and Sep 25, treat those results as suspect by OpenAI's own recommendation. If you did not test either model during that window, watch the record and rerun any forward-looking baseline that touches the post-fix capability.

The interesting observation is the placement: a Fix category changelog entry that explicitly tells users to rerun evaluations is the upper bound of what OpenAI will acknowledge in a public changelog without an accompanying incident post. The fix is shipped; the documentation timeline is intact; OpenAI did not need to publish a postmortem. The bar for rerun-on-the-record is what you get here.

Recommended action

1. Today (vision eval teams). If you benchmarked GPT-6 Sol or GPT-6 Luna on vision between Sep 22 and Sep 25, set those benchmarks aside and rerun against the current API state. Compare the new numbers to your pre-launch baselines before publishing any comparison. 2. Today (Codex computer-use operators). If you ran GPT-6 Sol or GPT-6 Luna in Codex's computer-use path between Sep 22 and Sep 25, retry a representative sample of those workflows on the current API state. Expect measurable differences on screenshot-driven actions. 3. Today (production agents). Audit decision logs for vision-grounded actions that originated from either model between Sep 22 and Sep 25. For actions with low confidence, rerun the decision with the fix in place before treating the original decision as authoritative. 4. This week (SDK / tool authors). Add a release note that documents the Sep 25 fix on GPT-6 Sol / GPT-6 Luna; if any of your example outputs were generated between Sep 22 and Sep 25, regenerate them. 5. This week (buyers comparing model families). Recompute any GPT-6 Sol or GPT-6 Luna vision-side comparison that was anchored to Sep 22–25 baselines; do not publish comparison charts based on those numbers without an explicit caveat. 6. Skip if not in scope. If you did not test either model on vision during the Sep 22–25 window, this is a record-keeping item rather than an urgent action.

Sources

  • https://developers.openai.com/api/docs/changelog — OpenAI developer changelog, verified at fetch 2026-09-26 20:08 UTC. The Sep 25 entry under the Fix category is the primary source for every claim about what shipped on Sep 25; the Sep 22 entry under the Feature category is the canonical record of the GPT-6 Sol and GPT-6 Luna launch.

Article history

  • Originally published: 2026-09-26 22:08 Berlin / 2026-09-26 20:08 UTC.
  • Last verified: 2026-09-26 20:08 UTC against the OpenAI changelog.
  • No corrections at this time.
Related Dispatches