GPT-6.1 Sol vs GPT-6 Sol: Better Agents, Same API Price?

A new AI model usually invites one question: how much smarter is it? GPT-6.1 Sol makes me ask a slightly less glamorous one—how much does it cost to finish a job?

OpenAI released GPT-6 Sol on September 22 and introduced GPT-6.1 Sol a week later. That is a fast upgrade cycle even by AI standards. The new name does not mean that GPT-6 Astra or Luna has been replaced.

OpenAI says the new Sol is stronger at agentic coding, computer use, and professional work. Yet its headline API input and output prices are unchanged from GPT-6 Sol. Cached input is the exception.

If one model can complete more tasks at the same per-token rate, the interesting story may be less about a chatbot leaderboard and more about the economics of cloud platforms, enterprise agents, and the hardware underneath them.

Previous Tech(EN) post: AT&T’s $3 Billion Corning Fiber Deal: Who Earns What?



Key Takeaways

  • The September 29 release is GPT-6.1 Sol, an upgrade to GPT-6 Sol—not an across-the-board renaming of the GPT-6 family. OpenAI describes better performance in coding agents, computer use, and professional workflows.
  • On OpenAI's standard API price list for short-context requests, both Sol models cost $2 per million uncached input tokens and $10 per million output tokens. Cached input changes from $0.20 to $0.10 per million.
  • Both list a 1.05-million-token context window and a 128,000-token maximum output. GPT-6.1 Sol removes the none reasoning setting; tool-using applications should use the Responses API.
  • AWS announced GPT-6.1 Sol's general availability on Bedrock on launch day. That is a verified distribution channel, not evidence that a new model has already generated measurable revenue for Amazon, Microsoft, or NVIDIA.
OpenAI API price comparison per million tokens: GPT-6 Sol and GPT-6.1 Sol both cost $2 standard input and $10 output; cached input falls from $0.20 to $0.10
OpenAI standard API rates, short context, per million tokens. A lower cached-input rate is not a 50% cut to a full agent's bill.

Original graphic based on OpenAI's GPT-6 Sol and GPT-6.1 Sol model pages, checked September 30, 2026.



The price that changed—and the two that did not

It would be easy to describe GPT-6.1 Sol as “Astra-like intelligence for less.” That is how OpenAI positions it against its premium Astra model. But the more useful comparison for an existing customer is with GPT-6 Sol. On that comparison, regular input and output token prices are flat. The directly visible discount is on cached input: 10 cents rather than 20 cents per million tokens.

That difference can matter to an agent repeatedly reading the same instructions, documents, or repository context. It does not mean the whole workload is 50% cheaper. An application still pays for uncached input, generated output, and any relevant tools; cache savings depend on how much of a request actually hits the cache. A pricing decision should be based on a representative workload rather than the biggest percentage in a model announcement.

The second potential saving is harder to price from a catalog. If a stronger model takes fewer wrong turns, retries less often, and needs less human correction, the cost per successful task can fall even at unchanged token rates. In an AWS post citing OpenAI's evaluation, GPT-6.1 Sol beat GPT-6 Sol's best DeepSWE v1.1 score by 6.4 percentage points and matched Astra at roughly one-fifth the cost per task. That is a vendor-reported result on a particular coding benchmark, not a guarantee for every enterprise workflow.



For developers, the migration is more than a model-name swap

Neither context capacity nor maximum output expands in the official model specifications: both Sol versions list 1.05 million tokens of context and 128,000 output tokens. The changes that could break an existing integration are more mundane. GPT-6 Sol allows reasoning effort none; GPT-6.1 Sol starts at low. OpenAI also directs tool-using GPT-6.1 Sol applications to the Responses API, while Chat Completions is supported without tool calls.

That matters because an enterprise agent is not just a model response. It is a chain of retrieval, tools, approvals, retries, and audit logs. A successful upgrade needs to test completed-task quality, latency, safety controls, and full cost under the same workload. The strongest demo is not necessarily the best production deployment.

Outside the API, OpenAI's product guidance says access to GPT-6.1 Sol depends on plan, client, and workspace settings. A model being announced is not the same as every user seeing it in every product on day one.



Related Companies

Amazon (AMZN, Nasdaq) has the clearest launch-day link. AWS says GPT-6.1 Sol is generally available on Bedrock, giving customers an existing enterprise channel for adoption. More production usage could support Bedrock revenue, but neither adoption volumes nor model-specific margins have been disclosed. Better efficiency could even mean fewer billable tokens per task unless broader usage offsets it.

Microsoft (MSFT, Nasdaq) has a different link: its amended OpenAI partnership identifies it as OpenAI's primary cloud partner and describes its nonexclusive model-and-product IP license. Those are real business connections, but I would not turn the release announcement into an unverified GPT-6.1 Sol deployment figure or a Microsoft revenue estimate. The model's confirmed availability and usage within each Microsoft product need their own evidence.

NVIDIA (NVDA, Nasdaq) has a longer-horizon infrastructure connection. Its OpenAI partnership announcement discusses future systems for training and running models. It does not identify an incremental GPU order caused by this week's release. More successful agents might drive more inference demand; improved efficiency might reduce compute used per task. The net hardware effect cannot be read from the model name alone.



Investment Watchpoints

Three numbers would tell us more than launch-day excitement: first, real completion rates and retries versus GPT-6 Sol at similar settings; second, production adoption through channels such as Bedrock and Microsoft's AI platforms; and third, whether that usage shows up in cloud growth, margins, or infrastructure commitments. A benchmark, a product listing, and recognized revenue are three separate milestones.

The counterargument is worth keeping close. If the new model is materially more efficient, customers might consume fewer tokens for the same work. If quality unlocks many new use cases, total usage could still rise. Both outcomes are plausible, and neither is a reason to assign immediate stock gains to every company in the AI supply chain. This is industry analysis, not a buy or sell recommendation.



Appendix. Why cached input deserves its own line

When a developer sends the same prompt prefix or reference material across requests, eligible reused input may be billed at a cached rate. The fraction that is actually cached depends on request design and cache behavior. That is why the new Sol's 50% cached-input rate reduction is useful but cannot be translated into a blanket 50% reduction in an agent's operating cost.



Sources and Update

Checked September 30, 2026. Model availability, prices, and supported features can change.