OpenAI DevDay Drops Sol, Ultrafast, and Cloud Codex for Developers

DevDay 2026 brings a cheaper GPT-6.1 Sol model, an Ultrafast tier at 300 tokens/sec, cloud Codex environments, voice-driven CLI, and plugin extensions inside ChatGPT.

·
·
OpenAI DevDay Drops Sol, Ultrafast, and Cloud Codex for Developers
  • GPT-6.1 Sol lands at a fifth of GPT-6 Astra's token price with near-Astra coding performance.
  • Ultrafast tier pushes up to 300 tokens/sec, 8x faster in Codex and 6x in the API.
  • Codex now runs in reusable cloud environments, steerable from phone or any device.
  • Refreshed Codex CLI adds voice control, an /agents view, and better worktree support.
  • New code review in ChatGPT desktop inspects diffs on GitHub and GitLab PRs.
  • Plugin extensions plus MCP Events let plugins claim sidebar UI and trigger on app events.

OpenAI DevDay 2026 puts agents in the cloud and adds a faster API lane

Among the more than 20 announcements in OpenAI’s DevDay recap, five directly affect how developers build and operate agents: a lower-cost reasoning model, a premium API service tier, cloud-hosted Codex, new review tools, and richer ChatGPT plugins. The releases target the recurring constraints of agent development: model cost, response latency, execution environments, human oversight, and integration with external systems.

Sol lowers the cost of long agent loops

GPT-6.1 Sol is OpenAI’s new mid-tier reasoning model for agentic coding, computer use, and multi-step workflows across applications. OpenAI positions its performance near GPT-6 Astra while charging one fifth of Astra’s standard input and output token prices.

Agent workloads multiply token usage as they plan, call tools, inspect results, and retry failed steps. Sol’s lower rates can support longer runs within the same budget, although teams should validate task quality and retry rates against their own evaluations before switching models. GPT-6.1 Sol is available now through the API and on Plus, Pro, Business, Enterprise, and Edu plans.

Ultrafast buys back response time

Ultrafast is a premium service tier that accelerates supported models. OpenAI reports up to eight times faster generation in Codex, reaching as many as 300 tokens per second, and up to six times faster generation through the API. GPT-6 Astra supports Ultrafast now, with Sol support planned.

API calls opt into the tier individually by setting service_tier to ultrafast, which allows applications to reserve the added cost for interactive requests while leaving batch and background jobs on the standard tier.

python
from openai import OpenAI

client = OpenAI()

resp = client.responses.create(
    model="gpt-6-astra",
    service_tier="ultrafast",
    input="Explain why the sky is blue in one sentence.",
)

print(resp.output_text)

OpenAI recommends using Ultrafast over WebSockets because a persistent connection avoids repeated connection setup and reduces network overhead across tool-heavy agent loops. The synchronous example above shows the request flag; production latency tests should use the WebSocket path described in the Ultrafast docs.

Default limits start at 500,000 tokens per minute for usage tiers 1 through 3 and rise to 5 million tokens per minute at tier 5. Processing is limited to US and global regions, with no EU regional endpoint currently available. Teams with residency requirements should resolve that constraint before adoption and compare the tier’s premium pricing with measured latency gains.

Codex leaves the laptop

Codex can now execute locally, in OpenAI’s cloud, or remotely under control from another device, including a phone. Reusable cloud development environments preserve approved settings, dependencies, and permissions, reducing startup time and giving teams a shared execution setup. A developer can launch a migration from a laptop, monitor it from a phone, and continue steering it from another computer.

The rebuilt Codex CLI adds several controls for longer and more concurrent sessions:

  • Voice input for starting and steering tasks from the terminal
  • An /agents view for delegating and tracking concurrent work
  • Improved Git worktree handling, prompt editing, and session resume
  • A cleaner terminal interface for long-running sessions

Codex reviews branches before merge

The new desktop review flow lets developers open a pull request or merge request in the ChatGPT Work or Codex desktop app, read a summary, inspect the diff, and question Codex about specific changes before posting feedback to GitHub or GitLab. It can also review an unshared branch before a pull request exists.

Cloud review mode can produce an initial set of findings while the developer is away, shortening the manual review pass when work resumes. OpenAI says the feature is available on every plan.

Plugins gain screens and triggers

Plugin extensions expose the platform OpenAI uses for its own ChatGPT integrations. Third-party plugins can now occupy a sidebar position, render interactive panels beside a conversation, and provide custom viewers for product-specific file formats. OpenAI has also added clearer feedback to the submission review process.

OpenAI also supports the proposed Model Context Protocol Events specification. MCP standardizes connections between AI clients and external tools; Events allows a connected application to send a notification that starts an automation. A project board could emit an event when a task appears, prompting ChatGPT to read linked documents and draft an implementation plan.

Event-driven integrations require controls beyond prompt design, including least-privilege credentials, event deduplication, bounded retries, and audit logs. Because MCP Events remains a proposed specification, integrations should isolate protocol handling so later schema changes remain manageable.

The agent stack stretches across surfaces

  1. Model economics: Sol’s lower token prices can extend planning and correction loops without increasing the model budget at the same rate.
  2. Request-level latency: Ultrafast makes generation speed a configurable API choice for endpoints where users are waiting.
  3. Persistent execution: Cloud environments and remote controls allow Codex tasks to continue beyond a local terminal session.
  4. Earlier review: Branch and cloud reviews move automated feedback ahead of the final human review pass.
  5. Event-driven integration: Plugin panels and MCP events give ChatGPT both an interface for third-party software and a mechanism for responding to external changes.

Roll out with benchmarks and guardrails

  1. Evaluate Sol on real tasks. Compare completion quality, retries, token usage, latency, and total cost with the current model before changing production defaults.
  2. Test Ultrafast over WebSockets. Measure median and tail time to first token, total completion time, tool-call latency, and cost on user-facing endpoints.
  3. Pilot cloud Codex with scoped access. Pin environment dependencies, limit repository credentials, and define when agents may write, commit, or push code.
  4. Design MCP event handlers for failure. Add idempotency, authorization checks, retry limits, and logs before allowing events to launch consequential workflows.
Trending
  • No trending articles

Comments

avatar

Next Reads