OpenAI Adds WebSocket Mode to Speed Up Agent Loops
07 Jul 2026
OpenAI has rolled out a WebSocket-based upgrade to its Responses API, alongside a new fast coding model called GPT-5.3-Codex-Spark, both aimed at cutting latency in agentic workflows.
What happened
Around November 2025, OpenAI kicked off a performance sprint focused on the Responses API. Part of that effort included a dedicated two-month sprint to build a new WebSocket mode, designed to speed up the back-and-forth communication that agent-based products rely on. OpenAI first tested the mode in alpha with a select group of coding agent startups before introducing GPT-5.3-Codex-Spark, described as a fast coding model built to take advantage of the new infrastructure.
The numbers
According to OpenAI, agent loops using the upgraded API ran 40% faster end-to-end, and alpha users reported similar gains—up to 40% improvements in their agentic workflows. On the inference side, speeds reportedly jumped from about 65 tokens per second (roughly what prior flagship models like GPT-5 and GPT-5.2 achieved) to nearly 1,000 tokens per second. Initial optimizations alone delivered close to a 45% improvement in time to first token. OpenAI says GPT-5.3-Codex-Spark hit its target of over 1,000 tokens per second, with bursts reaching as high as 4,000 TPS. The company also indicates that Codex users on GPT-5.3-Codex, GPT-5.4, and future models should see benefits from WebSocket mode.
Why founders should care
For founders building agentic or coding-agent products on the Responses API, these figures suggest latency could become less of a bottleneck—though the extent of real-world benefit is not yet certain. A 40% end-to-end speedup, if it holds outside alpha conditions, may make latency-sensitive agent applications more viable without needing to switch providers. Startups that gain early access to WebSocket mode could plausibly capture a timing advantage over competitors still on the older architecture, though it's unclear how quickly broader availability will follow. The launch of a dedicated fast model also hints that OpenAI sees speed as a differentiator for coding-agent tooling—something founders in that space may want to track as they plan future integrations.
Reasons for caution
All the performance figures reported here come from OpenAI itself, without independent verification. The 40% workflow improvement is based on an alpha test involving an undisclosed, limited set of startups, so it may not generalize to the broader developer base. The headline benchmarks—65 to nearly 1,000 tokens per second—could reflect best-case conditions rather than typical production performance. Adopting WebSocket mode may also introduce new integration work or create tighter dependency on OpenAI's evolving infrastructure.
What's still unclear
Several important details remain unaddressed. OpenAI hasn't said when WebSocket mode or GPT-5.3-Codex-Spark will move beyond alpha to general availability, nor has it disclosed pricing or cost implications. There's no independent benchmark confirming the reported speed gains, and OpenAI hasn't named which coding agent startups participated in the alpha or how they were chosen. The methodology behind the 40%, 45%, and tokens-per-second figures also hasn't been detailed, and it's not yet clear whether existing Responses API integrations will need code changes to adopt the new mode.
Bottom line
OpenAI's push to speed up the Responses API and introduce a fast coding model signals a clear bet on latency as a competitive lever for agentic products. Founders building in this space may want to watch for general availability and pricing details before committing engineering resources, while treating the current performance claims as promising but unverified.