All news
aiproductsaas

OpenAI Adds Sandbox Execution to Agents SDK

07 Jul 2026

OpenAI has expanded its Agents SDK with two new capabilities aimed at production-grade agent workflows: a model-native harness and native sandbox execution. The update signals a push toward more reliable, infrastructure-ready agents rather than prototype-stage tooling.

What's new

The updated harness now includes configurable memory, sandbox-aware orchestration, Codex-like filesystem tools, and standardized integrations. Alongside it, OpenAI introduced native sandbox execution support for seven third-party providers: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel.

A new Manifest abstraction describes an agent's workspace and supports four storage integrations: AWS S3, Google Cloud Storage, Azure Blob Storage, and Cloudflare R2. Together, these additions are designed to let developers stand up agent infrastructure with less custom engineering work.

Rollout details

The new harness and sandbox capabilities are launching first in Python, with TypeScript support planned for a later release. OpenAI says it is also working on additional agent capabilities — code mode and subagents — for both Python and TypeScript, though no timeline has been given for either.

The features are generally available to all customers through the API, priced under OpenAI's standard token and tool-use rates. No special access requirements were mentioned.

Early results: Oscar Health

Health insurer Oscar Health tested the updated SDK and reported that it enabled production-viable automation of a clinical records workflow that prior approaches could not reliably handle. Oscar Health's Rachael Burns is quoted describing the improvement in reliability for this specific workflow, though the report does not detail the workflow beyond its general description.

What's missing

Several details remain unclear from OpenAI's announcement:

  • No specific release date or version number for the Python launch
  • No timeline for TypeScript support, code mode, or subagents
  • No pricing differences disclosed between sandbox providers
  • No information on adoption beyond the Oscar Health example
  • No further specifics on the clinical records workflow itself

Why founders should care

For startups building on top of OpenAI's agent tooling, this update is likely to matter in a few ways:

  • Lower engineering overhead is plausible. Native sandboxing and filesystem tools may reduce the custom infrastructure work needed to move agents from prototype to production.
  • Multi-cloud flexibility could ease integration. Because the Manifest abstraction supports AWS, GCP, Azure, and Cloudflare storage, teams already using any of these providers may find onboarding simpler.
  • Python-first teams may have a temporary edge. Startups building primarily in Python could gain earlier access to advanced agent features, while TypeScript-based teams may need to wait for parity — a timing gap that could affect competitive positioning in agent-heavy product categories.
  • Reliability claims deserve scrutiny. The Oscar Health example suggests the SDK can now handle reliability-sensitive workflows, but with only one named customer case and no adoption data, founders should treat this as an early signal rather than proven track record.
  • Costs may become harder to predict at scale. Because pricing follows standard token and tool-use rates, high-volume agent workflows could see less predictable costs, particularly given the added sandbox and orchestration layers.

The risks

A few caveats are worth flagging. Native sandbox execution depends on seven third-party providers, which introduces potential dependency or vendor lock-in concerns. TypeScript teams will need to wait for a future release before accessing the same capabilities. And because pricing is based on token and tool-use consumption, cost predictability for heavy agent workloads is not guaranteed.

Overall, the update points to OpenAI positioning the Agents SDK for more demanding, production-level use cases — but with key details on timing, pricing, and broader adoption still undisclosed, founders evaluating the platform should watch for further specifics before committing significant engineering resources.

Sources