# The complete guide to the OpenAI Agents API

> The OpenAI Agents API runs the Codex agent harness for your app. A hands-on guide to sessions, events, tools, MCP, sandboxes, files, webhooks and costs.

Author: [Flavio Copes](https://flaviocopes.com/about/) | Published: 2026-10-01 | Topics: [AI](https://flaviocopes.com/tags/ai/) | Canonical: https://flaviocopes.com/agents-api/

The OpenAI Agents API lets your application use the agent that powers Codex. You send it a task, and OpenAI runs the whole agent loop: it calls the model, runs tools, executes commands in a sandbox if you give it one, and saves every step in a session you can come back to later.

It's in public beta, and it's a different thing from the Responses API you may already know. With the Responses API you build the agent loop yourself. With the Agents API the loop already exists, and your code steers it.

In this guide we'll build one small project, step by step: a **tutorial checker**. It reads a programming tutorial, finds the mistakes, runs the code blocks in a sandbox, and writes a corrected version you can download.

Each step adds one part of the API. By the end we'll have used sessions, events, structured output, saved agents, function tools, web search, MCP servers, hosted sandboxes, files, secrets, subagents and webhooks.

I picked this example because I publish a lot of tutorials, and the worst bug a tutorial can have is a code block that doesn't run.

## What is the Agents API?

A model on its own takes text in and gives text out. To get real work done, something has to wrap it in a loop.

That something sends the conversation to the model. It notices when the model wants to run a tool, runs it, and sends the result back. It repeats until the task is done. When the conversation gets too long, it compacts it. When the task is big, it can split it across helpers.

This wrapper is called a **harness**. Codex has one, and it's what lets [Codex](https://flaviocopes.com/codex/) work on its own for a long time.

The Agents API gives you that same harness as an API. OpenAI runs it on its servers. Your application creates sessions, sends messages, answers the agent when it asks for something, and reads the results.

OpenAI handles the agent loop and the model calls. It saves the conversation, every turn and every tool call. It compacts the context when it grows, starts subagents when you enable them, and can give the agent a sandbox to run code in.

You decide the model and the instructions, which tools the agent can use, and where it runs code. That can be nowhere, a sandbox hosted by OpenAI, or a machine you control. And of course you decide what your app does with the results.

## Agents API, Agents SDK or Responses API?

OpenAI now has three ways to build an agent, and they overlap. This is how they compare, following OpenAI's own comparison table:

| | Agents API | Agents SDK | Responses API |
| --- | --- | --- | --- |
| Who runs the agent loop | OpenAI, with the Codex harness | The SDK, inside your app | You write it |
| Where the state lives | In the session, saved by OpenAI | Your storage or SDK sessions | You keep the history or chain responses |
| Where code runs | OpenAI's sandbox, your sandbox, or nowhere | Your runtime | Your own environment |
| Effort to build an agent | Low | Medium | High |

Pick the Agents API for long-running tasks where you want OpenAI to manage the agent and keep its progress.

Pick the Agents SDK when you want the loop inside your own application, with your own storage and your own approval flow.

Pick the Responses API when you only need a model answer, or when you want to build everything yourself. If model calls, tokens and tool calling are new to you, the free [AI fundamentals course](https://flaviocopes.com/courses/ai-fundamentals/) covers them before you get here.

## The pieces: agents, environments, sessions and turns

Before any code, a few words we'll use everywhere.

An **agent** is a configuration: the model, the instructions, the tools and the reasoning settings. You can pass it inline every time, or save it once and reuse it by ID.

An **environment** is where the agent can run commands and work with files. There are three types:

- `none`: no filesystem and no shell. The agent can still think, call your functions, search the web and use MCP servers.
- `openai_hosted`: a sandbox OpenAI creates for the session, with Python, Node.js and common command line tools.
- `self_hosted`: a machine you run, connected to OpenAI by a small program called the executor.

A **session** is one conversation with one agent in one environment. It keeps the whole history, so you can come back hours later and send another message.

A **turn** is one round of work. You send a message to an idle session, the agent works until it's done, and the turn ends as completed, failed or cancelled.

While a turn runs, the session sends you **events**: text as it's written, tool calls, commands, status changes. Events are for live updates.

Everything important is also saved as **items**: messages, tool calls, command runs. You can fetch items later, even if you missed the events.

```mermaid
flowchart TB
  app["Your app"] -->|"message"| session["Session"]
  session --> turn["Turn"]
  turn -->|"model calls"| model["Model"]
  turn -->|"tool calls"| tools["Functions, MCP, web search"]
  turn -->|"commands and files"| env["Environment"]
  turn -->|"live events"| app
  turn --> items["Saved items"]
```

A session also has a status. It's `idle` when it waits for your next message, `in_progress` while a turn runs, `requires_action` when the agent waits for you (for example, for the result of one of your functions), and `failed` when something broke for good.

## What you need

To follow along you need:

- An OpenAI Platform account with API credits.
- An API key. If you create a restricted key, grant it `api.agents.read` and `api.agents.write` for sessions, plus `api.responses.write` for the model calls.
- A recent version of Node.js.
- The `openai` package from npm.

The API is in beta, so every request needs the `OpenAI-Beta: agents=v1` header. The official SDK adds it for you. If you use curl, add it yourself.

Let's create the project:

```bash
mkdir tutorial-checker
cd tutorial-checker
npm init -y
npm install openai
mkdir tutorials
```

Put your key in a `.env` file:

```
OPENAI_API_KEY=sk-proj-your-key-here
```

Add `.env` to your `.gitignore` right away. We'll run every script with `node --env-file=.env <file>.mjs`. Node loads the key from the file, so the same command works in any shell.

Now we need a tutorial to check. Save this as `tutorials/count-words.md`. It has two bugs on purpose:

````markdown
# Count the words in a file with Node.js

Read the file, split the text on spaces, and count the pieces.

```js
import fs from 'node:fs/promises'

const text = fs.readFileSync('notes.txt', 'utf8')
console.log(text.split(' ').length)
```

Run it with `node count.mjs`.
````

The first bug crashes the script: `node:fs/promises` has no `readFileSync`. The second one is sneakier. Splitting on a single space gives the wrong count as soon as the text has line breaks or two spaces in a row.

A good checker should find both.

## Step 1: your first session

Let's send the tutorial to an agent and ask for a review. Create `01-review.mjs`:

```js
import fs from 'node:fs/promises'
import OpenAI from 'openai'

const client = new OpenAI()
const tutorial = await fs.readFile('tutorials/count-words.md', 'utf8')

const events = await client.beta.agents.sessions.create({
  agent: { model: 'gpt-6.1-sol' },
  environment: { type: 'none' },
  input: `Review this tutorial and list every mistake a reader would hit.\n\n${tutorial}`,
  stream: true,
})

try {
  for await (const event of events) {
    if (event.type === 'agent.session.created') {
      console.log(`Session ${event.session.id}\n`)
    }
    if (event.type === 'agent.session.turn.output_text.delta') {
      process.stdout.write(event.delta)
    }
    if (event.type === 'agent.session.turn.completed') break
  }
} finally {
  events.controller.abort()
}
```

Everything lives under `client.beta.agents`. Here we call `sessions.create()` with four things:

- `agent`: the configuration. For now, just the model.
- `environment`: `none`, because reading a Markdown file doesn't need a sandbox.
- `input`: the first message. A session without an environment needs one at creation.
- `stream: true`: instead of returning the session object, the API returns a stream of events for the first turn.

Then we loop over the events. We print the session ID, print the text as it arrives, and stop when the turn completes. The `finally` block closes the connection.

Run it:

```bash
node --env-file=.env 01-review.mjs
```

You'll see the session ID first, then the review, a few words at a time. The wording changes on every run, but it should point out the missing `readFileSync` and the fragile `split(' ')`.

Keep that session ID. It's the handle to this conversation, and we'll send more messages to it in a moment.

A note on the model: the official examples mostly use `gpt-6-astra`. The Agents API docs also use `gpt-6.1-sol`, which costs a fifth as much per token. This guide uses `gpt-6.1-sol`, and you can swap the model name anywhere.

## Step 2: a helper that follows a turn

The first script is too optimistic. It only handles three event types and assumes everything goes well.

Real turns can fail. They can get cancelled. They can stop to ask your code for something. The stream can close before the turn ends. And if you turn on subagents later, their turns show up in the same stream, and a subagent finishing doesn't mean the main agent is done.

One thing trips everybody up at first: an `idle` session doesn't mean the task succeeded. It only means the session waits for input. You have to look at how the turn ended.

So let's write the event handling once, in a file we'll reuse in every step. Create `agent.mjs`:

```js
import OpenAI from 'openai'

export const client = new OpenAI()

export async function followTurn(events, { onAction } = {}) {
  const subagentTurns = new Set()
  const streamed = new Set()
  let answer = ''

  try {
    for await (const event of events) {
      switch (event.type) {
        case 'agent.session.created':
          console.log(`Session ${event.session.id}\n`)
          break

        case 'agent.session.turn.created':
          if (event.turn.subagent_id !== null) subagentTurns.add(event.turn_id)
          break

        case 'agent.session.turn.output_text.delta':
          if (subagentTurns.has(event.turn_id)) break
          streamed.add(event.item_id)
          process.stdout.write(event.delta)
          break

        case 'agent.session.turn.output_text.done':
          if (subagentTurns.has(event.turn_id)) break
          if (!streamed.has(event.item_id)) process.stdout.write(event.text)
          process.stdout.write('\n')
          break

        case 'agent.session.turn.item.done': {
          const item = event.item
          if (item.type === 'command_execution') {
            console.log(`\n$ ${item.command} (exit code ${item.exit_code})`)
          }
          if (item.type === 'message' && item.phase === 'final_answer' && !subagentTurns.has(item.turn_id)) {
            answer = item.content.map((part) => part.text).join('')
          }
          break
        }

        case 'agent.session.requires_action':
          if (onAction) await onAction(event.session)
          break

        case 'error':
          throw new Error(event.error.message)

        case 'agent.session.failed':
        case 'agent.session.environment.failed':
          throw new Error(`The session stopped: ${event.type}`)

        case 'agent.session.turn.failed':
        case 'agent.session.turn.cancelled':
          if (event.turn.subagent_id === null) {
            throw new Error(`${event.type}: ${event.turn.error?.message ?? 'no details'}`)
          }
          break

        case 'agent.session.turn.completed':
          if (event.turn.subagent_id === null) {
            return { turn: event.turn, usage: event.usage, answer }
          }
          break
      }
    }

    throw new Error('The stream closed before the turn ended. Check the saved session.')
  } finally {
    events.controller.abort()
  }
}
```

It looks long, but each case does one small job:

- `turn.created` tells us whether a turn belongs to the main agent or to a subagent. Subagent turns have a `subagent_id`, the main agent's turns have `null`. We remember the subagent ones so we can skip their text.
- `output_text.delta` is a piece of text as the model writes it. `output_text.done` carries the complete text of that part. The docs warn that deltas can be missing, so if we never printed any, we print the full text instead.
- `item.done` fires when an item is saved. Command runs show up here once we add a sandbox. The final answer is a message with `phase: 'final_answer'`. Progress notes the agent writes along the way have `phase: 'commentary'`.
- `requires_action` means the agent waits for us. We'll use it for function tools in step 8.
- The failure cases throw, so a broken turn never looks like a success.
- `turn.completed` for the main agent ends the loop and returns the turn, its token usage and the final answer.

If the loop ends without a verdict, we throw. A closed stream doesn't tell you what happened to the turn, and we'll see how to find out in step 4.

Now `01-review.mjs` gets shorter:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const tutorial = await fs.readFile('tutorials/count-words.md', 'utf8')

const events = await client.beta.agents.sessions.create({
  agent: { model: 'gpt-6.1-sol' },
  environment: { type: 'none' },
  input: `Review this tutorial and list every mistake a reader would hit.\n\n${tutorial}`,
  stream: true,
})

const { turn } = await followTurn(events)
console.log(`\nSession ID: ${turn.session_id}`)
```

## Step 3: continue, steer and cancel

A session remembers everything, so a follow-up message doesn't need to repeat the tutorial. Create `03-fix.mjs`:

```js
import { client, followTurn } from './agent.mjs'

const sessionId = process.argv[2]

const events = client.beta.agents.sessions.stream(sessionId, {
  input: 'Rewrite the code block so it works. Show only the fixed code.',
})

await followTurn(events)
```

Pass the session ID from step 1:

```bash
node --env-file=.env 03-fix.mjs sess_your_session_id
```

`sessions.stream()` is a helper in the JavaScript SDK. It opens the event stream first, then sends your message, so you never miss the first events of the turn. The session must be idle when you call it.

That order matters when you do it by hand. The event stream doesn't replay the past. If you send the message first and subscribe second, the opening events of the turn are gone.

What if a turn is already running and you change your mind? You send a message anyway. A message to an idle session starts a new turn. A message during a running turn **steers** it: the agent reads it and adjusts the work it's already doing.

Steering uses the lower-level `events.create()`:

```js
import { client } from './agent.mjs'

const sessionId = process.argv[2]

await client.beta.agents.sessions.events.create(sessionId, {
  events: [
    {
      type: 'agent.session.input.message',
      input: [
        {
          role: 'user',
          content: [{ type: 'input_text', text: 'Skip the style notes. Only report bugs.' }],
        },
      ],
    },
  ],
})
```

Not every turn accepts more input. When it can't, you get an `active_turn_not_steerable` error. Wait for the turn to end and send a normal follow-up instead.

To stop a turn, send a cancel event:

```js
await client.beta.agents.sessions.events.create(sessionId, {
  events: [{ type: 'agent.session.input.cancel' }],
})
```

The turn ends as `cancelled`, and our helper throws.

Be careful with one thing. Closing your stream, or pressing Ctrl+C on the script, does **not** stop the agent. The turn keeps running on OpenAI's side, and it keeps spending tokens. Only a cancel event stops it.

## Step 4: look at what the session saved

Events come and go. The session keeps the record. Create `04-inspect.mjs`:

```js
import { client } from './agent.mjs'

const sessionId = process.argv[2]

const session = await client.beta.agents.sessions.retrieve(sessionId)
console.log(`Status: ${session.status}`)
console.log('Usage so far:', session.usage)

for await (const turn of client.beta.agents.sessions.turns.list(sessionId, { order: 'asc' })) {
  console.log(`${turn.id} ${turn.status} ${turn.usage?.total_tokens ?? '?'} tokens`)
}

for await (const item of client.beta.agents.sessions.items.list(sessionId, { order: 'asc' })) {
  if (item.type === 'message') {
    const label = item.phase ? `${item.role}, ${item.phase}` : item.role
    console.log(`\n[${label}]`)
    console.log(item.content.map((part) => ('text' in part ? part.text : '')).join(''))
  } else {
    console.log(`\n[${item.type}]`)
  }
}
```

It prints three things.

The session: its status and total usage.

The turns: one line per round of work, with how it ended and how many tokens it used.

The items: every message and tool call the main agent made, oldest first. The list methods return pages, and `for await` walks through all of them for you.

Token usage has this shape:

```json
{
  "input_tokens": 5000,
  "input_tokens_details": { "cached_tokens": 1500 },
  "output_tokens": 900,
  "output_tokens_details": { "reasoning_tokens": 200 },
  "total_tokens": 5900
}
```

Cached tokens are part of `input_tokens`, and reasoning tokens are part of `output_tokens`. OpenAI calls these numbers best-effort. They can be `null` while accounting catches up, and they can change later. Use them to spot a runaway session, not as your bill.

This is also how you recover from a dropped connection. Streams don't replay missed events, so after a disconnect you:

1. Open a new stream and buffer what arrives.
2. Retrieve the session and its saved items while the stream stays open.
3. Rebuild your local state from the items, keyed by item ID.
4. Apply the buffered updates, skipping items that already reached their final state.
5. Go back to handling live events.

If the session says `requires_action`, its `required_actions` list tells you what's still pending. You don't need to resend the task.

## Step 5: configure the agent

So far the agent only had a model. Let's give it a job description. Create `05-configured.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const tutorial = await fs.readFile('tutorials/count-words.md', 'utf8')

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    instructions: [
      'You review programming tutorials written for beginners.',
      'Report the problems a reader would hit, the most serious first.',
      'Never invent APIs. If you are not sure a function exists, say so.',
    ].join('\n'),
    reasoning: { effort: 'high' },
    text: { verbosity: 'low' },
  },
  environment: { type: 'none' },
  metadata: { tutorial: 'count-words', source: 'cli' },
  input: `Review this tutorial.\n\n${tutorial}`,
  stream: true,
})

await followTurn(events)
```

Here's what the new fields do:

- `instructions` works like a system prompt. It applies to every turn of the session.
- `reasoning.effort` sets how hard the model thinks before it answers. The values go from `none` to `max`, and each model supports its own range. `gpt-6.1-sol` accepts `low` through `max` and defaults to `medium`. More effort means better answers on hard problems, more tokens and more time.
- `text.verbosity` asks for `low`, `medium` or `high` amounts of text. The default is `medium`.
- `metadata` is yours. Up to 16 key-value pairs, keys up to 64 characters, values up to 512. It's how you'll recognize a session later, for example inside a webhook.

You can also set `service_tier` to trade speed for price, and `reasoning.summary` to get a summary of the model's reasoning as events.

Settings can change during a session:

```js
await client.beta.agents.sessions.update(sessionId, {
  agent: { reasoning: { effort: 'xhigh' } },
})
```

The change applies to new turns. A turn that's already running keeps its old settings. You can switch `model`, `reasoning.effort` and `service_tier` this way, and the conversation history stays. If the new model doesn't support a setting you have, like a reasoning effort it doesn't offer, the update fails, so change both in the same request.

## Step 6: get JSON back

A review in prose is nice to read. If you want to save the issues in a database or show them in a UI, you want JSON.

The agent's `text.format` accepts a JSON Schema. The generated text must match it. Create `06-json.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const schema = {
  type: 'object',
  properties: {
    issues: {
      type: 'array',
      items: {
        type: 'object',
        properties: {
          severity: { type: 'string', enum: ['breaks', 'misleads', 'style'] },
          problem: { type: 'string' },
          fix: { type: 'string' },
        },
        required: ['severity', 'problem', 'fix'],
        additionalProperties: false,
      },
    },
  },
  required: ['issues'],
  additionalProperties: false,
}

const tutorial = await fs.readFile('tutorials/count-words.md', 'utf8')

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    text: { format: { type: 'json_schema', schema } },
  },
  environment: { type: 'none' },
  input: `Review this tutorial.\n\n${tutorial}`,
  stream: true,
})

const { answer } = await followTurn(events)
const { issues } = JSON.parse(answer)

for (const issue of issues) {
  console.log(`[${issue.severity}] ${issue.problem}\n  Fix: ${issue.fix}`)
}
```

The helper already kept the final answer, so we parse it and print one line per issue.

Structured output is in the API reference and in the SDK types. The guides only mention it in passing, as the "format" of the agent's responses. Keep the combined size of your input and your schema under 4 MiB, the request limit of the agent runtime.

## Step 7: save the agent

We keep repeating the same model and instructions. A saved agent stores them once. Create `07-save-agent.mjs`:

```js
import { client } from './agent.mjs'

const agent = await client.beta.agents.create({
  name: 'tutorial-checker',
  model: 'gpt-6.1-sol',
  instructions: [
    'You review programming tutorials written for beginners.',
    'Report the problems a reader would hit, the most serious first.',
    'Never invent APIs. If you are not sure a function exists, say so.',
  ].join('\n'),
  reasoning: { effort: 'high' },
})

console.log(agent.id)
```

Run it once and keep the ID. From now on a session can point to it with `agent_id`:

```js
const events = await client.beta.agents.sessions.create({
  agent_id: 'agent_your_agent_id',
  environment: { type: 'none' },
  input: `Review this tutorial.\n\n${tutorial}`,
  stream: true,
})
```

You can pass `agent_id` and `agent` together to change something for one session only:

```js
const events = await client.beta.agents.sessions.create({
  agent_id: 'agent_your_agent_id',
  agent: { reasoning: { effort: 'low' } },
  environment: { type: 'none' },
  input: `Quick check, bugs only.\n\n${tutorial}`,
  stream: true,
})
```

Two rules to remember. An override replaces the whole field instead of merging: pass `tools` and you replace the saved tool list. And the session copies the saved agent when it's created, so updating the saved agent with `client.beta.agents.update()` only affects new sessions.

## Step 8: let the agent call your code

Right now we paste the tutorial into the message. That doesn't scale. A real checker would fetch tutorials from wherever they live: a database, a CMS, a folder.

A **function tool** lets the agent ask your application for data. You describe the function with a name, a description and a JSON Schema for its arguments. When the agent wants to call it, the session pauses in `requires_action`, your code runs the function, and you send the result back.

Let's add a second tutorial first, so there's something to choose from. Save this as `tutorials/read-json.md`:

````markdown
# Read a JSON file in Node.js

```js
import fs from 'node:fs'

const config = JSON.parse(fs.readFileSync('config.json', 'utf8'))
console.log(config.port)
```
````

This one is correct. A good checker also needs to say when there's nothing to fix.

Now create `08-functions.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const getTutorial = {
  type: 'function',
  name: 'get_tutorial',
  description: 'Get the Markdown source of a tutorial by its slug, like count-words.',
  parameters: {
    type: 'object',
    properties: { slug: { type: 'string' } },
    required: ['slug'],
    additionalProperties: false,
  },
}

async function readTutorial(slug) {
  if (!/^[a-z0-9-]+$/.test(slug)) throw new Error(`Invalid slug: ${slug}`)
  return fs.readFile(`tutorials/${slug}.md`, 'utf8')
}

const handlers = {
  get_tutorial: (args) => readTutorial(args.slug),
}

async function answerCalls(session) {
  const results = []

  for (const action of session.required_actions) {
    if (action.type !== 'function_call') continue
    const ids = { turn_id: action.turn_id, call_id: action.call_id }

    try {
      const output = await handlers[action.name](action.arguments)
      results.push({ type: 'agent.session.input.tool_result', ...ids, success: true, output })
    } catch (error) {
      results.push({ type: 'agent.session.input.tool_result', ...ids, success: false, error: error.message })
    }
  }

  await client.beta.agents.sessions.events.create(session.id, { events: results })
}

const events = await client.beta.agents.sessions.create({
  agent: { model: 'gpt-6.1-sol', tools: [getTutorial] },
  environment: { type: 'none' },
  input: 'Review the count-words and read-json tutorials. Use get_tutorial to read them.',
  stream: true,
})

const { turn } = await followTurn(events, { onAction: answerCalls })
console.log(`\nSession ID: ${turn.session_id}`)
```

When the agent calls `get_tutorial`, the session emits `requires_action`. Its `required_actions` list holds one entry per pending call:

```json
{
  "type": "function_call",
  "turn_id": "turn_123",
  "call_id": "call_123",
  "name": "get_tutorial",
  "arguments": { "slug": "count-words" }
}
```

Our `answerCalls()` runs the matching handler and sends back a `tool_result` with the same `turn_id` and `call_id`. The output must be a string, so serialize objects with `JSON.stringify()`. If the function fails, we send `success: false` with an error message the agent can read. It might try another slug, or tell you the tutorial doesn't exist.

The agent may ask for both tutorials at once. That's why `answerCalls()` loops over every pending call and sends all the results in one request.

Always validate the arguments. The model picks them, so treat them like user input. Here the regex keeps the agent from reading `../../.env`.

### Let the SDK run the handlers

For follow-up turns, the `sessions.stream()` helper from step 3 can run your handlers for you:

```js
const events = client.beta.agents.sessions.stream(sessionId, {
  input: 'Check read-json again, and tell me if it handles a missing config.json.',
  toolHandlers: {
    get_tutorial: (args) => readTutorial(args.slug),
  },
})

await followTurn(events)
```

A handler returns a string, or an object the helper turns into JSON. If it throws, the helper tells the agent the tool failed. Handlers run one at a time.

### Functions with side effects

`get_tutorial` only reads. A function that writes, like `open_pull_request` or `send_email`, needs more care. If your process crashes after running it but before sending the result, the agent will still wait for that result.

Store each result by session ID, turn ID and call ID before you send it. After a restart, retrieve the session, look at `required_actions`, and resend the saved result instead of running the function twice. I wrote more about this in [how to let an AI agent take irreversible actions safely](https://flaviocopes.com/ai-agent-irreversible-actions-safely/).

### Many tools: tool search

Every function definition takes space in the model's context, on every call. With a few functions that's fine. With fifty, it adds up.

Add a `tool_search` tool and mark the rarely used functions with `defer_loading: true`. The agent sees them only when it searches for a tool that matches the task:

```js
tools: [
  { type: 'tool_search' },
  getTutorial,
  { ...getPublishingStats, defer_loading: true },
]
```

`getPublishingStats` here stands for any other function definition you have. You still send the full definition. Tool search only changes when it reaches the model.

### Programmatic tool calling

There's one more thing the harness does with your tools, and it's on by default. The model can write a small JavaScript program that calls your tools in a loop or in parallel, and processes their results before they reach its context.

The program runs in an isolated V8 runtime with no Node.js, no network and no filesystem. It can only reach the outside world through your tools. Your functions still run in your application as usual.

This helps when a task needs many related calls, like "check all fifty tutorials in this list". If you want every call to go through the model instead, disable it:

```js
tools: [{ type: 'programmatic_tool_calling', enabled: false }, getTutorial]
```

## Step 9: check facts with web search

Our reviewer only knows what the model learned in training. Node.js changes, and a tutorial can be wrong because an API moved.

Web search is a built-in tool. Add it to `tools` and the agent can search when it needs to. Create `09-web-search.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const tutorial = await fs.readFile('tutorials/count-words.md', 'utf8')

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    tools: [
      {
        type: 'web_search',
        mode: 'live',
        allowed_domains: ['nodejs.org'],
      },
    ],
  },
  environment: { type: 'none' },
  input: `Check every Node.js API in this tutorial against the official Node.js docs. Link the doc page for each one.\n\n${tutorial}`,
  stream: true,
})

await followTurn(events)
```

The options:

- `mode`: `live` searches the web, and it's the default. `cached` searches saved content without going online. `disabled` turns search off.
- `allowed_domains`: up to 100 domains. Here the agent can only read the Node.js docs, which keeps random blog posts out of its sources.
- `context_size`: `low`, `medium` (the default) or `high`, for how much search content the model receives.
- `location`: country, region, city and timezone, for local searches.

If you leave `web_search` out of `tools`, the agent can't search, even if your prompt asks it to.

Search isn't free: it's billed per call, plus the search content tokens. We'll look at prices near the end.

## Step 10: connect an MCP server

Tutorials often live in a GitHub repository, and GitHub has an MCP server. MCP is a standard way to expose tools to an agent. If it's new to you, start with [what MCP is](https://flaviocopes.com/what-is-mcp/).

An MCP tool entry tells the agent where the server is and which of its tools it may use:

```js
{
  type: 'mcp',
  server_label: 'github',
  transport: { type: 'http', server_url: 'https://api.githubcopilot.com/mcp/' },
  allowed_tools: ['get_file_contents', 'search_issues'],
  required: true,
}
```

`allowed_tools` limits what the agent can discover and call. Our checker only needs to read files and search issues, so it can't touch anything else, even with a token that allows more.

`required: true` makes the turn fail if the server can't start. By default, MCP servers are optional, and the agent carries on without them.

By default OpenAI connects to the server from its own infrastructure, so the URL must be reachable from the internet. The other option, `connection_origin: 'environment'`, makes the connection from your sandbox instead. That's how you reach a server on localhost or inside a private network. There's also a `stdio` transport, which starts the MCP server as a process inside the sandbox.

### Keep the token in a vault

The GitHub MCP server needs a token. You could pass it inline with `transport.authorization`, but then your code handles it on every session. A **vault** stores it on OpenAI's side, and sessions refer to it by ID.

Create a GitHub token with read-only access to the repository, add it to `.env` as `GITHUB_TOKEN`, then create `10-vault.mjs`:

```js
import { client } from './agent.mjs'

const vault = await client.beta.agents.vaults.create({ name: 'Tutorial checker' })

await client.beta.agents.vaults.credentials.create(vault.id, {
  name: 'GitHub MCP token',
  auth: {
    type: 'static_bearer',
    token: process.env.GITHUB_TOKEN,
    mcp_server_url: 'https://api.githubcopilot.com/mcp/',
  },
})

console.log(vault.id)
```

Run it once. The credential matches the MCP server by URL. Retrieving the vault later never returns the secret.

Now the session uses both. Create `10-mcp.mjs`, and replace the repository with one of yours:

```js
import { client, followTurn } from './agent.mjs'

const repo = 'your-github-name/your-docs'

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    tools: [
      {
        type: 'mcp',
        server_label: 'github',
        transport: { type: 'http', server_url: 'https://api.githubcopilot.com/mcp/' },
        allowed_tools: ['get_file_contents', 'search_issues'],
        required: true,
      },
    ],
  },
  environment: { type: 'none' },
  vault_ids: ['vault_your_vault_id'],
  input: `Read docs/count-words.md from the ${repo} repository. Review it, then search the repository's open issues for reports about that tutorial.`,
  stream: true,
})

await followTurn(events)
```

MCP calls appear as `mcp_call` items in the session, with the server label, the tool name, the arguments and the output.

If you work with MCP a lot, I also have a post on [building your own MCP server](https://flaviocopes.com/build-mcp-server/) and one on [the MCP servers I use](https://flaviocopes.com/mcp-servers-i-use/).

## Step 11: run the code in a hosted sandbox

Reading code is one thing. Running it is how you really know it works. That needs an environment.

An `openai_hosted` environment is a sandbox OpenAI creates for the session. It starts in `/workspace`, with Python, Node.js and common command line tools. The agent can run commands, create files and install things.

We'll upload the tutorial plus a `notes.txt` file for its code to read. The text has a line break, which is enough to expose the counting bug. Create `11-sandbox.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const tutorial = await fs.readFile('tutorials/count-words.md')
const notes = 'The quick brown fox\njumps over the lazy dog\n'

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    instructions: 'You check programming tutorials. Run every code block exactly as written before you judge it.',
  },
  environment: {
    type: 'openai_hosted',
    network: { access: 'disabled' },
    files: [
      { type: 'inline', path: '/workspace/count-words.md', data: tutorial.toString('base64') },
      { type: 'inline', path: '/workspace/notes.txt', data: Buffer.from(notes).toString('base64') },
    ],
  },
  input: [
    'Save the code block from /workspace/count-words.md as count.mjs and run it with Node.js, like a reader would.',
    'The file has 9 words. Fix the tutorial until the script prints 9.',
    'Write the corrected tutorial to /workspace/outputs/count-words.md.',
  ].join('\n'),
  stream: true,
})

const { turn } = await followTurn(events)
console.log(`\nSession ID: ${turn.session_id}`)
```

Let's go through the environment:

- `files` puts files in the sandbox before the agent starts. Inline files are base64-encoded. For bigger files, upload them with the Files API and pass `{ type: 'file_id', path, file_id }` instead.
- `network: { access: 'disabled' }` cuts the sandbox off from the internet. Network access is **on** by default, so turn it off whenever the agent runs code you don't trust. Tutorial code counts.
- `/workspace/outputs` is special. We'll download whatever lands there in the next step.

This time our helper prints each command the agent runs, with its exit code. You should see a first `node count.mjs` fail with a `TypeError`, then the agent fix the import and try again. The exact commands depend on the run.

Because the prompt gives the expected output, the agent has a real test to pass instead of an opinion to give. Give an agent a way to check its own work whenever you can.

### Configure the sandbox

The environment accepts more setup:

```js
environment: {
  type: 'openai_hosted',
  packages: {
    npm: ['typescript'],
    python: ['requests'],
  },
  setup_commands: [{ command: 'mkdir -p /workspace/reports' }],
  env: { NODE_ENV: 'test' },
  network: {
    access: 'restricted',
    allowed_domains: ['registry.npmjs.org'],
  },
}
```

- `packages` installs global npm packages, Python packages and system packages. Pin versions when the tutorial depends on one.
- `setup_commands` run in order before the agent starts, in `/workspace` unless you set a `cwd`. The packages and files are ready before they run. If a command exits with an error, the agent doesn't start.
- `env` sets environment variables. You can't set `PATH`, `OPENAI_API_KEY` or anything starting with `CODEX_`.
- `network` with `restricted` access allows up to 100 exact hostnames. Subdomains and redirect targets need their own entries.

The sandbox size is `container_size` in the environment: `small` (1 vCPU, 1 GB of memory), `medium` (2 vCPUs, 4 GB, the default) or `large` (4 vCPUs, 16 GB). It's in the docs, but the JavaScript SDK types I checked in September 2026 don't list it yet, so TypeScript may complain if you pass it.

Setting up the sandbox takes a moment. Your environment has an ID in `session.environment.id`, and you can check its state with a `GET` request to `/v1/agents/environments/{id}`. It goes from `provisioning` to `connected`, or to `failed` with an error.

### How long the sandbox lives

Files stay in the sandbox across turns, and OpenAI keeps a connected sandbox alive between turns. If activity and those keep-alives stop for an hour, the sandbox can be deleted, and you can't change that timeout.

So treat the sandbox as scratch space, and move the results out as soon as a turn completes. When you're done with a session, delete it to clean up the sandbox. If the delete returns `409` because setup or a turn is still running, wait a bit and try again.

## Step 12: download the results

When a turn completes, every file under `/workspace/outputs` becomes an **artifact**. Artifacts are saved copies. They can't change, and they survive after the sandbox is gone.

Create `12-download.mjs`:

```js
import fs from 'node:fs/promises'
import path from 'node:path'
import { client } from './agent.mjs'

const sessionId = process.argv[2]
await fs.mkdir('reports', { recursive: true })

for await (const artifact of client.beta.agents.sessions.artifacts.list(sessionId)) {
  const response = await client.beta.agents.sessions.artifacts.content(artifact.id, {
    session_id: sessionId,
  })
  const target = path.join('reports', path.basename(artifact.path))
  await fs.writeFile(target, Buffer.from(await response.arrayBuffer()))
  console.log(`${artifact.path} (${artifact.size_bytes} bytes) -> ${target}`)
}
```

Run it with the session ID from step 11. You get the corrected tutorial in `reports/count-words.md`.

Each artifact records its path, its size and the turn that produced it. We only keep the file name with `path.basename()`, so a strange path from the sandbox can't write outside the `reports` folder.

A few limits worth knowing:

- An upload request takes up to 50 files, with 5 MiB per inline file and 10 MiB of inline data in total. Files API uploads go up to 50 MiB.
- An artifact can be up to 200 MiB, and all the outputs of a session up to 500 MiB.
- You download one artifact per request. For many files, ask the agent to zip them into one.
- Artifacts outlive the sandbox, but download what you need before you delete the session.
- Only hosted sandboxes publish artifacts. With a self-hosted environment, the files stay on your machine.

## Step 13: give the sandbox a secret

Some tutorials call an API that needs a key. Say a tutorial shows how to read your GitHub profile with `fetch()`. To run it, the sandbox needs a GitHub token.

Putting the token in `env` would work, but then any code in the sandbox can read it. That includes the tutorial's code, and code the agent writes after reading a malicious web page.

A vault credential of type `environment_variable` solves this. The sandbox gets a placeholder in the variable. When a request goes out over HTTPS to an allowed host, a proxy swaps the placeholder for the real token. Printing the variable inside the sandbox shows the placeholder.

```js
await client.beta.agents.vaults.credentials.create(vaultId, {
  name: 'GitHub API token',
  auth: {
    type: 'environment_variable',
    secret_name: 'GITHUB_TOKEN',
    secret_value: process.env.GITHUB_TOKEN,
    networking: { type: 'limited', allowed_hosts: ['api.github.com'] },
  },
})
```

Then attach the vault to a hosted session, and allow the same host in the sandbox network:

```js
const events = await client.beta.agents.sessions.create({
  agent: { model: 'gpt-6.1-sol' },
  vault_ids: [vaultId],
  environment: {
    type: 'openai_hosted',
    network: { access: 'restricted', allowed_domains: ['api.github.com'] },
  },
  input: 'Run: curl https://api.github.com/user -H "Authorization: Bearer $GITHUB_TOKEN" and summarize the account.',
  stream: true,
})
```

The two lists do different jobs. `allowed_domains` lets the sandbox connect to a host. `allowed_hosts` on the credential lets the proxy hand the secret to that host. You need both.

This only works for HTTPS requests to ports 443 and 8443, and only in hosted sandboxes. The placeholder can't be used for local work like signing a request. For that, keep the secret in your app and expose the operation as a function tool.

For the bigger picture of agents and credentials, see [how to give AI agents access to passwords](https://flaviocopes.com/ai-agent-passwords/).

## Step 14: reuse instructions with skills and plugins

Our checker's instructions will grow as we add rules for each language. Putting all of that in `instructions` makes every turn pay for it, even when a tutorial needs none of it.

**Skills** fix this. A skill is a folder with a `SKILL.md` file: a name, a description, and instructions. The harness adds only the name and description to the context. The agent reads the full file when a task matches. The free [AI agent skills course](https://flaviocopes.com/courses/ai-agent-skills/) covers how to write good ones.

To give the agent a skill, put the skill folder in the sandbox and list its parent folder in `environment.capability_directories`. When the sandbox is ready, the harness searches those folders for `SKILL.md` files.

With a hosted sandbox, we can upload the skill with the other files. Say we wrote `skills/check-node-tutorial/SKILL.md` locally:

```js
const skill = await fs.readFile('skills/check-node-tutorial/SKILL.md')

const environment = {
  type: 'openai_hosted',
  files: [
    {
      type: 'inline',
      path: '/workspace/skills/check-node-tutorial/SKILL.md',
      data: skill.toString('base64'),
    },
  ],
  capability_directories: ['/workspace/skills'],
}
```

The paths must be absolute, the folders must exist in the sandbox, and a session can list up to 32 of them. Hosted sandboxes also accept `environment.skills`, with skills you uploaded to OpenAI by ID or as a base64 ZIP archive.

Review every skill before you hand it to an agent. A skill is instructions the agent will follow, so a bad one is a prompt injection you installed yourself.

A **plugin** is a bigger bundle: a `.codex-plugin/plugin.json` manifest, an optional `.mcp.json` with MCP servers, and a `skills/` folder. Hosted sessions take one ZIP archive per plugin in `environment.plugins`.

Existing sessions don't reload plugins. If you change a plugin, create a new session to pick it up.

If you create many sessions with the same packages, files, skills and plugins, an **environment template** saves that setup once. You pass its ID as `environment_template_id`. Fields you leave out come from the template, and a session can't loosen the template's network rules.

## Step 15: check many tutorials with subagents

One tutorial at a time is slow. With subagents, the main agent can hand each tutorial to a helper that works in parallel.

Let's add a third tutorial, `tutorials/fetch-repo.md`. It calls the GitHub API and forgets an `await`:

````markdown
# Get a GitHub repository with fetch()

```js
const response = await fetch('https://api.github.com/repos/nodejs/node')
const repo = response.json()
console.log(repo.stargazers_count)
```
````

It prints `undefined`, because `response.json()` returns a promise.

Now create `15-subagents.mjs`:

```js
import fs from 'node:fs/promises'
import { client, followTurn } from './agent.mjs'

const names = await fs.readdir('tutorials')
const files = []

for (const name of names) {
  const data = await fs.readFile(`tutorials/${name}`)
  files.push({ type: 'inline', path: `/workspace/tutorials/${name}`, data: data.toString('base64') })
}

const events = await client.beta.agents.sessions.create({
  agent: {
    model: 'gpt-6.1-sol',
    multi_agent: { enabled: true, max_concurrent_subagents: 3 },
  },
  environment: {
    type: 'openai_hosted',
    network: { access: 'restricted', allowed_domains: ['api.github.com'] },
    files,
  },
  input: [
    'Check every tutorial in /workspace/tutorials.',
    'Start one subagent per tutorial. Each subagent runs the code blocks of its tutorial and reports what breaks.',
    'Combine the results into /workspace/outputs/report.md, one section per tutorial.',
  ].join('\n'),
  stream: true,
})

const { turn } = await followTurn(events)
console.log(`\nSession ID: ${turn.session_id}`)
```

`multi_agent.enabled` gives the main agent tools to create subagents, send them messages, wait for them and interrupt them. You don't define those tools. `max_concurrent_subagents` caps how many run at once, and the default is 6.

The subagents share the sandbox with the main agent. There's one filesystem, not one per subagent. That's great for reading the same files. If two agents edit the same file, they have to coordinate, so give each one its own folder or its own output file.

Subagents get the MCP servers, their credentials and the web search settings. They don't get your function tools.

In the stream, subagent turns have a `subagent_id`. That's why our helper skips their text and only returns when the main turn completes. To see which agent ran a command, retrieve the command's turn and look at its `subagent_id`: it's `null` for the main agent.

Use subagents for independent pieces of work. For short tasks, or steps that depend on each other, one agent is simpler and cheaper. Each subagent spends its own tokens.

## Step 16: run it in the background with webhooks

A full check can take a while. You don't want a script waiting with an open stream. A **webhook** lets OpenAI call your server when the session changes state.

The events you can subscribe to:

| Event | When it fires |
| --- | --- |
| `agent.session.created` | A session is created |
| `agent.session.action_required` | The session needs a function result, an environment connection or a computer use approval |
| `agent.session.in_progress` | A turn starts |
| `agent.session.idle` | The session is ready for more input |
| `agent.session.failed` | The session failed |

First, create a webhook endpoint in the OpenAI Platform dashboard, select the Agents API events, and save the signing secret as `OPENAI_WEBHOOK_SECRET` in `.env`. If webhooks are new to you, I explained [how webhooks work](https://flaviocopes.com/webhooks/) in another post.

Install Express:

```bash
npm install express
```

Then start the work without streaming. `16-start.mjs` takes a tutorial name, uploads that file, creates the session and exits:

```js
import fs from 'node:fs/promises'
import { client } from './agent.mjs'

const slug = process.argv[2]
const tutorial = await fs.readFile(`tutorials/${slug}.md`)

const session = await client.beta.agents.sessions.create({
  agent_id: 'agent_your_agent_id',
  environment: {
    type: 'openai_hosted',
    network: { access: 'disabled' },
    files: [{ type: 'inline', path: `/workspace/${slug}.md`, data: tutorial.toString('base64') }],
  },
  metadata: { tutorial: slug },
  input: `Run every code block in /workspace/${slug}.md. Write what breaks and why to /workspace/outputs/${slug}-report.md.`,
})

console.log(`Started ${session.id}`)
```

Run it with `node --env-file=.env 16-start.mjs count-words`. The `metadata` travels with the session, so the webhook handler knows which tutorial a session belongs to.

And `16-webhooks.mjs` receives the events:

```js
import express from 'express'
import OpenAI from 'openai'

const client = new OpenAI({ webhookSecret: process.env.OPENAI_WEBHOOK_SECRET })
const app = express()

app.post('/webhooks/openai', express.raw({ type: 'application/json' }), async (request, response) => {
  const payload = request.body.toString('utf8')

  try {
    await client.webhooks.verifySignature(payload, request.headers)
  } catch {
    response.status(400).send('Invalid signature')
    return
  }

  const event = JSON.parse(payload)
  response.sendStatus(200)

  if (event.type === 'agent.session.idle') {
    const session = await client.beta.agents.sessions.retrieve(event.data.id)
    const page = await client.beta.agents.sessions.turns.list(session.id, { order: 'desc', limit: 1 })
    const lastTurn = page.data[0]
    console.log(`${session.metadata?.tutorial}: last turn ${lastTurn?.status}`)
  }

  if (event.type === 'agent.session.action_required') {
    console.log(`${event.data.id} is waiting for ${event.data.required_action.type}`)
  }
})

app.listen(8000, () => console.log('Listening on http://localhost:8000'))
```

The handler reads the raw body, because the signature is computed over the exact bytes OpenAI sent. `verifySignature()` rejects anything that didn't come from OpenAI. We parse the JSON ourselves, because the SDK's webhook event types didn't include the agent events yet when I checked in September 2026.

We answer `200` right away and do the work afterwards. In production you'd put slow work in a queue instead.

Look at the `idle` branch. `idle` only means the session is ready for more input. It doesn't say the last turn worked. So we fetch the latest turn and read its status. From there, the next step would be downloading the artifacts like in step 12.

The `action_required` payload only says what kind of action is pending. Retrieve the session to get the details, like the function arguments.

Two gaps to know about: deleting a session doesn't send a webhook, and it doesn't stop compute you started at a sandbox provider.

To test webhooks from your laptop, OpenAI needs a public URL. A tunnel gives you one in a few seconds. I explained how in my guide to [Cloudflare Quick Tunnels](https://flaviocopes.com/cloudflare-quick-tunnels/), which uses a webhook as its first example.

## Self-hosted sandboxes

Sometimes the code must run on your own machines. The data can't leave your network, or you need a GPU, or the project needs a special setup.

A `self_hosted` environment keeps the harness at OpenAI and moves the execution to you. You create the session:

```js
const session = await client.beta.agents.sessions.create({
  agent: { model: 'gpt-6.1-sol' },
  environment: {
    type: 'self_hosted',
    workspace_directory: '/workspace',
  },
})

console.log(session.environment.id, session.environment.remote_url)
```

Then, inside the machine, you run the executor that ships with the Codex CLI. The docs install it from the alpha release:

```bash
npm install -g @openai/codex@alpha

CODEX_API_KEY="$OPENAI_ENVIRONMENT_KEY" \
codex exec-server \
  --remote "<session.environment.remote_url>" \
  --environment-id "<session.environment.id>"
```

`CODEX_API_KEY` must be an **environment key**, created in the Agents section of the OpenAI Platform. It can only connect environments and can't do anything else with your account. Agent-generated code can read it, so never use your application key there.

The machine needs outbound access to `api.openai.com` and `wss://codex-cloud-environments.chatgpt.com`. When a message needs the executor and it's not connected, the API waits up to five minutes before the request fails. With webhooks, you get an `action_required` event of type `environment_connection`, which is your cue to start a machine.

Each session needs its own executor. You can reuse the machine image and the workspace directory across sessions.

If you don't want to run the machine yourself, OpenAI lists sandbox providers with setup guides: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop, Vercel and AWS. If you already use one of them, like [Vercel Sandbox](https://flaviocopes.com/vercel-sandbox/), that's the shortest path.

## Computer use

The agent can also use a browser in a hosted sandbox. You enable the desktop and add the tool:

```js
agent: {
  model: 'gpt-6.1-sol',
  tools: [{ type: 'computer_use' }],
},
environment: {
  type: 'openai_hosted',
  desktop: { enabled: true },
}
```

Every new website needs your approval. The session goes into `requires_action` with a `computer_use_approval_request`, and you answer with an event:

```js
await client.beta.agents.sessions.events.create(sessionId, {
  events: [
    {
      type: 'agent.session.input.computer_use_approval_request_result',
      request_id: requestId,
      response: { type: 'browser_origin_access', decision: 'approve' },
    },
  ],
})
```

The decision can be `approve`, `deny` or `cancel`. Sign-in pages arrive as a separate `browser_authentication` request, where you can submit credentials.

Approving a site lets the agent use it. It doesn't ask again before each click, so approve only sites where every action is safe.

For our checker, this could open a tutorial's demo page and confirm the button does what the text promises. I'd keep it for that kind of narrow job.

## Watch what the agent did

The OpenAI Platform shows your sessions at [platform.openai.com/logs?api=agents](https://platform.openai.com/logs?api=agents), in the **Agents** tab. Search a session by ID to see its turns, tool calls and subagents.

Tracing is on by default for new sessions. A trace shows the model responses, the tool calls and what each subagent did, with the token usage of each agent. Traces are built after a turn ends, so the answer can arrive before its trace. You can also download traces as OpenTelemetry (OTLP) JSON and open them in another tracing tool.

The session items from step 4 are the programmatic version of the same information.

## Errors and retries

Errors come in two places: HTTP errors when a request fails, and errors inside a turn or a session.

The turn error has a `code` that tells you what to do:

- `rate_limit_exceeded`, `server_overloaded`, `request_timeout`, `connection_failed`: temporary. Retry after a delay.
- `flex_unavailable`: the cheaper flex tier has no capacity. Retry later, or switch the session to `service_tier: 'default'`.
- `context_length_exceeded`: the conversation doesn't fit the model's context anymore. Start a new session with a short summary of where you were.
- `session_budget_exceeded`: the session used up its budget. Start a new session to continue.
- `credit_balance_exhausted`, `usage_limit_exceeded`: a billing problem. Retrying faster won't help.
- `invalid_request`, `authentication_error`, `resource_not_found`: fix the request or the model name first.

When you retry, check before you repeat. A failed turn may already have changed files or called your functions. The recovery OpenAI recommends is:

1. Retrieve the session, the turn and the saved items. If the turn is still running, keep following it. If it completed, use the result.
2. Check what the failed turn already did.
3. Wait, honoring `Retry-After` when the response has it, with a growing delay and a maximum number of attempts.
4. Once the session is idle, send a follow-up asking it to continue only the unfinished work.

The SDK already retries some failed HTTP requests on its own. The loop above is for turns.

## What it costs

There's no extra fee for the Agents API itself. You pay for the tokens, the tools and the sandbox time your agents use.

These are the standard prices per million tokens for the two models in this guide, checked on September 30, 2026:

| Model | Input | Cached input | Output |
| --- | --- | --- | --- |
| `gpt-6-astra` | $10.00 | $1.00 | $50.00 |
| `gpt-6.1-sol` | $2.00 | $0.10 | $10.00 |

Built-in tools and sandboxes, checked the same day:

- Web search: $10 per 1,000 calls, plus the search content tokens at the model's rate.
- Hosted sandboxes use the container rates: $0.03 for 1 GB, $0.12 for 4 GB and $0.48 for 16 GB, per 20-minute session. Eligible sessions are billed by the minute, with a 5-minute minimum.

Agents make many model calls per task, and each call sends the conversation so far. That's where the money goes. Prompt caching helps, because calls that share the same prefix pay the cached rate for it. Keep instructions and tool definitions stable, and put new details in follow-up messages instead of changing the instructions.

A high cache rate is not the same as a cheap task, though. Cached tokens still cost something, and a long session keeps sending a long history. When I compare setups, I'd compare the cost of finishing the same task, not the cache percentage.

When OpenAI announced hosted sandboxes, their own staff warned in the forum to calculate container costs before spinning up many of them. Take that advice. Test with one session, look at its usage and the billing dashboard, then scale.

## Security checklist

An agent with a sandbox runs code that nobody reviewed. Plan for that code to be hostile, because sometimes the input is.

- **Keep your application key out of the sandbox.** Code the agent writes can read anything in its environment, including environment keys.
- **Turn the network off, or restrict it.** Hosted sandboxes have network access by default. Allow only the hosts a task needs.
- **Use vault credentials for secrets**, so the sandbox sees a placeholder and not the token.
- **Restrict MCP tools** with `allowed_tools`, and use read-only tokens where you can.
- **Validate function arguments** like you'd validate a form field.
- **Separate workloads.** Use different sessions for different users, and a dedicated OpenAI project for the application.
- **Guard irreversible actions.** A function that deletes, pays or publishes needs a human or a hard rule in front of it.

I also wrote about [running coding agents unattended](https://flaviocopes.com/agents-unattended/), which covers the same risks from the other side.

## Limits and when not to use it

The Agents API is a beta, and when I checked in September 2026 it had real limits:

- Data residency is United States only.
- It doesn't support Zero Data Retention, even with a self-hosted sandbox. If your contract requires ZDR, this API is out for now.
- A hosted sandbox can disappear after an hour without activity, taking unsaved files with it.
- The event stream doesn't replay. You need the recovery flow from step 4 in any serious app.
- It's a beta, so names and fields can change. Some fields reached the SDK types before the guides, and `container_size` is in the guides but not in the types yet.

And there are jobs it doesn't fit:

- **A single model answer.** If you ask one question and get one reply, the Responses API is simpler and has fewer moving parts.
- **Chat inside your app where you want full control** of the loop, the storage and the approvals. That's what the Agents SDK is for.
- **Coding on your own laptop.** A local coding agent sees your real files, your tools and your dev server. For that I use a local agent, not an API.

## How I would use it

I haven't put the Agents API into anything I run yet. But I know exactly where I'd start, and it's the example of this guide.

This site has more than 1,500 tutorials, and hundreds of them date back to 2018. Code that old breaks quietly when an API changes or a package moves.

I'd build the checker as a nightly job. A small script picks a handful of old tutorials, starts one hosted session per tutorial with the network off, and asks the agent to run every code block with a recent Node.js. The webhook lands in a Cloudflare Pages Function, like the Paddle webhook that already delivers course access on this site. When a session goes idle, the function downloads the report and adds the tutorial to a list I review by hand.

The agent wouldn't change the tutorials. It would tell me which ones are broken and why. I'd keep the fixing in my own hands.

The parts of the API that make this work are the ones you can't easily build yourself: the sandbox, the saved session and the webhooks. If I only needed a model answer, one model call would be enough. The [app idea generator](https://flaviocopes.com/tools/app-idea-generator/) on this site works that way, with a small open model on Cloudflare Workers AI.

What I wouldn't do is move my day-to-day coding here. When I work on this site I want the agent on my machine, next to my dev server, one task at a time, where I can watch it. The Agents API is for work that should happen without me watching.

## Where to go next

The official docs are the best reference, and they change often during the beta:

- The [Agents API overview](https://developers.openai.com/api/docs/guides/agents-api/overview) and the [quickstart](https://developers.openai.com/api/docs/guides/agents-api/quickstart).
- [Events and items](https://developers.openai.com/api/docs/guides/agents-api/sessions/events), for everything the stream can send.
- [OpenAI-hosted sandboxes](https://developers.openai.com/api/docs/guides/agents-api/environments/openai-hosted) and [files and artifacts](https://developers.openai.com/api/docs/guides/agents-api/environments/files).
- The [API reference](https://developers.openai.com/api/reference/resources/beta/subresources/agents), for every field.
- The [Agents API examples in the OpenAI Cookbook](https://github.com/openai/openai-cookbook/tree/main/examples/agents_api).
- The [launch announcement in the OpenAI developer forum](https://community.openai.com/t/introducing-the-agents-api-and-hosted-sandboxes/1396481).

If you want to stream an agent's answer to a browser, [Server-Sent Events](https://flaviocopes.com/server-sent-events/) is the simplest way, and I showed it with a model in [streaming LLM responses with SSE](https://flaviocopes.com/streaming-llm-responses-sse/).
