It is tempting to summarize this layer as “the SDK wraps the agent.” The source suggests a narrower and more useful reading. SDKs and app-server are adapters: they translate external calls into the thread, turn, setting, event, and notification data structures already used by the Codex runtime. They do not reimplement conversation history, permission checks, or rollout.

The core idea of this chapter is: the public interface is a runtime adapter, not a second agent. app-server turns JSON-RPC requests into typed ClientRequest values, then routes them to request processors. SDKs wrap that protocol into friendlier calls such as Thread.run(), runStreamed(), or CLI JSONL events.

Reading goal. This chapter tracks only external entry: transport, initialization, typed requests, thread and turn entry points, notification fanout, and SDK ergonomics. After reading it, you should be able to separate the app-server public API, the core runtime execution, and the convenience layer added by each SDK.

Evidence boundary. Source snapshot: 5c5308fc9a9e. Transport and connection lifecycle follow the official App Server documentation; concrete fields, dispatch, processors, listener conversion, and SDK routing follow the pinned public snapshot. Documentation can lag the implementation: the current request enum has removed thread/rollback, while historical markers retain replay support. Hosted services and private backends are outside this analysis.

1. Start by splitting “embed Codex”

Embedding requires a conversation, an execution, and a stream of events. Official documentation calls these Thread, Turn, and Item. App-server exchanges JSON-RPC messages over stdio JSONL, with WebSocket, Unix-socket and in-process entry points also visible in source. A turn/start response confirms input submission; completion notifications deliver the final execution state.

External goal app-server shape Runtime module Why a prompt API is too small
Open a conversation. thread/start with cwd, model, permissions, environments, and config. ThreadManager creates a CodexThread. The runtime must establish identity, authority, rollout, and listeners.
Send user input. turn/start with input, overrides, and additional context. CodexThread.start_or_steer_turn(...) asks core to start or steer. The input affects model view, tools, usage, events, and persistence.
Render progress. turn/started, item/*, and turn/completed notifications. The listener maps core EventMsg into ServerNotification. Progress is a stream of facts, not only a final response.
Resume or fork. thread/resume, thread/fork. The rollout and recovery records from Part X. Continuation depends on replayable evidence, not only visible messages.

This preserves the same semantics across IDEs, SDKs, CLI commands, and in-process hosts can differ in ergonomics. Once they enter the runtime, they should converge on the same typed requests, the same thread and turn implementations, and the same event stream.

The smallest replayable shape is a sequence of requests and notifications, not one long answer. This abbreviated JSONL keeps only the fields needed to connect the documented primitives to the source modules below:

This is a simplified new-turn JSONL shape. Ordinary responses, timestamps and optional fields are omitted; t1 stands for the id returned by thread/start. Notifications and responses may interleave.

{"id":1,"method":"initialize","params":{"clientInfo":{"name":"example_client","version":"1.0.0"}}}
{"method":"initialized"}
{"id":2,"method":"thread/start","params":{}}
{"id":3,"method":"turn/start","params":{"threadId":"t1","input":[{"type":"text","text":"fix failing test","text_elements":[]}]}}
{"method":"turn/started","params":{"threadId":"t1","turn":{"id":"u1","status":"inProgress","items":[],"itemsView":"notLoaded","error":null}}}
{"method":"item/completed","params":{"threadId":"t1","turnId":"u1","item":{"type":"agentMessage","id":"m1","text":"Fixed the test."}}}
{"method":"turn/completed","params":{"threadId":"t1","turn":{"id":"u1","status":"completed","items":[],"itemsView":"notLoaded","error":null}}}

The useful part is the order. A connection is initialized first. A thread then establishes identity and runtime settings. A turn carries one user input. Runtime progress returns through notifications. SDK convenience methods wrap this sequence; they do not replace the lifecycle.

2. app-server turns a transport into an initialized session

The first check is not turn/start; it is initialize. The official lifecycle documentation says each transport connection must send one initialize request, then an initialized notification. Requests issued before that handshake are rejected. This gives the server a per-connection client identity, capability set, and notification opt-out configuration.

The source follows that order. process_request deserializes JSON-RPC into a typed ClientRequest. In-process embedders can bypass JSON deserialization through process_client_request, but the comment says that it preserves identical semantics by delegating to handle_client_request. That shared handler treats Initialize specially, then sends all other requests through initialized dispatch.

public edge:
JSON-RPC request
  -> serde ClientRequest
  -> initialize check
  -> experimental API check
  -> serialization scope
  -> request processor

Initialized requests also pass through serialization scopes. The protocol macro generates ClientRequest.serialization_scope(); app-server maps those scopes to global, thread, path, command, process, filesystem watch, or OAuth queues. Shared reads can proceed together; mutating work uses exclusive access. The public edge therefore protects request ordering before core state is touched.

3. thread/start creates a runtime container

ThreadStartParams carries model, provider, service tier, cwd, workspace roots, approval policy, sandbox, permissions, instructions, personality, environments, dynamic tools, and capability roots. Read together, these fields show that thread/start is not only “create a chat.” It establishes the baseline that later turns will inherit.

When message_processor sees ClientRequest::ThreadStart, it enters thread_processor.thread_start. The processor rejects incompatible permissions and sandbox fields, parses environments, builds config overrides, then spawns thread_start_task. The task ultimately calls ThreadManager.start_thread_with_options(StartThreadOptions { ... }), passing initial history, dynamic tools, service naming, tracing, environment selection, and extension initialization into core.

The response is only half the startup story. app-server also auto-attaches a conversation listener, updates thread watch state, sends ThreadStartResponse, and then emits thread/started. The caller is subscribed to the event stream from the beginning.

4. turn/start asks core to start or steer

Once a thread exists, turn/start is the entry point that actually drives the agent. TurnStartParams carries the target thread, input items, optional Responses API metadata, additional context, environment and cwd overrides, workspace roots, approval and sandbox choices, permission profile, model and reasoning settings, output schema, personality, and collaboration mode. A turn is therefore not just appended text; it may also update the runtime settings used by following turns.

turn_start_inner loads the thread, checks direct input and provider configuration, then validates input size. Normal input becomes TurnInput::UserInput; a separate toolOutput path accepts a named standalone tool result and rejects nonempty input alongside it. The settings builder still rejects conflicting permissions and sandboxPolicy. It then creates a TurnInputRequest, calls start_or_steer_turn, and waits for core’s routing decision.

turn/start shape:
input + overrides + additionalContext
  -> TurnInput + ThreadSettingsOverrides
  -> TurnInputRequest
  -> CodexThread.start_or_steer_turn()
  -> Started | Steered | NotSubmitted

The result distinguishes Started, Steered, and NotSubmitted. The first two return an InProgress turn with itemsView = notLoaded; steering returns the existing turn id. Declined input returns an error, including when the server is draining. Success confirms acceptance, not completion: notifications still report completion, failure, or interruption.

5. Progress returns as notifications

A running turn may emit model deltas, tool activity, permission requests, plan updates, token usage, thread status changes, and completion. A final response object would be too late for interactive clients and too narrow for approval or interrupt flows.

App-server relays core events to TUI and other clients, while rollout supports later turn and item reconstruction
Live client notifications and later reconstruction from rollout follow separate paths.

The listener performs this routing. ensure_conversation_listener gets the thread from ThreadManager, subscribes the connection to thread state, and starts a listener task. The task reads conversation.next_event(), updates thread-local state, finds subscribed connections, and calls apply_bespoke_event_handling. That handler turns EventMsg::TurnStarted into ServerNotification::TurnStarted, item deltas into server notifications through item_event_to_server_notification, and completion into ServerNotification::TurnCompleted.

One reader must consume stdio stdout, which mixes JSON-RPC responses and notifications. Python’s MessageRouter routes responses by request id and gives each turn consumer an independent cursor. Multiple handles pointing at the same turn do not steal one another’s events. It captures cursors before sending a request and buffers early events; only prefixes consumed by every subscriber are pruned. A fresh independent subscription starts at the next event, not the beginning of turn history. Transport failure wakes all blocked waiters.

6. Python SDK wraps app-server stdio

The Python SDK follows app-server most closely. The CodexClient docstring describes a typed JSON-RPC client over stdio. Startup resolves a Codex executable, runs codex app-server --listen stdio://, opens stdin/stdout/stderr, and starts reader and stderr-drain threads. Initialization sends initialize with clientInfo and capabilities.experimentalApi, then sends initialized.

The low-level client exposes app-server methods such as thread_start, thread_resume, thread_fork, turn_start, turn_interrupt, and turn_steer. The high-level Codex object turns those into application ergonomics: construction starts and initializes the runtime; thread_start() returns a Thread; Thread.turn() builds TurnStartParams and returns a TurnHandle; Thread.run() consumes the stream and collects a TurnResult from completed items, usage, and turn/completed.

SDK layer Caller sees Still backed by Why this layer exists
CodexClient Typed requests, notifications, wait and stream helpers. stdio JSON-RPC plus app-server methods. One stdout stream is not consumed by competing readers.
Codex Login, account, and thread methods. A runtime connection initialized during construction. Callers do not hand-write the handshake.
Thread run(), turn(), read(). turn/start plus notification stream. Synchronous run collects events without changing runtime semantics.
AsyncCodexClient Async thread, turn, and stream helpers. The synchronous client executed in worker threads. Blocking reads do not monopolize the event loop.

7. TypeScript SDK wraps a lighter CLI route

The TypeScript SDK README is explicit: it wraps the codex CLI from @openai/codex, spawning the CLI and exchanging JSONL events over stdin/stdout. That is a different API from Python. Python sits next to app-server v2; the TypeScript SDK currently drives codex exec --experimental-json.

CodexExec.run builds that command: it starts with exec --experimental-json, then appends config overrides, model, sandbox, working directory, additional directories, output schema, reasoning effort, network access, web search, approval policy, and optionally resume <threadId>. It spawns the child process, writes input to stdin, reads stdout line by line, and yields each JSONL line.

Thread.runStreamedInternal normalizes caller input into prompt text and images, calls CodexExec.run, parses each line as a ThreadEvent, stores the thread id when it sees thread.started, and yields events. run() then buffers item.completed and turn.completed into a final result. The API is simpler, but it also exposes fewer operations than app-server v2.

8. In-process hosting still reuses app-server semantics

in_process.rs is useful because its file-level comment states the design plainly. The module replaces socket/stdio transports with bounded in-memory channels while still running the existing MessageProcessor and outbound routing. Incoming requests are typed ClientRequest values, but responses still return through the same JSON-RPC result envelope used by stdio and websocket transports.

start() completes initialize and initialized before returning a handle. Later requests still enter process_client_request, and initialized state, experimental capability, and notification opt-outs are mirrored to outbound state. Even inside one process, Codex keeps the app-server semantics instead of creating a hidden shortcut around initialization and request ordering.

9. What to carry forward

The external integration path is now complete. Part I followed one request into the runtime. The middle chapters inspected context, tools, permissions, client event conversion, extensions, hooks, cache discipline, and rollout recovery. This chapter shows how external clients enter that same runtime through app-server and SDK adapters. The next chapter returns to long-lived state: how stable lessons from old rollouts become memory for future threads.

Integration need Question to ask Mechanism Do not read it as
IDE or rich client. Do you need all thread, turn, event, and control operations? codex app-server plus JSON-RPC notifications. A synchronous answer API.
Python automation. Do you need app-server v2 methods and typed notifications? Python SDK over app-server stdio. Direct chat-text read/write.
Node or TypeScript task runner. Is CLI JSONL enough for the workflow? TypeScript SDK over codex exec --experimental-json. A full app-server client.
Same-process host. Can it preserve app-server semantics? in_process typed requests plus the same result envelope. A bypass around initialization, queues, and notifications.

A public interface for an agent runtime should not be only a convenient mouthpiece for the model. It should bring external callers into the runtime’s existing thread, turn, permission, rollout, and event mechanisms. The more the runtime can do, the more valuable that shared implementation becomes.

Sources