It is tempting to summarize this layer as “the SDK wraps the agent.” The source suggests a narrower and more useful reading. SDKs and app-server are adapters: they translate external calls into the thread, turn, setting, event, and notification data structures already used by the Codex runtime. They do not reimplement conversation history, permission checks, or rollout.
The core idea of this chapter is: the public interface
is a runtime adapter, not a second agent.
app-server turns JSON-RPC requests into typed
ClientRequest values, then routes them to request
processors. SDKs wrap that protocol into friendlier calls such
as Thread.run(), runStreamed(), or CLI
JSONL events.
Reading goal. This chapter tracks only external entry: transport, initialization, typed requests, thread and turn entry points, notification fanout, and SDK ergonomics. After reading it, you should be able to separate the app-server public API, the core runtime execution, and the convenience layer added by each SDK.
Evidence boundary. Source snapshot: 5c5308fc9a9e. Transport and connection lifecycle follow the official App Server documentation; concrete fields, dispatch, processors, listener conversion, and SDK routing follow the pinned public snapshot. Documentation can lag the implementation: the current request enum has removed thread/rollback, while historical markers retain replay support. Hosted services and private backends are outside this analysis.
1. Start by splitting “embed Codex”
Embedding requires a conversation, an execution, and a stream of events. Official documentation calls these Thread, Turn, and Item. App-server exchanges JSON-RPC messages over stdio JSONL, with WebSocket, Unix-socket and in-process entry points also visible in source. A turn/start response confirms input submission; completion notifications deliver the final execution state.
| External goal | app-server shape | Runtime module | Why a prompt API is too small |
|---|---|---|---|
| Open a conversation. | thread/start with cwd, model, permissions, environments, and config. |
ThreadManager creates a CodexThread. |
The runtime must establish identity, authority, rollout, and listeners. |
| Send user input. | turn/start with input, overrides, and additional context. |
CodexThread.start_or_steer_turn(...) asks core to start or steer. |
The input affects model view, tools, usage, events, and persistence. |
| Render progress. | turn/started, item/*, and turn/completed notifications. |
The listener maps core EventMsg into ServerNotification. |
Progress is a stream of facts, not only a final response. |
| Resume or fork. | thread/resume, thread/fork. |
The rollout and recovery records from Part X. | Continuation depends on replayable evidence, not only visible messages. |
This preserves the same semantics across IDEs, SDKs, CLI commands, and in-process hosts can differ in ergonomics. Once they enter the runtime, they should converge on the same typed requests, the same thread and turn implementations, and the same event stream.
The smallest replayable shape is a sequence of requests and notifications, not one long answer. This abbreviated JSONL keeps only the fields needed to connect the documented primitives to the source modules below:
This is a simplified new-turn JSONL shape. Ordinary responses, timestamps and optional fields are omitted; t1 stands for the id returned by thread/start. Notifications and responses may interleave.
{"id":1,"method":"initialize","params":{"clientInfo":{"name":"example_client","version":"1.0.0"}}}
{"method":"initialized"}
{"id":2,"method":"thread/start","params":{}}
{"id":3,"method":"turn/start","params":{"threadId":"t1","input":[{"type":"text","text":"fix failing test","text_elements":[]}]}}
{"method":"turn/started","params":{"threadId":"t1","turn":{"id":"u1","status":"inProgress","items":[],"itemsView":"notLoaded","error":null}}}
{"method":"item/completed","params":{"threadId":"t1","turnId":"u1","item":{"type":"agentMessage","id":"m1","text":"Fixed the test."}}}
{"method":"turn/completed","params":{"threadId":"t1","turn":{"id":"u1","status":"completed","items":[],"itemsView":"notLoaded","error":null}}}
The useful part is the order. A connection is initialized first. A thread then establishes identity and runtime settings. A turn carries one user input. Runtime progress returns through notifications. SDK convenience methods wrap this sequence; they do not replace the lifecycle.
2. app-server turns a transport into an initialized session
The first check is not turn/start; it is
initialize. The official lifecycle documentation says
each transport connection must send one initialize
request, then an initialized notification. Requests
issued before that handshake are rejected. This gives the server a
per-connection client identity, capability set, and notification
opt-out configuration.
The source follows that order. process_request
deserializes JSON-RPC into a typed ClientRequest.
In-process embedders can bypass JSON deserialization through
process_client_request, but the comment says that it
preserves identical semantics by delegating to
handle_client_request. That shared handler treats
Initialize specially, then sends all other requests
through initialized dispatch.
public edge:
JSON-RPC request
-> serde ClientRequest
-> initialize check
-> experimental API check
-> serialization scope
-> request processor
Initialized requests also pass through serialization scopes. The
protocol macro generates ClientRequest.serialization_scope();
app-server maps those scopes to global, thread, path, command,
process, filesystem watch, or OAuth queues. Shared reads can
proceed together; mutating work uses exclusive access. The public
edge therefore protects request ordering before core state is
touched.
3. thread/start creates a runtime container
ThreadStartParams carries model, provider, service
tier, cwd, workspace roots, approval policy, sandbox, permissions,
instructions, personality, environments, dynamic tools, and
capability roots. Read together, these fields show that
thread/start is not only “create a chat.” It
establishes the baseline that later turns will inherit.
When message_processor sees
ClientRequest::ThreadStart, it enters
thread_processor.thread_start. The processor rejects
incompatible permissions and sandbox
fields, parses environments, builds config overrides, then spawns
thread_start_task. The task ultimately calls
ThreadManager.start_thread_with_options(StartThreadOptions { ... }),
passing initial history, dynamic tools, service naming, tracing,
environment selection, and extension initialization into core.
The response is only half the startup story. app-server also
auto-attaches a conversation listener, updates thread watch
state, sends ThreadStartResponse, and then emits
thread/started. The caller is subscribed to the
event stream from the beginning.
4. turn/start asks core to start or steer
Once a thread exists, turn/start is the entry point
that actually drives the agent. TurnStartParams
carries the target thread, input items, optional Responses API
metadata, additional context, environment and cwd overrides,
workspace roots, approval and sandbox choices, permission profile,
model and reasoning settings, output schema, personality, and
collaboration mode. A turn is therefore not just appended text; it
may also update the runtime settings used by following turns.
turn_start_inner loads the thread, checks direct input and provider configuration, then validates input size. Normal input becomes TurnInput::UserInput; a separate toolOutput path accepts a named standalone tool result and rejects nonempty input alongside it. The settings builder still rejects conflicting permissions and sandboxPolicy. It then creates a TurnInputRequest, calls start_or_steer_turn, and waits for core’s routing decision.
turn/start shape:
input + overrides + additionalContext
-> TurnInput + ThreadSettingsOverrides
-> TurnInputRequest
-> CodexThread.start_or_steer_turn()
-> Started | Steered | NotSubmitted
The result distinguishes Started, Steered, and NotSubmitted. The first two return an InProgress turn with itemsView = notLoaded; steering returns the existing turn id. Declined input returns an error, including when the server is draining. Success confirms acceptance, not completion: notifications still report completion, failure, or interruption.
5. Progress returns as notifications
A running turn may emit model deltas, tool activity, permission requests, plan updates, token usage, thread status changes, and completion. A final response object would be too late for interactive clients and too narrow for approval or interrupt flows.
The listener performs this routing. ensure_conversation_listener
gets the thread from ThreadManager, subscribes the
connection to thread state, and starts a listener task. The task
reads conversation.next_event(), updates thread-local
state, finds subscribed connections, and calls
apply_bespoke_event_handling. That handler turns
EventMsg::TurnStarted into
ServerNotification::TurnStarted, item deltas into
server notifications through item_event_to_server_notification,
and completion into ServerNotification::TurnCompleted.
One reader must consume stdio stdout, which mixes JSON-RPC responses and notifications. Python’s MessageRouter routes responses by request id and gives each turn consumer an independent cursor. Multiple handles pointing at the same turn do not steal one another’s events. It captures cursors before sending a request and buffers early events; only prefixes consumed by every subscriber are pruned. A fresh independent subscription starts at the next event, not the beginning of turn history. Transport failure wakes all blocked waiters.
6. Python SDK wraps app-server stdio
The Python SDK follows app-server most closely. The
CodexClient docstring describes a typed JSON-RPC
client over stdio. Startup resolves a Codex executable, runs
codex app-server --listen stdio://, opens
stdin/stdout/stderr, and starts reader and stderr-drain threads.
Initialization sends initialize with
clientInfo and capabilities.experimentalApi,
then sends initialized.
The low-level client exposes app-server methods such as
thread_start, thread_resume,
thread_fork, turn_start,
turn_interrupt, and turn_steer. The
high-level Codex object turns those into application
ergonomics: construction starts and initializes the runtime;
thread_start() returns a Thread;
Thread.turn() builds TurnStartParams and
returns a TurnHandle; Thread.run()
consumes the stream and collects a TurnResult from
completed items, usage, and turn/completed.
| SDK layer | Caller sees | Still backed by | Why this layer exists |
|---|---|---|---|
CodexClient |
Typed requests, notifications, wait and stream helpers. | stdio JSON-RPC plus app-server methods. | One stdout stream is not consumed by competing readers. |
Codex |
Login, account, and thread methods. | A runtime connection initialized during construction. | Callers do not hand-write the handshake. |
Thread |
run(), turn(), read(). |
turn/start plus notification stream. |
Synchronous run collects events without changing runtime semantics. |
AsyncCodexClient |
Async thread, turn, and stream helpers. | The synchronous client executed in worker threads. | Blocking reads do not monopolize the event loop. |
7. TypeScript SDK wraps a lighter CLI route
The TypeScript SDK README is explicit: it wraps the
codex CLI from @openai/codex, spawning
the CLI and exchanging JSONL events over stdin/stdout. That is a
different API from Python. Python sits next to app-server v2;
the TypeScript SDK currently drives codex exec --experimental-json.
CodexExec.run builds that command: it starts with
exec --experimental-json, then appends config
overrides, model, sandbox, working directory, additional
directories, output schema, reasoning effort, network access, web
search, approval policy, and optionally resume <threadId>.
It spawns the child process, writes input to stdin, reads stdout
line by line, and yields each JSONL line.
Thread.runStreamedInternal normalizes caller input
into prompt text and images, calls CodexExec.run,
parses each line as a ThreadEvent, stores the thread
id when it sees thread.started, and yields events.
run() then buffers item.completed and
turn.completed into a final result. The API is
simpler, but it also exposes fewer operations than app-server
v2.
8. In-process hosting still reuses app-server semantics
in_process.rs is useful because its file-level
comment states the design plainly. The module replaces
socket/stdio transports with bounded in-memory channels while
still running the existing MessageProcessor and
outbound routing. Incoming requests are typed
ClientRequest values, but responses still return
through the same JSON-RPC result envelope used by stdio and
websocket transports.
start() completes initialize and
initialized before returning a handle. Later requests
still enter process_client_request, and initialized
state, experimental capability, and notification opt-outs are
mirrored to outbound state. Even inside one process, Codex keeps
the app-server semantics instead of creating a hidden shortcut
around initialization and request ordering.
9. What to carry forward
The external integration path is now complete. Part I followed one request into the runtime. The middle chapters inspected context, tools, permissions, client event conversion, extensions, hooks, cache discipline, and rollout recovery. This chapter shows how external clients enter that same runtime through app-server and SDK adapters. The next chapter returns to long-lived state: how stable lessons from old rollouts become memory for future threads.
| Integration need | Question to ask | Mechanism | Do not read it as |
|---|---|---|---|
| IDE or rich client. | Do you need all thread, turn, event, and control operations? | codex app-server plus JSON-RPC notifications. |
A synchronous answer API. |
| Python automation. | Do you need app-server v2 methods and typed notifications? | Python SDK over app-server stdio. | Direct chat-text read/write. |
| Node or TypeScript task runner. | Is CLI JSONL enough for the workflow? | TypeScript SDK over codex exec --experimental-json. |
A full app-server client. |
| Same-process host. | Can it preserve app-server semantics? | in_process typed requests plus the same result envelope. |
A bypass around initialization, queues, and notifications. |
A public interface for an agent runtime should not be only a convenient mouthpiece for the model. It should bring external callers into the runtime’s existing thread, turn, permission, rollout, and event mechanisms. The more the runtime can do, the more valuable that shared implementation becomes.
Sources
- Official App Server protocol, initialization, and lifecycle
core-apipublic facade re-exporting thread and runtime primitivesClientRequestmacro and serialization scopeServerNotificationmacro and JSON-RPC notification conversion- JSON-RPC and typed requests enter the same processing path
- initialize check, experimental API check, and serialization queue dispatch
ThreadStart,ThreadResume, andThreadForkdispatchTurnStart,TurnSteer, andTurnInterruptdispatch- request serialization queue keys and exclusive/shared read access
ThreadStartParamspublic shapeThreadResumeParamsresume variantsThreadForkParamsfork variantsTurnStartParamspublic shapethread_start_innerandthread_start_taskturn_start_innerbuildsTurnInputRequest- turn settings override validation
- conversation listener subscription
- listener task reads core events and projects them
TurnStartedandTurnCompletenotification handling- item delta events become
ServerNotification TurnCompletednotification emission- in-process app-server runtime host design note
- in-process start performs initialize / initialized
- in-process request still enters
process_client_request - Python SDK
CodexClientconfig, app-server startup, and initialization - Python SDK thread start/resume/fork wrappers
- Python SDK turn start, wait, and stream helpers
- Python SDK
MessageRouterroutes responses and notifications - Python SDK high-level
Codexandthread_start - Python SDK
Thread.run/Thread.turn - Python SDK collects
TurnResult - TypeScript SDK README: CLI JSONL wrapper, streaming, and resume
- TypeScript SDK
CodexExec.runbuildscodex exec --experimental-json - TypeScript SDK spawns CLI and reads stdout lines
- TypeScript SDK
Thread.runStreamed/run