Create Response
Create a model response, in OpenAI's Responses format. **Routing modes**, controlled by the `agent` field, exactly as on `/v1/chat/completions`: - **Gateway mode** (`agent` omitted): a pure LLM proxy, no system prompt and no tools. Rate limiting and budget enforcement still apply. - **Orchestrator mode** (`agent: "asteria"`): the full Asteria agent, with RAG, web search, code execution and sub-agents. - **Custom agent mode** (`agent: "<slug>"`): a shared custom agent by slug. **Which model**: send `model: "default"` to use the organization's default chat model and its fallback chain. Pinning a real model id is brittle, since an org admin can remove it later. **`store` defaults to off, and OpenAI's defaults it on.** That inversion is the one thing to know before chaining. An unstored call retains nothing, so a chain built on it has nothing to continue from. `store: true` keeps the response and its history: fetch it back with `GET /v1/responses/{id}`, erase it with `DELETE /v1/responses/{id}`, and it expires on its own after the retention window (30 days by default, configurable per project and per organization). An organization can also disable storage outright, in which case `store: true` is overridden: the response comes back with `store: false` and an `X-Asteria-Warning` header saying so. **`previous_response_id` continues a stored response.** Its history is replayed for you, so send only what is new in `input`. The new turn joins the same conversation, which is what lets a chain keep one scratchpad and one sandbox workspace across turns, and the whole chain's retention runs from its most recent turn rather than its first. An id that names no stored response is a `404` that says so and names `store: true`, never an empty chain answered with a `200`. Streaming is not available here yet; `/v1/chat/completions` streams today. **Storing changes which tools the run has.** A stored call has a conversation, so the tools that need one are offered: `scratchpad_*`, `measure_*` and the background sandbox verbs. An unstored call has none of them, though `sandbox_run` still answers inside the turn either way. `add_to_user_bio` and `search_past_conversations` are never available on this surface, which has no signed-in user. **Non-standard parameters need `extra_body`.** `agent`, `collection_uuid`, `enable_thinking` and `enable_fallback` are ours, and an OpenAI SDK builds typed request params, so passing one as a keyword is a client-side `TypeError` that never reaches us. In `openai-python`: ```python client.responses.create( model="default", input="Hello", extra_body={"agent": "asteria"}, ) ```
Authorization
ApiKey Your API key (sk-ast-...)
In: header
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Request body for POST /v1/responses.
Every class docstring and description= in this request's schema is
published: it reaches the OpenAPI spec and docs-site renders it into a
public page, where check-content.mjs bans the em dash. Comments beside
them are ours and are not published; the two are a line apart and only one
is a CI gate.
Response Body
application/json
application/json
curl -X POST "https://example.com/v1/responses" \ -H "Content-Type: application/json" \ -d '{ "input": "string" }'nullList Agents
List available agents for the caller's organization. Returns the built-in Asteria orchestrator plus any shared custom agents that have a slug. The `id` field in each entry is the slug to pass as `agent` in POST /v1/chat/completions.
Retrieve Response
Retrieve a stored response by id. Only responses created with `store: true` exist to be fetched, and only until they expire (30 days by default, configurable per project). The read is scoped to the organisation that owns the response: there is no signed-in user on `/v1`, so the organisation is the only principal a scope can be built from. **What comes back is the response, not the request.** `output`, `status`, `model` and `usage` are stored facts about the turn and are returned verbatim. The parameters the create call echoed back (`instructions`, `temperature`, `top_p`, `max_output_tokens`, `text`) are the caller's own request and are not stored, so they come back null here. If you need them alongside the output, keep them where you keep the id.