Asteria Docs
Chat & Discovery

Chat Completions

OpenAI-compatible chat completions endpoint. Supports both streaming (SSE) and non-streaming responses. **Routing modes**, controlled by the `model` + `agent` fields: - **Gateway mode** (`agent` omitted): pure LLM proxy, no system prompt, no tools. Rate limiting and budget enforcement still apply. - **Orchestrator mode** (`agent: "asteria"`): full Asteria agent with RAG, web search, code execution, and sub-agents. - **Custom agent mode** (`agent: "<slug>"`): invoke a shared custom agent by slug. The agent's system prompt and tools are applied. **Which model**: send `model: "default"` to use the organization's default chat model and its fallback chain, which is the robust choice, because a pinned model id can be removed by an org admin later and start returning `400`. Omitting the field entirely does the same thing, but most OpenAI-compatible clients cannot omit it, so `"default"` is the reachable form (GH #1329). **This endpoint is stateless.** Nothing is retained between calls, exactly as with OpenAI Chat Completions: send the conversation back in `messages`. State lives in `POST /v1/responses`. A `store` parameter is accepted for compatibility and does nothing. It creates no thread, and it does not retain the completion either, since there is no `GET /v1/chat/completions/{id}` to read one back from. Setting it is harmless, but keep your own copy of the response. The one exception is `thread_id`, an Asteria extension: pass one returned by an earlier call and the request runs inside that conversation, which is what the conversation-scoped tools need. Without it the orchestrator simply does not offer them, rather than offering tools that would refuse: - `scratchpad_*` and `measure_*`, which are keyed on a conversation. - `sandbox_start` / `sandbox_check` / `sandbox_stop`. A background job is delivered into a conversation after the turn ends. `sandbox_run` answers inside the turn and is always available. - `add_to_user_bio` and `search_past_conversations`. An API key is an organisation, not a person, so these are never available here. Everything else works either way, including the sandbox workspace: files that `fetch_url` or `crawl_site` bring in are visible to `sandbox_run` within the same request. Without a `thread_id` that workspace is discarded when the response completes, so nothing carries to the next call. **Non-standard parameters need `extra_body`.** `agent`, `thread_id`, `collection_uuid`, `enable_thinking` and `enable_fallback` are ours, and an OpenAI SDK builds typed request params, so passing one as a keyword is a client-side `TypeError` that never reaches us. In `openai-python`: ```python client.chat.completions.create( model="default", messages=[{"role": "user", "content": "Hello"}], extra_body={"agent": "asteria", "thread_id": thread_id}, ) ```

POST
/v1/chat/completions
Authorization<token>

Your API key (sk-ast-...)

In: header

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Request body for /v1/chat/completions.

Response Body

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/chat/completions" \  -H "Content-Type: application/json" \  -d '{    "messages": [      {        "role": "system",        "content": "string"      }    ]  }'
{  "id": "string",  "object": "chat.completion",  "created": 0,  "model": "asteria",  "choices": [    {      "index": 0,      "message": {        "role": "system",        "content": "string",        "name": "string"      },      "finish_reason": "stop"    }  ],  "usage": {    "prompt_tokens": 0,    "completion_tokens": 0,    "total_tokens": 0  },  "thread_id": "string"}