Create Message
Create a message, in Anthropic's Messages format. **Point the Anthropic SDK at us and change nothing else.** Authenticate with `x-api-key` carrying your Asteria key, exactly as the SDK sends it, or with `Authorization: Bearer` as on the rest of `/v1`. `anthropic-version` is accepted and echoed. **Routing modes**, controlled by the `agent` field, exactly as on the other two dialects: - **Gateway mode** (`agent` omitted): a pure LLM proxy, no system prompt and no tools. Rate limiting and budget enforcement still apply. - **Orchestrator mode** (`agent: "asteria"`): the full Asteria agent, with RAG, web search, code execution and sub-agents. - **Custom agent mode** (`agent: "<slug>"`): a shared custom agent by slug. **Which model**: send `model: "default"` to use the organization's default chat model and its fallback chain. Pinning a real model id is brittle, since an org admin can remove it later. The model that answers need not be an Anthropic one: the dialect you speak and the provider that serves you are independent here, which is the point of this endpoint. **`max_tokens` is required**, as it is on Anthropic's own API. A request without one is refused rather than given a ceiling it did not choose. **`stream: true`** speaks the Anthropic streaming protocol: `message_start`, then `content_block_start` / `content_block_delta` / `content_block_stop` around the text, then `message_delta` carrying `stop_reason` and the final usage, then `message_stop`. `ping` arrives on a quiet run. There is **no `[DONE]` sentinel**, which is Chat Completions' shape and not this one. A run that never opened its stream ends in a bare `error` event; one that opened and then failed closes its block and ends in `error` too. Hanging up does not cancel the run: it finishes server side, is billed and is logged, exactly as if you had stayed. **This endpoint is stateless.** Nothing is stored and there is no id to fetch back; send the prior turns in `messages`. `/v1/responses` is where state lives. **`system` is a top-level field**, not a message in the array. A `system` role inside `messages` is refused, as it is upstream. **Content blocks**: `text`, `image` and `document` are carried, and `cache_control` is honoured as a caching boundary at the end of the message that marks it (at most four per request, as upstream). `tool_use`, `tool_result`, `thinking` and `redacted_thinking` are refused by name rather than accepted and dropped, as are the `tools`, `tool_choice` and `thinking` request fields and `citations` on a document: there is no client-side tool loop on any `/v1` dialect, and a `200` that silently discarded a tool result is worse than a clear refusal. Anything carried with a documented degradation says so in an `X-Asteria-Warning` header.
Authorization
ApiKey Your API key (sk-ast-...)
In: header
Header Parameters
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
POST /v1/messages, plus this platform's three extensions.
agent, collection_uuid and enable_fallback are ours and are spelled the
same way on all three dialects, so a developer who learned them on one does
not relearn them here. model accepts "default" for the org's chain, as
everywhere else on /v1.
Response Body
application/json
application/json
curl -X POST "https://example.com/v1/messages" \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "user", "content": "string" } ], "max_tokens": 1 }'nullRemove Response
Delete a stored response and the conversation holding its history. Erases the turn's messages too, which is the point: the response object is the only handle this API gives out, so deleting it and leaving the messages would keep data the caller has no way to reach or remove. A second delete of the same id is a 404, exactly like an unknown one.
Create Embeddings
Create embeddings for one or more texts. **Model.** `model` names an embedding model enabled for your organization, or `default` for the organization's primary one. Unlike chat, there is no fallback: two embedding models produce vectors in different latent spaces, so quietly answering from a second model would return vectors that cannot be compared with the ones you already have. **Dimensions.** Vectors come back at the model's native width. Pass `dimensions` to truncate, which only models supporting Matryoshka representation honour; one that does not will report so. **Encoding.** `float` returns JSON numbers, `base64` returns each vector as base64-encoded little-endian float32. Both are supported, and the OpenAI SDKs request `base64` by default. **Not supported.** `input` must be text: arrays of token ids are refused, because decoding them needs the embedding model's own tokenizer. `user` is accepted and unused; API traffic is attributed to the key and its project.