Asteria Docs

Extensions

The non-standard fields Asteria Cloud adds to the OpenAI and Anthropic APIs, and how to send them through the official SDKs.

Asteria Cloud speaks three standard dialects: OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages. Point an official SDK at our base_url and everything you already know keeps working.

A few capabilities have no home in those specifications, because no other provider has them: choosing an agent, pointing a run at one of your collections, and telling an embedding model whether it is reading a document or a query. Those arrive as extra fields on the request body.

You do not need an Asteria SDK to use them. Every official SDK has a pass-through for exactly this, and the sections below show it for each one.

Request fields

FieldChat CompletionsResponsesMessagesEmbeddings
agentyesyesyesno
collection_uuidyesyesyesno
enable_fallbackfalsefalsefalseno
thread_idyesnonono
enable_thinkingfalsefalsenono
input_typenonono"document"

Where a cell shows a value, that is the default when you omit the field.

agent

Selects what answers you. Omit it for gateway mode, where the request goes straight to the model with no tools and no system prompt, exactly as it would at the model's own provider. Set it to asteria for the orchestrator, with the full tool set. Set it to one of your own agent slugs to run that agent.

See Getting started for what each mode can do.

collection_uuid

Points the run at one of your collections, so retrieval searches that collection rather than choosing among everything the organization has. Only meaningful when agent is set: gateway mode has no retrieval to steer.

enable_fallback

Whether a failing model may be retried against your organization's fallback chain.

It is off by default on all three dialects, which is the opposite of the in-app behaviour and deliberate: an API caller pinning a model usually means it, and silently answering from a different one is the kind of surprise that is expensive to debug. Set it to true to opt in.

thread_id

Chat Completions only. Resumes a conversation created earlier through this API.

Chat Completions is otherwise stateless: omit it and no conversation is created, which is the drop-in behaviour you want when porting. The cost of omitting it is that tools needing a conversation (the scratchpad, measures, background sandbox jobs) are unavailable. On Responses the equivalent is store: true plus previous_response_id; Messages is stateless with no equivalent.

enable_thinking

Chat Completions and Responses. Asks the model to reason before it answers, which trades latency and output tokens for quality on harder questions. Messages has no equivalent field.

Not every model can think, and the models your organization has enabled decide whether yours can. On a model that does not support thinking, the field is ignored and the request is answered normally, with nothing in the response to say so: neither the OpenAI nor the Anthropic response format has a field for a non-fatal notice, so inventing one would only be discarded by your SDK. Your organization's administrator can see which models support thinking, and the response's model field tells you which one answered.

Thinking overrides temperature. When thinking is enabled, providers accept only their own default temperature and reject any other value, so we drop the one you sent rather than fail the request. Send enable_thinking or a temperature, not both.

input_type

Embeddings only. Whether these texts are documents being indexed or queries being matched against an index.

Most models ignore it: OpenAI's embeddings are symmetric, so a query and a document embed identically. Voyage and Cohere are asymmetric and embed the two differently, and passing the wrong one costs retrieval quality with no error to warn you. If your organization's embedding model is one of those, set this.

Sending them

OpenAI Python SDK

extra_body merges into the request body.

from openai import OpenAI

client = OpenAI(api_key="sk-ast-...", base_url="https://api.asteria-labs.com/v1")

response = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "What changed in the Q4 deck?"}],
    extra_body={
        "agent": "asteria",
        "collection_uuid": "018f2c1a-...",
        "enable_fallback": True,
    },
)

Embeddings take theirs the same way:

vectors = client.embeddings.create(
    model="default",
    input=["how do I reset my password"],
    extra_body={"input_type": "query"},
)

OpenAI TypeScript SDK

The TypeScript client has no extra_body. Extra keys go at the top level of the request object, with a cast, because the SDK's types describe OpenAI's schema rather than ours.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ASTERIA_API_KEY,
  baseURL: "https://api.asteria-labs.com/v1",
});

const response = await client.chat.completions.create({
  model: "default",
  messages: [{ role: "user", content: "What changed in the Q4 deck?" }],
  agent: "asteria",
  collection_uuid: "018f2c1a-...",
} as OpenAI.ChatCompletionCreateParamsNonStreaming & AsteriaExtensions);

Every extension is optional, so declare them that way and the same type works for gateway mode, where you omit agent entirely:

type AsteriaExtensions = {
  agent?: string;
  collection_uuid?: string;
  enable_fallback?: boolean;
  thread_id?: string;
};

Anthropic SDK

extra_body works the same way as in the OpenAI Python client. Authenticate with your Asteria key in x-api-key, which is the header the SDK already sends.

from anthropic import Anthropic

client = Anthropic(api_key="sk-ast-...", base_url="https://api.asteria-labs.com")

message = client.messages.create(
    model="default",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Summarize the Q4 results"}],
    extra_body={"agent": "asteria"},
)

Plain HTTP

They are ordinary body fields. Nothing special is required.

curl -X POST https://api.asteria-labs.com/v1/chat/completions \
  -H "Authorization: Bearer sk-ast-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "default",
    "agent": "asteria",
    "collection_uuid": "018f2c1a-...",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Reading extra response fields

Responses can carry fields the official SDKs do not know about. They are not lost: both the Python OpenAI and Anthropic clients keep unrecognized fields on model_extra.

response = client.chat.completions.create(...)
response.model_extra          # {'thread_id': None}
response.model_extra.get("thread_id")

In TypeScript the parsed object simply carries the field, and a cast makes it visible to the type checker.

The X-Asteria-Warning header

Some requests are honoured with a caveat. Rather than fail a call we can mostly serve, we serve it and say what we changed, in an X-Asteria-Warning response header.

The case you are most likely to meet first is a system message in agent mode. Sending one is standard practice against OpenAI, so it comes across in almost every port. An Asteria agent defines its own system prompt, so yours is ignored rather than merged, and the header says so:

X-Asteria-Warning: system message ignored: agent 'asteria' defines its own system prompt

The call still succeeds and the answer is still good, which is exactly why this is worth reading: nothing else tells you that the instructions you carefully wrote had no effect. In gateway mode there is no agent prompt to collide with and your system message is honoured normally.

Another: POST /v1/responses with store: true against an organization that has response storage disabled. The response comes back with store: false echoed, and the header explains that the request was overridden rather than honoured.

The header is advisory. A request that produces one is a successful request, so nothing breaks if you ignore it, but logging it in development will tell you about a mismatch between what you asked for and what your organization allows.

raw = client.chat.completions.with_raw_response.create(
    model="default",
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={"agent": "asteria"},
)
warning = raw.headers.get("x-asteria-warning")
if warning:
    print("heads up:", warning)

completion = raw.parse()

What we deliberately do not add

There is no Asteria SDK, and that is the point rather than a gap. The value of this API is that your existing client works unchanged. Everything above is reachable through the official SDKs today, and a wrapper would mostly duplicate them while adding a version of our own for you to keep in step.

If you want the fields typed, the shapes are small enough to declare locally: the table at the top of this page is the complete list.

Cookbooks

Integration recipes for common Asteria use cases: drop-in replacement, shared agents, working with files, Power Automate, and programmatic project setup.

Chat Completions

OpenAI-compatible chat completions endpoint. Supports both streaming (SSE) and non-streaming responses. **Routing modes**, controlled by the `model` + `agent` fields: - **Gateway mode** (`agent` omitted): pure LLM proxy, no system prompt, no tools. Rate limiting and budget enforcement still apply. - **Orchestrator mode** (`agent: "asteria"`): full Asteria agent with RAG, web search, code execution, and sub-agents. - **Custom agent mode** (`agent: "<slug>"`): invoke a shared custom agent by slug. The agent's system prompt and tools are applied. **Which model**: send `model: "default"` to use the organization's default chat model and its fallback chain, which is the robust choice, because a pinned model id can be removed by an org admin later and start returning `400`. Omitting the field entirely does the same thing, but most OpenAI-compatible clients cannot omit it, so `"default"` is the reachable form (GH #1329). **This endpoint is stateless.** Nothing is retained between calls, exactly as with OpenAI Chat Completions: send the conversation back in `messages`. State lives in `POST /v1/responses`. A `store` parameter is accepted for compatibility and does nothing. It creates no thread, and it does not retain the completion either, since there is no `GET /v1/chat/completions/{id}` to read one back from. Setting it is harmless, but keep your own copy of the response. The one exception is `thread_id`, an Asteria extension: pass one returned by an earlier call and the request runs inside that conversation, which is what the conversation-scoped tools need. Without it the orchestrator simply does not offer them, rather than offering tools that would refuse: - `scratchpad_*` and `measure_*`, which are keyed on a conversation. - `sandbox_start` / `sandbox_check` / `sandbox_stop`. A background job is delivered into a conversation after the turn ends. `sandbox_run` answers inside the turn and is always available. - `add_to_user_bio` and `search_past_conversations`. An API key is an organisation, not a person, so these are never available here. Everything else works either way, including the sandbox workspace: files that `fetch_url` or `crawl_site` bring in are visible to `sandbox_run` within the same request. Without a `thread_id` that workspace is discarded when the response completes, so nothing carries to the next call. **Non-standard parameters need `extra_body`.** `agent`, `thread_id`, `collection_uuid`, `enable_thinking` and `enable_fallback` are ours, and an OpenAI SDK builds typed request params, so passing one as a keyword is a client-side `TypeError` that never reaches us. In `openai-python`: ```python client.chat.completions.create( model="default", messages=[{"role": "user", "content": "Hello"}], extra_body={"agent": "asteria", "thread_id": thread_id}, ) ```