Asteria Docs

Getting started

Authentication, routing modes, SDK setup, and error reference for the Asteria Cloud API.

Authentication

All /v1/ requests require a Bearer token in the Authorization header.

Key prefixScopeUsed for
sk-ast-...ProjectChat completions, model/agent discovery
sk-ast-adm-...Admin/v1/organization/ management routes

Project keys are created per-project in the platform under API Keys. Admin keys are created by org admins under Settings → Admin API Keys.

Quick start

curl -X POST https://api.asteria-labs.com/v1/chat/completions \
  -H "Authorization: Bearer sk-ast-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "agent": "asteria",
    "messages": [{"role": "user", "content": "Hello, what can you do?"}]
  }'

No model is needed to get started. Omit it and your organization's default chat model is used. Pin one when you want a specific model.

SDK setup

The Asteria Cloud API is OpenAI-compatible. Use the official OpenAI SDKs with a custom base_url.

Python

from openai import OpenAI

client = OpenAI(
    api_key="sk-ast-...",
    base_url="https://api.asteria-labs.com/v1",
)

response = client.chat.completions.create(
    model="claude-opus-4-6",
    extra_body={"agent": "asteria"},
    messages=[{"role": "user", "content": "Summarize our Q4 results"}],
    stream=True,
)

for chunk in response:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "sk-ast-...",
  baseURL: "https://api.asteria-labs.com/v1",
});

const stream = await client.chat.completions.create({
  model: "claude-opus-4-6",
  // @ts-ignore (Asteria Cloud extension field)
  agent: "asteria",
  messages: [{ role: "user", content: "Summarize our Q4 results" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || "");
}

Routing modes

Every request to POST /v1/chat/completions runs in one of three modes, controlled by model and agent.

modelagentModeBehavior
"default"(omitted)GatewayPure LLM proxy on the organization's default chat model
"claude-opus-4-6"(omitted)GatewayPure LLM proxy: your system prompt, no Asteria tools
"default""asteria"OrchestratorFull Asteria agent on the organization's default chat model
"claude-opus-4-6""asteria"OrchestratorFull Asteria agent: RAG, web search, code execution
"gpt-4o""hr-bot-h26dx"Custom agentShared custom agent by slug

agent selects the agent layer and is optional.

Choosing the model

Send model: "default" and the request runs on your organization's default chat model. The same default the Asteria app uses. Set it to a real id to pin a specific model; GET /v1/models lists what your project may use, and lists real models only.

Leaving model out does the same as "default". Reach for the sentinel anyway: the OpenAI SDK and everything built on it always send the field, so an omitted model is only available if you assemble the request yourself. The name is reserved and can never collide with a real model.

Two failure cases both return 400 invalid_request_error:

  • the organization has no default chat model configured, and model was "default" or omitted ("No default chat model is configured for this organization");
  • the default model resolves to one your project's allow-list excludes ("The organization default model is not allowed for this project").

"default" does not opt you into the fallback chain: that still requires "enable_fallback": true, whether the model is pinned or defaulted.

Gateway mode

No system prompt is injected, no tools are loaded. You control the full conversation.

{
  "messages": [
    {"role": "system", "content": "You are a concise financial analyst."},
    {"role": "user", "content": "What is our cash position?"}
  ]
}

Orchestrator mode

{
  "agent": "asteria",
  "messages": [{"role": "user", "content": "Summarize our Q4 board deck"}]
}

Any role: "system" message is ignored with a non-fatal warning field in the response.

Custom agent mode

Call a shared custom agent by slug. Discover slugs with GET /v1/agents.

{
  "model": "gpt-4o",
  "agent": "hr-bot-h26dx",
  "messages": [{"role": "user", "content": "How many vacation days do I have left?"}]
}

Opt in to the fallback chain with "enable_fallback": true.

agent, collection_uuid and enable_fallback are not part of the OpenAI specification. They are ordinary body fields, and the official SDKs pass them through: see Extensions for a snippet per SDK and the full list of what we add.

Multimodal content

{
  "role": "user",
  "content": [
    {"type": "input_text", "text": "What's in this document?"},
    {"type": "input_image", "url": "https://example.com/chart.png"},
    {"type": "input_file", "file_url": "https://example.com/report.pdf"}
  ]
}

Attachments can be sent three ways, and the transport decides which limit applies:

TransportHowCounts against
By id{"type": "input_file", "file_id": "file-abc123"}nothing per request. The bytes are already stored
By URL"file_url", or "url" on an imagethe fetch budget: 32 MB per attachment, 32 MB and 10 fetches per request
Inline base64"file_data", or "image_url"/"url" as a data: URIthe 32 MB request body limit

Base64 inflates by about a third, so a 24 MB PDF sent inline is already near the body limit.

Above that, or for anything you will send more than once, upload it to POST /v1/files and reference the id. That endpoint takes multipart/form-data and accepts up to 250 MB, because the bytes stream to storage instead of riding inside a JSON body. file_id works on input_image too.

curl https://api.asteria-labs.com/v1/files \
  -H "Authorization: Bearer sk-ast-YOUR_KEY" \
  -F purpose=assistants \
  -F file=@report.pdf

Files belong to the organisation rather than the key that uploaded them, and persist until you delete them, so a file_id in a deployed integration keeps working. See the cookbook recipe for the full loop.

Streaming

Set "stream": true for Server-Sent Events:

data: {"id":"chatcmpl-...","choices":[{"delta":{"content":"Hello"},"index":0}]}
data: [DONE]

Limits

The maximum request body is 32 MB; over that returns 413. Attachments fetched by URL are capped at 32 MB each, 32 MB total, and 10 fetches per request. Also: 200 messages per request, 100 content parts per message, 1 MB per text part.

POST /v1/files has its own, larger ceiling of 250 MB per file. It is a separate limit rather than an exception to the one above: those bytes arrive as multipart and stream to storage, so they never have to be held in memory the way a base64 body does. Stored files count against your organisation's storage quota instead.

Rate limits and budgets

When rate-limited: HTTP 429 with Retry-After: <seconds>.

When the project budget is exhausted:

{"error": {"type": "budget_exceeded", "message": "Monthly budget exceeded..."}}

A single run is also capped at a fixed amount of provider spend, so an agent looping on an expensive tool is stopped mid-run:

{"error": {"type": "run_cost_limit_exceeded", "message": "This response was stopped because it reached the $5.00 spend limit for a single request..."}}

Retrying the same request will be stopped the same way. Ask something narrower, or ask your administrator to raise the ceiling. On a streaming call the sentence arrives as a final content chunk instead, since the response has already started.

Error reference

All errors use the OpenAI shape: {"error": {"message": "...", "type": "..."}}.

StatusTypeCause
400invalid_request_errorBlocked model, no organization default when model is omitted, malformed body
400run_cost_limit_exceededOne run reached the per-run spend ceiling and was stopped
401authentication_errorInvalid or missing API key
402credits_exhaustedThe organisation is out of Asteria Cloud credits
402project_credit_limit_reachedThis project reached its monthly credit limit; the organisation still has credit
404invalid_request_errorUnknown agent slug
413invalid_request_errorRequest body over 32 MB
429rate_limit_exceededRPM limit hit
429budget_exceededMonthly budget exhausted
500internal_errorServer-side failure
503server_errorThe API is temporarily unreachable, briefly the case during a deploy. Safe to retry

503 is the one status the API can return without your request having reached the application, so it carries no request-specific detail. Every OpenAI-compatible SDK retries it by default; if you handle retries yourself, back off and retry rather than failing the call.