Getting started
Authentication, routing modes, SDK setup, and error reference for the Asteria Cloud API.
Authentication
All /v1/ requests require a Bearer token in the Authorization header.
| Key prefix | Scope | Used for |
|---|---|---|
sk-ast-... | Project | Chat completions, model/agent discovery |
sk-ast-adm-... | Admin | /v1/organization/ management routes |
Project keys are created per-project in the platform under API Keys. Admin keys are created by org admins under Settings → Admin API Keys.
Quick start
curl -X POST https://api.asteria-labs.com/v1/chat/completions \
-H "Authorization: Bearer sk-ast-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"agent": "asteria",
"messages": [{"role": "user", "content": "Hello, what can you do?"}]
}'No model is needed to get started. Omit it and your organization's default chat
model is used. Pin one when you want a specific model.
SDK setup
The Asteria Cloud API is OpenAI-compatible. Use the official OpenAI SDKs with a custom base_url.
Python
from openai import OpenAI
client = OpenAI(
api_key="sk-ast-...",
base_url="https://api.asteria-labs.com/v1",
)
response = client.chat.completions.create(
model="claude-opus-4-6",
extra_body={"agent": "asteria"},
messages=[{"role": "user", "content": "Summarize our Q4 results"}],
stream=True,
)
for chunk in response:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")TypeScript
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "sk-ast-...",
baseURL: "https://api.asteria-labs.com/v1",
});
const stream = await client.chat.completions.create({
model: "claude-opus-4-6",
// @ts-ignore (Asteria Cloud extension field)
agent: "asteria",
messages: [{ role: "user", content: "Summarize our Q4 results" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}Routing modes
Every request to POST /v1/chat/completions runs in one of three modes, controlled by model and agent.
model | agent | Mode | Behavior |
|---|---|---|---|
"default" | (omitted) | Gateway | Pure LLM proxy on the organization's default chat model |
"claude-opus-4-6" | (omitted) | Gateway | Pure LLM proxy: your system prompt, no Asteria tools |
"default" | "asteria" | Orchestrator | Full Asteria agent on the organization's default chat model |
"claude-opus-4-6" | "asteria" | Orchestrator | Full Asteria agent: RAG, web search, code execution |
"gpt-4o" | "hr-bot-h26dx" | Custom agent | Shared custom agent by slug |
agent selects the agent layer and is optional.
Choosing the model
Send model: "default" and the request runs on your organization's default chat
model. The same default the Asteria app uses. Set it to a real id to pin a
specific model; GET /v1/models lists what your project may use, and lists real
models only.
Leaving model out does the same as "default". Reach for the sentinel anyway:
the OpenAI SDK and everything built on it always send the field, so an omitted
model is only available if you assemble the request yourself. The name is
reserved and can never collide with a real model.
Two failure cases both return 400 invalid_request_error:
- the organization has no default chat model configured, and
modelwas"default"or omitted ("No default chat model is configured for this organization"); - the default model resolves to one your project's allow-list excludes ("The organization default model is not allowed for this project").
"default" does not opt you into the fallback chain: that still requires
"enable_fallback": true, whether the model is pinned or defaulted.
Gateway mode
No system prompt is injected, no tools are loaded. You control the full conversation.
{
"messages": [
{"role": "system", "content": "You are a concise financial analyst."},
{"role": "user", "content": "What is our cash position?"}
]
}Orchestrator mode
{
"agent": "asteria",
"messages": [{"role": "user", "content": "Summarize our Q4 board deck"}]
}Any role: "system" message is ignored with a non-fatal warning field in the response.
Custom agent mode
Call a shared custom agent by slug. Discover slugs with GET /v1/agents.
{
"model": "gpt-4o",
"agent": "hr-bot-h26dx",
"messages": [{"role": "user", "content": "How many vacation days do I have left?"}]
}Opt in to the fallback chain with "enable_fallback": true.
agent, collection_uuid and enable_fallback are not part of the OpenAI
specification. They are ordinary body fields, and the official SDKs pass them
through: see Extensions for a snippet per SDK and the full
list of what we add.
Multimodal content
{
"role": "user",
"content": [
{"type": "input_text", "text": "What's in this document?"},
{"type": "input_image", "url": "https://example.com/chart.png"},
{"type": "input_file", "file_url": "https://example.com/report.pdf"}
]
}Attachments can be sent three ways, and the transport decides which limit applies:
| Transport | How | Counts against |
|---|---|---|
| By id | {"type": "input_file", "file_id": "file-abc123"} | nothing per request. The bytes are already stored |
| By URL | "file_url", or "url" on an image | the fetch budget: 32 MB per attachment, 32 MB and 10 fetches per request |
| Inline base64 | "file_data", or "image_url"/"url" as a data: URI | the 32 MB request body limit |
Base64 inflates by about a third, so a 24 MB PDF sent inline is already near the body limit.
Above that, or for anything you will send more than once, upload it to POST /v1/files and reference the id. That endpoint takes multipart/form-data and accepts up to 250 MB, because the bytes stream to storage instead of riding inside a JSON body. file_id works on input_image too.
curl https://api.asteria-labs.com/v1/files \
-H "Authorization: Bearer sk-ast-YOUR_KEY" \
-F purpose=assistants \
-F file=@report.pdfFiles belong to the organisation rather than the key that uploaded them, and persist until you delete them, so a file_id in a deployed integration keeps working. See the cookbook recipe for the full loop.
Streaming
Set "stream": true for Server-Sent Events:
data: {"id":"chatcmpl-...","choices":[{"delta":{"content":"Hello"},"index":0}]}
data: [DONE]Limits
The maximum request body is 32 MB; over that returns 413. Attachments fetched by URL are capped at 32 MB each, 32 MB total, and 10 fetches per request. Also: 200 messages per request, 100 content parts per message, 1 MB per text part.
POST /v1/files has its own, larger ceiling of 250 MB per file. It is a separate limit rather than an exception to the one above: those bytes arrive as multipart and stream to storage, so they never have to be held in memory the way a base64 body does. Stored files count against your organisation's storage quota instead.
Rate limits and budgets
When rate-limited: HTTP 429 with Retry-After: <seconds>.
When the project budget is exhausted:
{"error": {"type": "budget_exceeded", "message": "Monthly budget exceeded..."}}A single run is also capped at a fixed amount of provider spend, so an agent looping on an expensive tool is stopped mid-run:
{"error": {"type": "run_cost_limit_exceeded", "message": "This response was stopped because it reached the $5.00 spend limit for a single request..."}}Retrying the same request will be stopped the same way. Ask something narrower, or ask your administrator to raise the ceiling. On a streaming call the sentence arrives as a final content chunk instead, since the response has already started.
Error reference
All errors use the OpenAI shape: {"error": {"message": "...", "type": "..."}}.
| Status | Type | Cause |
|---|---|---|
| 400 | invalid_request_error | Blocked model, no organization default when model is omitted, malformed body |
| 400 | run_cost_limit_exceeded | One run reached the per-run spend ceiling and was stopped |
| 401 | authentication_error | Invalid or missing API key |
| 402 | credits_exhausted | The organisation is out of Asteria Cloud credits |
| 402 | project_credit_limit_reached | This project reached its monthly credit limit; the organisation still has credit |
| 404 | invalid_request_error | Unknown agent slug |
| 413 | invalid_request_error | Request body over 32 MB |
| 429 | rate_limit_exceeded | RPM limit hit |
| 429 | budget_exceeded | Monthly budget exhausted |
| 500 | internal_error | Server-side failure |
| 503 | server_error | The API is temporarily unreachable, briefly the case during a deploy. Safe to retry |
503 is the one status the API can return without your request having reached the application, so it carries no request-specific detail. Every OpenAI-compatible SDK retries it by default; if you handle retries yourself, back off and retry rather than failing the call.