Cookbooks
Integration recipes for common Asteria use cases: drop-in replacement, shared agents, working with files, Power Automate, and programmatic project setup.
Drop-in OpenAI replacement
Switch from OpenAI to Asteria by changing two things: the API key and the base URL. The model field now selects the LLM; agent selects the behavior layer.
Python
from openai import OpenAI
client = OpenAI(
api_key="sk-ast-YOUR_KEY",
base_url="https://api.asteria-labs.com/v1",
)
# Orchestrator mode: RAG + tools
response = client.chat.completions.create(
model="claude-opus-4-6",
extra_body={"agent": "asteria"},
messages=[{"role": "user", "content": "Summarize the attached report."}],
stream=True,
)
for chunk in response:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")TypeScript
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "sk-ast-YOUR_KEY",
baseURL: "https://api.asteria-labs.com/v1",
});
const completion = await client.chat.completions.create({
model: "claude-opus-4-6",
// @ts-ignore (Asteria Cloud extension)
agent: "asteria",
messages: [{ role: "user", content: "What are our top 5 risks?" }],
});
console.log(completion.choices[0].message.content);Environment variables
export OPENAI_API_KEY="sk-ast-YOUR_KEY"
export OPENAI_BASE_URL="https://api.asteria-labs.com/v1"
# OpenAI SDK picks these up automatically
# You still need to pass agent in the request bodyEmbed a shared agent in a web app
Create a custom agent once in the platform, share it, then call it by slug from any application.
1. Create and share the agent
In the platform, open the Agent Builder:
- Set name, system prompt, tools.
- Set visibility to Shared.
- Save. The slug appears below the agent name (e.g.,
hr-bot-h26dx).
Discover available agents via GET /v1/agents.
2. Backend proxy (Node.js + Express)
Keep the API key server-side:
import express from "express";
import OpenAI from "openai";
const app = express();
app.use(express.json());
const asteria = new OpenAI({
apiKey: process.env.ASTERIA_API_KEY,
baseURL: process.env.ASTERIA_BASE_URL,
});
app.post("/api/chat", async (req, res) => {
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
const stream = await asteria.chat.completions.create({
model: "gpt-4o",
// @ts-ignore
agent: "hr-bot-h26dx",
messages: [{ role: "user", content: req.body.message }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) res.write(`data: ${JSON.stringify({ content })}\n\n`);
}
res.write("data: [DONE]\n\n");
res.end();
});
app.listen(3000);Work with files
Two halves of the same surface: files you upload, and files the agent produces. Both live at /v1/files, both are addressed by file_id, and both count against your organisation's storage quota.
Upload once, ask many times
Sending a 20 MB report inline on every question pays for the same bytes every time. Upload it once instead and reference the id, which also lifts the ceiling from 32 MB to 250 MB.
from openai import OpenAI
client = OpenAI(api_key="sk-ast-YOUR_KEY", base_url="https://api.asteria-labs.com/v1")
# Upload once. The id is stable and has no expiry.
with open("annual-report.pdf", "rb") as fh:
uploaded = client.files.create(file=fh, purpose="assistants")
# Ask as many times as you like. No bytes on the wire after the first call.
for question in ["What was revenue?", "List the stated risks.", "Who signed off?"]:
answer = client.chat.completions.create(
model="default",
extra_body={"agent": "asteria"},
messages=[{
"role": "user",
"content": [
{"type": "input_text", "text": question},
{"type": "input_file", "file_id": uploaded.id},
],
}],
)
print(question, "->", answer.choices[0].message.content)model: "default" follows your organisation's configured chat model, so you do not have to name one.
Collect what the agent produced
In orchestrator mode the agent also writes files. crawl_site reads a whole site into one document, fetch_url saves what it downloads, and sandbox_export saves whatever a script produced. You do not ask for this and you do not have to handle it: when a tool produces a file, its result names the id, and the file is waiting at /v1/files.
Anything the agent wrote carries purpose="assistants_output", which is how you tell it apart from your own uploads.
before = {f.id for f in client.files.list(purpose="assistants_output").data}
client.chat.completions.create(
model="default",
extra_body={"agent": "asteria"},
messages=[{"role": "user", "content": "Read https://docs.example.com and summarise their auth model."}],
)
# Whatever the run produced, ready to download or feed into the next call.
produced = [
f for f in client.files.list(purpose="assistants_output").data
if f.id not in before
]
for f in produced:
print(f.id, f.filename, f.bytes)Two things to know:
- The
file_idis the handle, not the workspace path. A tool result may mention a path like/data/crawl_docs.example.com.md. That workspace lives only for the request that created it; the id keeps working afterwards. purposeis the filter.assistants_outputis what the agent wrote,assistantsis what you uploaded. Listing without a filter returns both.
TypeScript
import OpenAI from "openai";
import { createReadStream } from "fs";
const client = new OpenAI({
apiKey: process.env.ASTERIA_API_KEY,
baseURL: "https://api.asteria-labs.com/v1",
});
const uploaded = await client.files.create({
file: createReadStream("annual-report.pdf"),
purpose: "assistants",
});
const answer = await client.chat.completions.create({
model: "default",
// @ts-ignore (Asteria Cloud extensions: agent, and the input_* content parts)
agent: "asteria",
// @ts-ignore
messages: [{
role: "user",
content: [
{ type: "input_text", text: "What was revenue?" },
{ type: "input_file", file_id: uploaded.id },
],
}],
});
console.log(answer.choices[0].message.content);
// What the agent wrote during that run:
const produced = await client.files.list({ purpose: "assistants_output" });Housekeeping
Files persist until you delete them, as they do on OpenAI and Anthropic, so nothing breaks under a deployed integration on a timer. That also means output accumulates if a scheduled job runs daily. Delete what you no longer need:
for f in client.files.list(purpose="assistants_output").data:
client.files.delete(f.id)Identical content stored twice is kept once and counted once, but each upload is a separate id you delete independently. A key from another organisation gets a 404 for your ids, the same answer an id that never existed gets.
Build a search index with embeddings
/v1/embeddings is the one endpoint here that runs no agent. It takes text and
returns vectors, priced per input token, and it is OpenAI-compatible so the
official client works unchanged.
Index documents, then search them
Two calls, and the difference between them is the point.
import numpy as np
from openai import OpenAI
client = OpenAI(api_key="sk-ast-YOUR_KEY", base_url="https://api.asteria-labs.com/v1")
docs = [
"Expenses over 500 EUR need director approval.",
"Laptops are replaced every three years.",
"Parental leave is 16 weeks at full pay.",
]
# Indexing: these are documents.
indexed = client.embeddings.create(model="default", input=docs)
matrix = np.array([d.embedding for d in indexed.data])
# Searching: this is a query. On an asymmetric model that distinction
# is what makes the match good.
query = client.embeddings.create(
model="default",
input="how long is maternity leave",
extra_body={"input_type": "query"},
)
q = np.array(query.data[0].embedding)
scores = matrix @ q / (np.linalg.norm(matrix, axis=1) * np.linalg.norm(q))
print(docs[int(scores.argmax())])Vectors come back at the model's native width, whatever that is: a request
against google/gemini-embedding-2 returns 3072 numbers, not a size we chose
for you. Your index, your dimensions.
Choosing a model and a width
# `default` follows your organization's assigned embedding model.
client.embeddings.create(model="default", input="hello")
# Or name one your organization has assigned.
client.embeddings.create(model="google/gemini-embedding-2", input="hello")
# Truncate, if the model supports it, to trade a little accuracy for storage.
client.embeddings.create(model="default", input="hello", dimensions=256)There is no fallback chain here, unlike chat. Two embedding models produce vectors in different latent spaces, so quietly answering from a second one would hand you vectors that cannot be compared against the ones you already stored. A failure is a failure, and you get to decide what to do about it.
Notes
encoding_formatacceptsfloatandbase64. The Python SDK asks forbase64on its own when numpy is installed and decodes it for you, so you normally do not think about this at all.inputtakes a string or an array of strings, up to 2048 of them per request. Arrays of token ids are not accepted: decoding them needs the embedding model's own tokenizer, and guessing it would silently embed different text.- Embeddings are priced per input token and have no completion side, so the
usageblock reportsprompt_tokensand an equaltotal_tokens.
Call from Power Automate
Use Power Automate's HTTP action. Non-streaming only, so set stream: false.
HTTP action:
| Field | Value |
|---|---|
| Method | POST |
| URI | https://api.asteria-labs.com/v1/chat/completions |
| Headers | Authorization: Bearer sk-ast-YOUR_KEY, Content-Type: application/json |
Body (custom agent):
{
"model": "claude-opus-4-6",
"agent": "hr-bot-h26dx",
"messages": [{"role": "user", "content": "@{triggerBody()?['message']}"}]
}Parse JSON the response: the reply is at choices[0].message.content.
Tips: Set timeout ≥ 120 s. Wrap in a Scope with "Configure run after" to handle 429 errors.
Programmatic project setup
Use the Admin API to provision projects, keys, and members from CI or automation scripts.
Requires an admin key (sk-ast-adm-...). Create one under Settings → Admin API Keys.
import httpx
BASE = "https://api.asteria-labs.com/v1/organization"
HEADERS = {"Authorization": "Bearer sk-ast-adm-YOUR_KEY"}
# Create a project
project = httpx.post(
f"{BASE}/projects",
headers=HEADERS,
json={"name": "New Team Project"},
).raise_for_status().json()
project_id = project["id"]
# Create an API key for it
key = httpx.post(
f"{BASE}/projects/{project_id}/api-keys",
headers=HEADERS,
json={"name": "ci-pipeline", "rate_limit_rpm": 30},
).raise_for_status().json()
print("API key (save this!):", key["key"]) # shown only once
# Add a team member
httpx.post(
f"{BASE}/projects/{project_id}/users",
headers=HEADERS,
json={"user_id": 42, "role": "editor"},
).raise_for_status()
# Query org usage for the current month
from datetime import date
today = date.today()
usage = httpx.get(
f"{BASE}/usage",
headers=HEADERS,
params={"start_date": date(today.year, today.month, 1).isoformat(), "end_date": today.isoformat()},
).raise_for_status().json()
print(f"Requests: {usage['total_requests']}, Cost: ${usage['estimated_cost_usd']:.2f}")