Asteria Docs

Cookbooks

Integration recipes for common Asteria use cases: drop-in replacement, shared agents, working with files, Power Automate, and programmatic project setup.

Drop-in OpenAI replacement

Switch from OpenAI to Asteria by changing two things: the API key and the base URL. The model field now selects the LLM; agent selects the behavior layer.

Python

from openai import OpenAI

client = OpenAI(
    api_key="sk-ast-YOUR_KEY",
    base_url="https://api.asteria-labs.com/v1",
)

# Orchestrator mode: RAG + tools
response = client.chat.completions.create(
    model="claude-opus-4-6",
    extra_body={"agent": "asteria"},
    messages=[{"role": "user", "content": "Summarize the attached report."}],
    stream=True,
)
for chunk in response:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "sk-ast-YOUR_KEY",
  baseURL: "https://api.asteria-labs.com/v1",
});

const completion = await client.chat.completions.create({
  model: "claude-opus-4-6",
  // @ts-ignore (Asteria Cloud extension)
  agent: "asteria",
  messages: [{ role: "user", content: "What are our top 5 risks?" }],
});
console.log(completion.choices[0].message.content);

Environment variables

export OPENAI_API_KEY="sk-ast-YOUR_KEY"
export OPENAI_BASE_URL="https://api.asteria-labs.com/v1"
# OpenAI SDK picks these up automatically
# You still need to pass agent in the request body

Embed a shared agent in a web app

Create a custom agent once in the platform, share it, then call it by slug from any application.

1. Create and share the agent

In the platform, open the Agent Builder:

  1. Set name, system prompt, tools.
  2. Set visibility to Shared.
  3. Save. The slug appears below the agent name (e.g., hr-bot-h26dx).

Discover available agents via GET /v1/agents.

2. Backend proxy (Node.js + Express)

Keep the API key server-side:

import express from "express";
import OpenAI from "openai";

const app = express();
app.use(express.json());

const asteria = new OpenAI({
  apiKey: process.env.ASTERIA_API_KEY,
  baseURL: process.env.ASTERIA_BASE_URL,
});

app.post("/api/chat", async (req, res) => {
  res.setHeader("Content-Type", "text/event-stream");
  res.setHeader("Cache-Control", "no-cache");
  res.setHeader("Connection", "keep-alive");

  const stream = await asteria.chat.completions.create({
    model: "gpt-4o",
    // @ts-ignore
    agent: "hr-bot-h26dx",
    messages: [{ role: "user", content: req.body.message }],
    stream: true,
  });

  for await (const chunk of stream) {
    const content = chunk.choices[0]?.delta?.content;
    if (content) res.write(`data: ${JSON.stringify({ content })}\n\n`);
  }
  res.write("data: [DONE]\n\n");
  res.end();
});
app.listen(3000);

Work with files

Two halves of the same surface: files you upload, and files the agent produces. Both live at /v1/files, both are addressed by file_id, and both count against your organisation's storage quota.

Upload once, ask many times

Sending a 20 MB report inline on every question pays for the same bytes every time. Upload it once instead and reference the id, which also lifts the ceiling from 32 MB to 250 MB.

from openai import OpenAI

client = OpenAI(api_key="sk-ast-YOUR_KEY", base_url="https://api.asteria-labs.com/v1")

# Upload once. The id is stable and has no expiry.
with open("annual-report.pdf", "rb") as fh:
    uploaded = client.files.create(file=fh, purpose="assistants")

# Ask as many times as you like. No bytes on the wire after the first call.
for question in ["What was revenue?", "List the stated risks.", "Who signed off?"]:
    answer = client.chat.completions.create(
        model="default",
        extra_body={"agent": "asteria"},
        messages=[{
            "role": "user",
            "content": [
                {"type": "input_text", "text": question},
                {"type": "input_file", "file_id": uploaded.id},
            ],
        }],
    )
    print(question, "->", answer.choices[0].message.content)

model: "default" follows your organisation's configured chat model, so you do not have to name one.

Collect what the agent produced

In orchestrator mode the agent also writes files. crawl_site reads a whole site into one document, fetch_url saves what it downloads, and sandbox_export saves whatever a script produced. You do not ask for this and you do not have to handle it: when a tool produces a file, its result names the id, and the file is waiting at /v1/files.

Anything the agent wrote carries purpose="assistants_output", which is how you tell it apart from your own uploads.

before = {f.id for f in client.files.list(purpose="assistants_output").data}

client.chat.completions.create(
    model="default",
    extra_body={"agent": "asteria"},
    messages=[{"role": "user", "content": "Read https://docs.example.com and summarise their auth model."}],
)

# Whatever the run produced, ready to download or feed into the next call.
produced = [
    f for f in client.files.list(purpose="assistants_output").data
    if f.id not in before
]
for f in produced:
    print(f.id, f.filename, f.bytes)

Two things to know:

  • The file_id is the handle, not the workspace path. A tool result may mention a path like /data/crawl_docs.example.com.md. That workspace lives only for the request that created it; the id keeps working afterwards.
  • purpose is the filter. assistants_output is what the agent wrote, assistants is what you uploaded. Listing without a filter returns both.

TypeScript

import OpenAI from "openai";
import { createReadStream } from "fs";

const client = new OpenAI({
  apiKey: process.env.ASTERIA_API_KEY,
  baseURL: "https://api.asteria-labs.com/v1",
});

const uploaded = await client.files.create({
  file: createReadStream("annual-report.pdf"),
  purpose: "assistants",
});

const answer = await client.chat.completions.create({
  model: "default",
  // @ts-ignore (Asteria Cloud extensions: agent, and the input_* content parts)
  agent: "asteria",
  // @ts-ignore
  messages: [{
    role: "user",
    content: [
      { type: "input_text", text: "What was revenue?" },
      { type: "input_file", file_id: uploaded.id },
    ],
  }],
});
console.log(answer.choices[0].message.content);

// What the agent wrote during that run:
const produced = await client.files.list({ purpose: "assistants_output" });

Housekeeping

Files persist until you delete them, as they do on OpenAI and Anthropic, so nothing breaks under a deployed integration on a timer. That also means output accumulates if a scheduled job runs daily. Delete what you no longer need:

for f in client.files.list(purpose="assistants_output").data:
    client.files.delete(f.id)

Identical content stored twice is kept once and counted once, but each upload is a separate id you delete independently. A key from another organisation gets a 404 for your ids, the same answer an id that never existed gets.


Build a search index with embeddings

/v1/embeddings is the one endpoint here that runs no agent. It takes text and returns vectors, priced per input token, and it is OpenAI-compatible so the official client works unchanged.

Index documents, then search them

Two calls, and the difference between them is the point.

import numpy as np
from openai import OpenAI

client = OpenAI(api_key="sk-ast-YOUR_KEY", base_url="https://api.asteria-labs.com/v1")

docs = [
    "Expenses over 500 EUR need director approval.",
    "Laptops are replaced every three years.",
    "Parental leave is 16 weeks at full pay.",
]

# Indexing: these are documents.
indexed = client.embeddings.create(model="default", input=docs)
matrix = np.array([d.embedding for d in indexed.data])

# Searching: this is a query. On an asymmetric model that distinction
# is what makes the match good.
query = client.embeddings.create(
    model="default",
    input="how long is maternity leave",
    extra_body={"input_type": "query"},
)
q = np.array(query.data[0].embedding)

scores = matrix @ q / (np.linalg.norm(matrix, axis=1) * np.linalg.norm(q))
print(docs[int(scores.argmax())])

Vectors come back at the model's native width, whatever that is: a request against google/gemini-embedding-2 returns 3072 numbers, not a size we chose for you. Your index, your dimensions.

Choosing a model and a width

# `default` follows your organization's assigned embedding model.
client.embeddings.create(model="default", input="hello")

# Or name one your organization has assigned.
client.embeddings.create(model="google/gemini-embedding-2", input="hello")

# Truncate, if the model supports it, to trade a little accuracy for storage.
client.embeddings.create(model="default", input="hello", dimensions=256)

There is no fallback chain here, unlike chat. Two embedding models produce vectors in different latent spaces, so quietly answering from a second one would hand you vectors that cannot be compared against the ones you already stored. A failure is a failure, and you get to decide what to do about it.

Notes

  • encoding_format accepts float and base64. The Python SDK asks for base64 on its own when numpy is installed and decodes it for you, so you normally do not think about this at all.
  • input takes a string or an array of strings, up to 2048 of them per request. Arrays of token ids are not accepted: decoding them needs the embedding model's own tokenizer, and guessing it would silently embed different text.
  • Embeddings are priced per input token and have no completion side, so the usage block reports prompt_tokens and an equal total_tokens.

Call from Power Automate

Use Power Automate's HTTP action. Non-streaming only, so set stream: false.

HTTP action:

FieldValue
MethodPOST
URIhttps://api.asteria-labs.com/v1/chat/completions
HeadersAuthorization: Bearer sk-ast-YOUR_KEY, Content-Type: application/json

Body (custom agent):

{
  "model": "claude-opus-4-6",
  "agent": "hr-bot-h26dx",
  "messages": [{"role": "user", "content": "@{triggerBody()?['message']}"}]
}

Parse JSON the response: the reply is at choices[0].message.content.

Tips: Set timeout ≥ 120 s. Wrap in a Scope with "Configure run after" to handle 429 errors.


Programmatic project setup

Use the Admin API to provision projects, keys, and members from CI or automation scripts.

Requires an admin key (sk-ast-adm-...). Create one under Settings → Admin API Keys.

import httpx

BASE = "https://api.asteria-labs.com/v1/organization"
HEADERS = {"Authorization": "Bearer sk-ast-adm-YOUR_KEY"}

# Create a project
project = httpx.post(
    f"{BASE}/projects",
    headers=HEADERS,
    json={"name": "New Team Project"},
).raise_for_status().json()
project_id = project["id"]

# Create an API key for it
key = httpx.post(
    f"{BASE}/projects/{project_id}/api-keys",
    headers=HEADERS,
    json={"name": "ci-pipeline", "rate_limit_rpm": 30},
).raise_for_status().json()
print("API key (save this!):", key["key"])  # shown only once

# Add a team member
httpx.post(
    f"{BASE}/projects/{project_id}/users",
    headers=HEADERS,
    json={"user_id": 42, "role": "editor"},
).raise_for_status()

# Query org usage for the current month
from datetime import date
today = date.today()
usage = httpx.get(
    f"{BASE}/usage",
    headers=HEADERS,
    params={"start_date": date(today.year, today.month, 1).isoformat(), "end_date": today.isoformat()},
).raise_for_status().json()
print(f"Requests: {usage['total_requests']}, Cost: ${usage['estimated_cost_usd']:.2f}")