Models
Ordering the chat chain that decides your default, and setting the three task models, with the constraints that will reject your first attempt.
Models in the console. Nothing in the product works until this page has something in it.
On credits, some of this page is read-only.
Your providers and models are provisioned for you and cannot be added, edited or deleted here: those actions are refused, because the provider is part of your plan rather than an account of your own. Your organisation also arrives with the recommended models already assigned to every job, so nothing below is a setup step you have to complete.
What remains yours is which of those models does what, and both sections apply as written: the chat chain (which is also how you pick the default) and the task models. Skip Adding a provider and Adding a model. See How billing works.
Two layers: providers (an account somewhere, with a credential) and models (what that provider offers you).
If you are setting up for the first time, the order that works is: add a provider, add one chat model and put it at the top of the chat chain, add an embedding model, then come back for the rest.
Adding a provider
Add Provider, pick one, give it a credential. Keys are encrypted at rest.
Most providers want an API key. Two are different:
Google Vertex takes either a service account JSON or the environment's own credentials, plus a GCP project id and a region.
Custom is any OpenAI-compatible endpoint: a self-hosted model, a gateway, something on your own infrastructure. You give it a URL.
Test connection before saving. It is faster than discovering the key is wrong when a user does.
A provider always runs on the credential you gave it. If one is missing, the provider stops rather than reaching for anything else.
That safeguard is deliberate: it guarantees the traffic on your bill is traffic you authorised, and that a configuration mistake shows up immediately instead of quietly running somewhere you did not intend.
Deleting a provider deletes its models too.
Adding a model
Under a provider, Add Model. The fields that matter:
Model ID is the provider's own identifier. Leave the field and pricing, context window and name autofill from a public catalogue, filling only what is empty. Review before saving: the catalogue is good, not authoritative, and a wrong price silently distorts every cost figure you will later rely on.
Context window drives automatic conversation compaction. Set it correctly or long conversations compact at the wrong point.
Thinking mode: Off (unsupported), Optional (users get the toggle), Always (the model reasons on every request). This is what puts the brain icon in the message box.
Built-in web search marks a model that can search on its own, which pairs with the "prefer the model's own search" setting in Organization.
Hidden from chat selector keeps a model out of the user-facing dropdown. Use it for anything not meant for conversation.
Pricing is required. Without it a model still runs, and every usage and cost figure it touches is wrong by omission.
Chat models and the chain
Chat models are not chosen one at a time. You order them into a chain of up to three, and that ordering decides everything:
The model at the top is your organisation's default. It is what new conversations start on, and it is what an API request gets when it omits model, so applications that name no model follow whatever sits first here. See Gateway mode. There is no separate "make this the default" switch: promoting a model to the top of the chain is that switch.
On credits the chain reads the same way as it does anywhere else. Each model carries a posture label instead of a price, Cheap and fast, Balanced or Best intelligence, so a choice never depends on reading a price card. The model your organisation was provisioned with is badged Recommended, so you can always see what you have moved away from. Your members see a shorter form of the same label in their own picker.
The models under it are the fallback. When the model above fails, the next is tried, and the user is told which model actually answered rather than being left to wonder.
The chain covers conversations in the app. API callers opt into it per request with enable_fallback, so an application that does not ask for it sees the failure instead. That is deliberate: a caller may prefer a clean error to an answer from a model it did not choose.
Each model is tested when you save, so a chain that cannot serve is refused at that point rather than at somebody's next message.
Put a flash or lite tier model at the top.
Today's flash-class models are sufficiently capable for the large majority of what this product does, and the top of this chain is used in a great many places: every conversation in the organisation, plus internal work that has no model of its own, such as naming conversations, summarising, and generating a workflow from a description.
Because it is used so widely, its cost and its latency compound across every message every member sends. A frontier model in this slot means paying frontier prices to title a conversation, for output nobody reads.
Members who want something heavier can switch model in a conversation, and can set their own personal default from the chat model picker. Neither changes what the organisation falls back to.
Worth ordering the chain before you need it: a provider outage becomes a degradation instead of an outage. Put a model from a different provider below the top one, or an outage takes the whole chain with it.
On credits it works differently, and better. Your models are reached through a gateway that already routes each one across several upstream providers of its own, so a single provider going down is absorbed before your chain is ever consulted. What your chain adds on top is cover for the model itself: put a different model family underneath, Gemini under GPT say, and a bad release or a model-wide outage still has somewhere to go.
Three more things to know:
Your users are told when it happens. A banner appears above the message box naming the model that was unavailable and the one that answered instead, which they can dismiss. So a fallback is a visible degradation rather than a silent one, and somebody may well ask you about it before you have noticed.
Watch the pattern, not the incident. One banner is the chain doing its job. A chain being used repeatedly means the model at the top is unhealthy, and that shows up in the Dashboard alert and the fallback breakdown in Usage, both worth a weekly glance.
Embedding and reranking cannot have chains, and the reason is not a missing feature. See below.
Choosing which models your team sees
Every model has a switch: Offer this model to users. Off, it stays configured and disappears from everyone's picker.
On credits this is how you shape the list. Your organisation is provisioned with the whole catalogue, and the recommended models are already on. Turning a few more on, or a few off, is the normal way to decide what your team is offered without configuring anything.
Task models
Three slots, for background jobs rather than conversation. Each is chosen separately, and each has a different reason for existing.
Embeddings
Turns your documents into vectors so a collection can be searched by meaning rather than by keyword. It runs when a document is added and again for every question asked of a collection.
Without one, collections cannot be indexed at all. Users get "Embedding model not configured", which they cannot act on. Configure this before anyone tries to make a collection.
Three constraints, and the first will reject your first attempt if you skip it:
The model must output 1024 dimensions. That is the canonical width the whole index is built on, and it is checked when you save: a model of another width is rejected there and then rather than failing later.
Most providers offer a 1024-wide option or a way to request one. If the model you want only emits 1536 or 768, it is not usable here.
Batch size is how many texts go in one API call, and every provider caps it differently (Voyage allows 1000, Cohere 96). Set it to what your provider documents.
If you do not know the cap, be conservative: 100 or below. Too high and the provider rejects whole batches, which surfaces as indexing that fails for large documents and works for small ones. Too low only costs a little throughput, and indexing is a background job.
No fallback chain, and it could not have one. A collection is searched with the model that indexed it, so falling back to another model would search one index with another model's vectors: confident nonsense rather than an error.
Reranking
Takes the candidates retrieval found and reorders them by how well they actually answer the question. Optional, and it earns its place on large collections where the first pass returns plenty of plausible-looking material.
Has its own batch size, with the same advice: match the provider's documented cap, and stay at or below 100 if you do not know it.
No fallback chain, for the same reason as embeddings.
Image generation
Used when someone asks for an image. Optional: without it, that request simply is not available.
It must be a model that generates images, not one that merely reads them. A vision-capable chat model can describe a picture and cannot produce one. What you want here is an image model: Google's Gemini image models (the ones marketed as Nano Banana) or OpenAI's image models.
Cost per image is optional and worth filling in, because image models are priced per image rather than per token, so leaving it empty means image generation contributes nothing to your cost figures.
What is not on this page
Two jobs people expect to find a slot for, and there is none to set:
Authoring Workshop documents and decks runs on the chat model of the conversation that asked for it. If a member wants a document written by a stronger model, they switch model in the conversation before asking.
Drafting a workflow from a description runs on your default chat model, and falls through its chain if that model is unavailable.
Both follow the chat model deliberately: the person asking already chose a model for the conversation, and quietly answering on a different one would make the cost and the writing style disagree with what they picked.
Changing the embedding model
The one change on this page with real consequences, and the console explains itself well when you attempt it.
Switching re-embeds everything already indexed: document chunks, knowledge graph entities, conversation summaries. You are shown the counts and an estimated duration before committing.
The design is careful, and worth trusting:
- New vectors are built alongside the old ones and swapped in only when the whole corpus is ready.
- Search stays available throughout. The current index keeps answering.
- You can cancel at any point, and nothing indexed is lost.
- If it fails, your existing vectors are untouched and you stay on the previous model.
You cannot change the embedding model while a switch is running.
Do this deliberately: on a large corpus it is hours of work, and the reason to do it is a genuinely better model, not curiosity.
A working setup, end to end
For an organisation starting from nothing:
| Slot | Pick |
|---|---|
| Top of the chat chain | A flash or lite tier model. Capable enough for the large majority of work, used almost everywhere, and your default by virtue of being first |
| Under it | The same tier from a different provider, or a different model family on credits |
| Embeddings | A 1024-dimension model, batch size at the provider's cap or 100 |
| Reranking | Optional. Add it when collections get large |
| Image generation | Optional. Only if your members need generated images |