Docs · models
Models & seats
PAVONA checks model names in the order below. If a name is ambiguous, it lists the matches so you can choose. A seat is a role assigned to a model.
How model names are matched
Enter a model name or alias. PAVONA checks these options in order:
| Kind of name | Example | What is selected |
|---|---|---|
| Agent mode | openclaw, agent, priest, harness | PAVONA's tool loop on an available local model. These aliases select a local model. Cloud models selected by name can also use tools, subject to the same permission checks. |
| Installed local | llama-3.1-8b, or a unique prefix like llama | A model on this machine or one of your other machines, at the server where it measured best. If a prefix matches more than one model, PAVONA refuses and names every candidate. |
| Catalog id | anthropic/claude-sonnet-5 | A paid model over OpenRouter, checked against the live catalog first, so a typo cannot be billed. Upper and lower case do not matter when you type it. The id that is billed is the one the catalog spells. |
| Short name | claude, grok, astra | Matched against the catalog by id, by suffix and by whole word. For example, the word "astra" selects the one model whose name contains it. A pricing variant of the same model counts as the same model, so it does not make the name ambiguous. A word that matches nothing gives an error that names what would work, and your typo is never sent to a paid API. On the day a new model is listed, a launch that finds no match re-reads the live catalog once before refusing, in case the cached copy is out of date. |
Where the model list comes from
The Models screen probes your machines to see which runtimes answer. It tries Ollama first (its default port, or the one its environment names), then any OpenAI-compatible runtime on the network, such as LM Studio or vLLM. The list is built from the runtimes that answered. A runtime that is down is listed as down instead of being left out. /models in the console reads the same probe.
Putting a model on a machine
The panel's Discover tab and the console's /catalog read the same catalogue of installable models: the Ollama library and Hugging Face GGUF files, each shown with its actual size. /pull llama3.1 downloads a model onto this machine and shows the progress. Name a host as well to put it on any machine of yours that can download. Only Ollama can be told to fetch over HTTP. For other runtimes the console explains how to get a model onto them (for LM Studio, through its own Discover tab). /hosts lists every machine and whether it can download. /adopt adds a second computer running Ollama as a model server, and checks it at the moment you add it.
The default model
The model the panel starts with is set in a text file you can edit: upgrades\PRIEST-MODEL.txt. It is read on every turn, so after you change it the next message uses the new model without a restart. The word auto is a rule. It resolves to the best model that measured as capable across your machines at that moment.
Paid cloud models and the spending limit
A cloud model is used per call through OpenRouter with your key, and every call goes through the same spending check. The month's spending is compared, per step, against the limit you set in the Money panel. Each step's cost is read from the usage figures in the response and recorded per model. When the limit is reached, paid models stop and the figures are shown. Local inference has no provider fee and is not subject to the cloud spending limit. Cloud providers bill your account separately, and their billing may lag the figures shown in PAVONA.
The same tools for every model
Local and cloud models use the same tool handler and permission checks. Available tools depend on your configuration. Cloud requests are billed by the provider for each step.