Reference
OpenAI-compatible API
The daemon serves /v1/models and /v1/chat/completions on the
OpenAI chat-completions wire. Anything that already speaks it — Open WebUI, LobeChat, Continue,
Zed, the openai client libraries, curl — can talk to Prometheus without a
Prometheus-specific client. Beacon stays the cockpit; this is the second door.
Endpoints
| Route | What it does |
|---|---|
GET /v1/models | The model catalog as OpenAI-shaped rows. The id is the key you pass as model: local for the configured primary, or a cloud preset. Only models the daemon can currently reach are listed. |
POST /v1/chat/completions | One agent turn. Send the whole conversation; get the reply as a completion object, or as Server-Sent Events in the OpenAI chunk shape when "stream": true. |
Both sit behind the same bearer-token middleware as /api/. There is no separate key
for this surface.
Authentication
The web API token, as a bearer. oara token show prints it; it also lives in the
daemon's env file as PROMETHEUS_API_TOKEN. See
Tokens and the open web API before you expose the port beyond
localhost or your tailnet — this surface is exactly as public as the rest of the API.
curl http://localhost:8005/v1/chat/completions \
-H "Authorization: Bearer $PROMETHEUS_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'
Pointing a client at it is the same three values everywhere: base URL
http://<daemon-host>:8005/v1, API key = the token, model = a key from
/v1/models. Open WebUI calls these an “OpenAI API connection”; Continue and
Zed call it an OpenAI-compatible provider.
What is different from OpenAI, on purpose
| Where | What happens, and why |
|---|---|
| state | Stateless per request. Each call runs one agent turn in a fresh openai:<id> session that persists nothing to the daemon's conversation memory. Like OpenAI, the client owns the history and sends all of it every call. The daemon's own sessions, memory extraction and retention are untouched; it does not learn from a compat client. |
| tools | The tools are the daemon's. A request carrying tools, functions, tool_choice or function_call is refused with 400 tools_unsupported. Refusing is honest; ignoring would let a client believe its tools were in play. The model uses Prometheus's tools, gated by the security gate, and the client only ever sees text deltas. A tool that needs an approval blocks the turn exactly as it would for Beacon — the operator approves it there. |
| sampling | Generation settings are ignored. temperature, top_p, max_tokens, n, stop, logprobs. The agent loop owns them. They are accepted and dropped rather than refused, so clients that always send them keep working. |
| system | system messages are appended, not substituted. They land under an “Instructions from the connecting client” heading after the daemon's own system prompt, and never replace the identity, tool and safety text the loop is built on. |
| model | model is a catalog key, not a model name. It applies a per-request override and clears it afterwards. An unknown key is 404 model_not_found, with the fix in the message: pick one from GET /v1/models. |
| content | Text only. Message content must be a string or text parts. Image and audio parts are refused with 400 unsupported_content. The gateways handle inbound media; this surface does not. |
One extension
"mode": "chat" in the request body runs the turn with no tools at all — a plain
model reply. The default, "agent", is the full loop. Anything else is
400 invalid_mode.
Replies also carry a prometheus object beside the standard fields: the session_id the turn ran under and the mode it used. Standard clients ignore it; it is there for people reading logs.
Errors
Every refusal uses OpenAI's error envelope — {"error": {"message", "type", "code",
"param"}} — so client libraries raise it as their own error type and the
message says what to change. The codes you will meet:
| Status and code | Meaning |
|---|---|
400 tools_unsupported | Remove the client-side tool fields. |
400 messages_required | messages must be a non-empty list. |
400 last_message_not_user | The final message must be from the user. |
400 invalid_role | Roles are system, user, assistant. |
400 unsupported_content | Text only. |
404 model_not_found | Not a key from /v1/models. |
503 loop_unavailable | The web server answered but the agent loop is not wired to it yet — usually the daemon is still starting. oara doctor will say. |
502 turn_failed (or a more specific kind) | The turn started and the loop died — a provider error, most often. Non-streaming replies get the status; a stream that has already sent headers reports it as its last event instead. |
{"error": …} event and then [DONE] instead. Read the
whole stream, not just the first chunk, and treat that event as the failure it is.