Get started
First run
oara setup is the one canonical path. It probes for your
inference server, writes the config, and refuses to write one that cannot work.
Four ways to run it
| Command | What it does |
|---|---|
| oara setup | The rich wizard: identity files, gateway, smoke test. The default, and what most people want. |
| --fast | Probe, write yaml and an env template, done. Three questions. No identity generation, no smoke test. |
| --noninteractive | Implies --fast. Zero questions: first detected server, CLI gateway. |
| --gateway-only | Add or change a messaging gateway on an install that is already set up. |
Prometheus previously shipped two competing init paths with different outputs and
no guidance about which to use. They are one subcommand now; the old
--setup flag and the prometheus-init script are thin
forwards to it.
The fast path, end to end
This is a real run against a llama.cpp server on 8080. Nothing is edited for effect except the tool-registration log, which is noted below.
oara setup --fast --noninteractive ┌─ Prometheus setup (fast) ─────────────────────────────────────┐ Config will be written to ~/.prometheus/prometheus.yaml └───────────────────────────────────────────────────────────────┘ Probing for local inference servers … Local inference: 1 server(s) detected: • llama.cpp @ http://localhost:8080 2ms (1 models, first: gemma-3-27b-it-Q4_K_M) Env template written to ~/.config/prometheus/env ==================================================================== CONNECT A CLIENT (Beacon) Beacon is the desktop cockpit for this daemon (chat, coding runs, documents, dashboards). Get it: https://oara.ai/beacon Address: <this machine>:8005 (or this machine's Tailscale / LAN address, port 8005) Token: minted on first daemon start — re-print with `oara token show` ==================================================================== Setup complete. Next steps: 1. Chat now: prometheus 2. Always-on: oara daemon Health check anytime: oara doctor
prometheus rather than
oara; both work, and the alias is on its way out.
What it wrote
One config file. It is plain YAML and you are meant to read it.
model: provider: llama_cpp base_url: http://localhost:8080 model: gemma-3-27b-it-Q4_K_M grammar_enforcement: true max_tool_iterations: 500 context: effective_limit: 24000 compression_trigger: 0.75 reserved_output: 2000 security: permission_mode: default workspace_root: ~/.prometheus/workspace web: enabled: true api_port: 8005 ws_port: 8010 (gateways, deferred tool loading and learning settings omitted — see Configuration)
Plus an env template at ~/.config/prometheus/env. Secrets go there,
never in the YAML and never on a command line.
No server found
If nothing is listening, setup tells you exactly what it checked and writes nothing:
Probing for local inference servers … Local inference: no servers detected on standard ports. Checked llama.cpp:8080, Ollama:11434, LM Studio:1234, vLLM:8000. Already running a server on another port or machine? oara setup --probe-url http://host:port To run a local model, install one of: Ollama (easiest): curl -fsSL https://ollama.com/install.sh | sh ollama pull qwen3:8b # or any model that fits your hardware llama.cpp (fastest): git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp && cmake -B build && cmake --build build -j ./build/bin/llama-server -m models/your-model.gguf -c 32768 --port 8080 Then run `oara setup` again — it will detect the server. No config was written — nothing usable was detected, and Prometheus refuses to write a config that is known to be broken.
It exits 2 — not zero, so a script notices, and not a crash either.
No file is created. Three ways forward, and the output above already gives you the
first one:
Start a server
Any of the four on its usual port, then re-run.
Point at a server somewhere else
A non-standard port, another machine, a CI stub. Both API shapes are tried at the URL you give.
oara setup --fast --probe-url http://gpu-box:8080
Use a cloud provider instead
Seven have presets. Each knows its own default environment variable and model.
| --provider | Reads | Default model |
|---|---|---|
| anthropic | ANTHROPIC_API_KEY | claude-sonnet-4-6 |
| openai | OPENAI_API_KEY | gpt-4o |
| deepseek | DEEPSEEK_API_KEY | deepseek-v4-flash |
| kimi | MOONSHOT_API_KEY | kimi-k2.6 |
| glm | ZAI_API_KEY | glm-5.2 |
| mimo | MIMO_API_KEY | mimo-v2.5-pro |
| qwen | QWEN_API_KEY | qwen3.7-max |
export ANTHROPIC_API_KEY=... oara setup --provider anthropic
--provider reads it from your environment or from the env file, and
there is no flag to pass it inline. If the variable is not set, setup says so,
tells you where the env file is, and writes nothing:
“a cloud config without its key is known-broken, and Prometheus refuses
to write a config that cannot work.” Use --api-key-env VAR
to read from a different variable, and --model to override the preset.
Then check it
oara doctor reads the config you just wrote
and tells you what is actually true about this machine.