Providers
Ollama
Uses the Ahena CLI (beta). Install it with npm install -g @ahena/cli, or prefix commands with npx @ahena/cli. See Installation. The dashboard covers projects, connections, Doctor, the Stack Graph, plans and approvals without it.
Overview
Ollama runs models on your machine, your LAN or Tailscale. Ahena treats it as a local
provider: the CLI verifies and checks it on your machine and reports the results; Ahena's
servers never contact the endpoint. It generates an adapter behind the common ahena.ai
interface, so development can use Ollama while production uses a hosted provider.
Package: @ahena/provider-ollama · Category: ai · Locality: local · Default model: llama3.2
Connection
OLLAMA_HOST (not secret), e.g. http://localhost:11434. Optional OLLAMA_API_KEY for an
authenticating proxy; it's read from your local environment when Doctor runs.
ahena connect ollama --set OLLAMA_HOST=http://localhost:11434
Permissions
None at Ahena. Ollama itself has no authentication: never expose it to the internet without an authenticating proxy (Doctor fails public plain-http endpoints).
Capabilities
| Capability | What Ahena does |
|---|---|
text-generation |
Generates src/ahena/ai/providers/ollama.ts behind ahena.ai.generate(). |
local-inference |
Checks endpoint, exposure, model, capabilities, context length and latency on your machine. |
Limitations
- Ahena can't check it from the dashboard or CI; checks run where the CLI runs.
- Local models differ from hosted ones in quality, speed and context length.
Manual steps
- Install Ollama and start it (
ollama serve). - Pull your model:
ollama pull llama3.2. - Set
AI_PROVIDER=ollamain development if several adapters exist.
Doctor checks
| Id | Severity | Meaning |
|---|---|---|
ollama.endpoint |
PASS / FAIL | Reachable (with version), or how to start it. |
ollama.exposure |
INFO / WARNING / FAIL | LAN/Tailscale (INFO); public https (WARNING: needs an auth proxy); public plain http (FAIL). |
ollama.model |
PASS / FAIL | Model pulled. |
ollama.capabilities / ollama.context |
PASS / FAIL / WARNING | Model can generate text; context length (WARNING under 8K). |
ollama.latency |
PASS / WARNING / FAIL | One-token generation, from Ollama's own timings, excluding load time. |
Disconnect behavior
Removes the connection. Nothing on your machine is changed.