Documentation menu

Providers

Ollama

Uses the Ahena CLI (beta). Install it with npm install -g @ahena/cli, or prefix commands with npx @ahena/cli. See Installation. The dashboard covers projects, connections, Doctor, the Stack Graph, plans and approvals without it.

Overview

Ollama runs models on your machine, your LAN or Tailscale. Ahena treats it as a local provider: the CLI verifies and checks it on your machine and reports the results; Ahena's servers never contact the endpoint. It generates an adapter behind the common ahena.ai interface, so development can use Ollama while production uses a hosted provider.

Package: @ahena/provider-ollama · Category: ai · Locality: local · Default model: llama3.2

Connection

OLLAMA_HOST (not secret), e.g. http://localhost:11434. Optional OLLAMA_API_KEY for an authenticating proxy; it's read from your local environment when Doctor runs.

ahena connect ollama --set OLLAMA_HOST=http://localhost:11434

Permissions

None at Ahena. Ollama itself has no authentication: never expose it to the internet without an authenticating proxy (Doctor fails public plain-http endpoints).

Capabilities

Capability What Ahena does
text-generation Generates src/ahena/ai/providers/ollama.ts behind ahena.ai.generate().
local-inference Checks endpoint, exposure, model, capabilities, context length and latency on your machine.

Limitations

  • Ahena can't check it from the dashboard or CI; checks run where the CLI runs.
  • Local models differ from hosted ones in quality, speed and context length.

Manual steps

  1. Install Ollama and start it (ollama serve).
  2. Pull your model: ollama pull llama3.2.
  3. Set AI_PROVIDER=ollama in development if several adapters exist.

Doctor checks

Id Severity Meaning
ollama.endpoint PASS / FAIL Reachable (with version), or how to start it.
ollama.exposure INFO / WARNING / FAIL LAN/Tailscale (INFO); public https (WARNING: needs an auth proxy); public plain http (FAIL).
ollama.model PASS / FAIL Model pulled.
ollama.capabilities / ollama.context PASS / FAIL / WARNING Model can generate text; context length (WARNING under 8K).
ollama.latency PASS / WARNING / FAIL One-token generation, from Ollama's own timings, excluding load time.

Disconnect behavior

Removes the connection. Nothing on your machine is changed.