Skip to main content
Dorgu’s AI features are optional, off by default, and bring-your-own-key. Detection, diagnosis, and rule-based remediation all work with no key at all. The AI layer adds root-cause enhancement and the ordered, multi-step remediation plans described in Self-healing.

Why it is off, and what that costs you

The AI planner is off by default, and for now that is also the recommendation. aiRemediation.enabled defaults to false, llm.provider defaults to empty, and the install commands in these docs leave both alone. A default install spends nothing on inference and sends nothing out of your cluster.Rules always run. Detection, rule-based diagnosis, and rule-based remediation are local, free, deterministic, and are the floor the whole product stands on. They are what the clean-room runs have actually healed with.The AI planner is an addition on top of that floor, and across three consecutive clean-room evaluations, meaning a stranger installing from these published docs and nothing else, it has been the weaker path: it invented specifics in the first, wrote a worse plan than the rules in the second, and in the third produced nine remediations of which none could change a workload. Operator v0.11.0 is the release that fixed that last one, and it is too recent to have an evaluation behind it yet.This is not a retreat from AI. It is where the two halves have earned their place: deterministic detection and remediation, with AI for explanation. The setting that says exactly that is a provider with no planner, which gives you AI root-cause prose while the numbers that reach your cluster stay the rule engine’s:
AI setup has the full walkthrough for both, and what each one costs you.
There are two knobs, not one, and they are worth separating before you set either:
The guardrails do not care which wrote the plan. Blast radius, the no-new-keys rule, rate limits, and human approval all apply identically to an AI plan and a rule-based one, and since operator v0.11.0 a value the model asked for that a guardrail refuses is replaced by one Dorgu computed, recorded on spec.steps[].safety. Turning the planner on widens what Dorgu will say. It does not widen what Dorgu will do.

What you are opting into

  • BYO key. Dorgu does not proxy your traffic. You supply an Anthropic API key and pay Anthropic directly.
  • Anthropic only, operator-side. The operator’s AI diagnosis accepts claude or gemini; the AI remediation planner is Anthropic-only today and is skipped unless llm.provider=claude.
  • The key stays in your cluster. It lives in a Kubernetes Secret in your namespace and is injected into the operator pod as the ANTHROPIC_API_KEY environment variable via secretKeyRef. The chart never passes it as a container argument.
  • Your cluster data goes to Anthropic only when AI is enabled. With no provider configured, nothing leaves the cluster.
  • No key, no problem. Rule-based detection, diagnosis, and remediation are the floor and are always on.

Enable it

1

Create the Secret

Create the Secret out of band so the key never passes through Helm values:
If the namespace does not exist yet, create it first with kubectl create namespace dorgu-system.
2

Install or upgrade the operator

This is the canonical install command plus the four AI flags, which are the only ones AI needs. Detection and metrics-server integration are already on by default; add --set websocket.enabled=true if you also want live dorgu watch streaming.Drop the last flag, aiRemediation.enabled=true, for AI diagnosis with rule-written plans. That is the setting the default recommends, and it is the one to start from.
aiRemediation.enabled=true needs operator v0.11.0 or newer to be worth setting. On v0.10.0 and older, a plan diagnosing a resource change could be persisted with an empty patch and an action type of notification, so dorgu remediation approve printed “No resource change to apply” and healed nothing. Nine of nine AI-planned remediations were unappliable in clean-room run #4. Check what you are running with helm list -n dorgu-system before you turn the planner on. See versions.
Omitting --version resolves the newest published chart. To pin an exact version, see the installation guide.
3

Verify

You are looking for two startup lines:
Both must appear. AI diagnosis enabled on its own means the planner did not start — check that aiRemediation.enabled=true and llm.provider=claude.
Detection is already on. Since chart 0.8.0 healthCheck.enabled defaults to true, so there is no longer a flag to remember here. If you explicitly set healthCheck.enabled=false, AI configuration gives you a correctly configured operator with nothing to diagnose.
AI is deliberately the one thing you have to opt into. Detection and rule-based diagnosis are local and free, so they are on out of the box. Inference costs money and sends incident context to a third party, so that consent stays explicit: no llm.provider, no key, no aiRemediation.enabled, no API calls and no spend.

Values reference

Prefer llm.existingSecret in any shared or production cluster. llm.apiKey never reaches the pod spec — the chart still delivers it via secretKeyRef — but it is recorded in the Helm release, so helm get values and helm template will print the raw key to anyone with access. existingSecret keeps the key out of Helm entirely.
If neither llm.existingSecret nor a chart-managed Secret is configured, the chart renders no env: block, the operator finds no ANTHROPIC_API_KEY, and it logs LLM provider configured but no API key found, AI diagnosis disabled — then runs rule-based.

Egress

The operator calls the Anthropic API directly from its pod. On a private cluster it needs outbound HTTPS — through a NAT gateway, an egress proxy, or whatever your network policy allows. If egress is blocked, AI planning fails and logs AI remediation planning failed, falling back to rules; the loop keeps working on the deterministic path.

Turn it off

Either way the self-healing loop keeps running rule-based. Existing RemediationAction objects with planSource: ai-anthropic remain valid and reviewable; new proposals come back as rule-based. To remove the key from the cluster as well:
That Secret is yours, not Helm’s, so helm uninstall will not remove it. You created it out of band with kubectl create secret so the key would never pass through Helm values, and the price of that is that removing the release leaves the key in the cluster. Whether you are turning AI off or removing Dorgu entirely, delete it explicitly. The full teardown, which leads with exactly this, is Uninstall.

CLI-side LLMs are separate

The operator’s AI is unrelated to the CLI’s LLM configuration. dorgu generate can use OpenAI, Anthropic, Gemini, or Ollama from your laptop; none of that configuration reaches the operator, and the operator’s key is never used by the CLI. See LLM providers for the full comparison.

Self-healing

What the AI planner actually receives and produces

Helm values

Every chart value in one place