Skip to main content
This guide installs the Dorgu Operator with the self-healing loop running on its deterministic path, then walks a real incident from detection to healed. No API key is needed at any point.

Step 1: Install the operator

The self-healing loop is on by default as of chart 0.8.0. AI is the part you opt into, and it is not part of this walkthrough.
Omitting --version resolves the newest published chart. To pin an exact chart version, see the installation guide. Add --set healthCheck.interval=30s to tighten the loop for the demo; the default is 60s. Metrics-server integration is on by default and adds container-level usage signals.
To install the operator without the detection loop, add --set healthCheck.enabled=false. It will still validate personas and discover the cluster, but it will never open an incident or propose a fix.
Confirm the loop started:
Look for Phase 2a/2b health check and remediation enabled with the interval. If it is absent, detection is off.

Step 1b: Optional, turn on AI

The AI planner is off by default, and for now that is also the recommendation. aiRemediation.enabled defaults to false, llm.provider defaults to empty, and the install commands in these docs leave both alone. A default install spends nothing on inference and sends nothing out of your cluster.Rules always run. Detection, rule-based diagnosis, and rule-based remediation are local, free, deterministic, and are the floor the whole product stands on. They are what the clean-room runs have actually healed with.The AI planner is an addition on top of that floor, and across three consecutive clean-room evaluations, meaning a stranger installing from these published docs and nothing else, it has been the weaker path: it invented specifics in the first, wrote a worse plan than the rules in the second, and in the third produced nine remediations of which none could change a workload. Operator v0.11.0 is the release that fixed that last one, and it is too recent to have an evaluation behind it yet.This is not a retreat from AI. It is where the two halves have earned their place: deterministic detection and remediation, with AI for explanation. The setting that says exactly that is a provider with no planner, which gives you AI root-cause prose while the numbers that reach your cluster stay the rule engine’s:
AI setup has the full walkthrough for both, and what each one costs you.
Everything in this guide works without it. When you do want it, AI setup is the walkthrough, and the startup lines to look for are AI diagnosis enabled and AI remediation planning enabled.

Step 2: Confirm cluster discovery

The operator auto-creates a cluster-scoped ClusterPersona named dorgu-cluster on startup and populates it with nodes, add-ons, capacity, Kubernetes version, and platform type — typically within seconds.
dorgu health should report your nodes Ready, resource saturation, control-plane health, and Active Incidents: 0.
On a managed cluster (EKS, GKE, AKS) the control plane is not visible as pods, so it is reported as Healthy (external/managed). That is expected.
If the cluster already runs applications, dorgu health also prints an Unmonitored section naming the Deployments no persona covers. That is the blind spot, and step 3b closes it. To manage the ClusterPersona yourself via GitOps instead, install with --set operator.autoCreateClusterPersona=false and apply your own.

Step 3: Give the operator something to watch

The operator only acts on applications that have an ApplicationPersona, and it correlates a signal to a persona by name prefix, so the persona name must be a prefix of the pod names.
Or apply one directly:
Always declare resources.limits. Without a current limit the remediation proposer has no baseline to compute a bounded change from, and skips the persona with persona has no resource limits configured.
The operator validates the running Deployment against the persona and reports issues in .status.validation.issues — missing resource requests, an unenforced runAsNonRoot, a missing liveness probe. The phase reflects the overall state (Active, Degraded, or Failed), and it re-validates every 60 seconds or immediately when the Deployment changes. See validation.
No label is required on the Deployment. The operator resolves a persona to its workload through an ordered chain: the app.kubernetes.io/name label on the Deployment, then the app label, then metadata.name, then spec.selector.matchLabels. Helm, kustomize, and most hand-written YAML label the pod template only, and that resolves through the last rung. When it cannot resolve, the persona reports NoDeployment or AmbiguousDeployment and names what it tried: see troubleshooting.

Step 3b: Brownfield, a cluster that already has apps

Writing a persona by hand is fine for one app. It is not how you onboard a cluster that already runs twenty of them, and until a persona exists the operator sees nothing: no persona means no incident, which means no proposed fix. dorgu persona import reads the live Deployments and synthesizes a persona for each from what is already in the spec: resources, probes, replicas, ports, image, and ownership labels. No local source, no Dockerfile, no relabelling.
Pay attention to the inferred-limits warning. Where a container declares no resource limits, import derives them and says which ones it invented. The remediation proposer skips any persona without limits, so a persona with limits nobody chose will heal against the wrong numbers. Review them before you rely on them.
The full walkthrough, including breaking an imported app and healing it, is in the CLI quickstart. Requires CLI v0.9.0 or newer.
On a brownfield cluster, expect most of these workloads to be owned. For a Deployment that Helm, ArgoCD, or Flux reconciles, Dorgu detects and diagnoses exactly as it does for anything else, but it is recommendation only: it names the owner and tells you what to change in your source of truth rather than patching the Deployment. Patching an owned workload claims those fields away from its owner, and the next helm upgrade then fails outright rather than quietly reverting.Read the ownership model before your first approval on a real cluster. The guard needs CLI v0.10.0 or newer: an older CLI ignores the ownership record and patches an owned Deployment regardless. See versions.

Step 4: Break something and watch the loop

The fastest way to see self-healing work is a deliberate OOM. The CLI quickstart has a ready-to-apply manifest: a container that allocates ~90 MB against a 64Mi limit, plus a matching persona. Once it is crash-looping, one detection cycle later:

Step 5: Approve, and watch it heal

The operator patches the ApplicationPersona spec — the app’s desired state — and moves the remediation to Applying. The CLI then patches the Deployment with your credentials so the pod actually recovers. That last part happens only when nothing else owns the Deployment. The demo app is applied with kubectl apply, so it is unmanaged and Dorgu heals it. On an owned workload the CLI declines instead, exits 4, and hands you the owner-shaped steps: nothing is approved and nothing changes. See the ownership model.
The remediation stays in Applying / Verifying for the verification window — 10 minutes by default — before it reports Completed. At the end of that window the operator re-runs detection and rolls the change back automatically if health regressed. Your pod is healthy well before then.
The operator never creates or modifies Deployments, Services, or any other workload resource. It updates Persona, Incident, and Remediation CRDs and recommends; applying workload changes stays with the CLI, ArgoCD, Flux, or kubectl. That is RBAC, not convention: its ClusterRole grants only get, list, watch on Deployments, and no access to Secrets. Check it yourself on security and permissions.

Next steps

Ownership model

Which workloads Dorgu patches, which it only recommends for, and why

Security and permissions

The operator’s ClusterRole, and the verbs it deliberately lacks

Self-healing in depth

Signals, guardrails, verification, and troubleshooting

AI setup

Key handling, verification, and turning AI off

Validation details

What the operator validates and how severity works

Configuration

Flags for webhooks, ArgoCD, Prometheus, and WebSocket

dorgu persona import

Create personas from the Deployments you already run

Uninstall

Complete teardown. helm uninstall on its own leaves your API key behind