Step 1: Install the operator
The self-healing loop is on by default as of chart 0.8.0. AI is the part you opt into, and it is not part of this walkthrough.--version resolves the newest published chart. To pin an exact chart version, see the installation guide. Add --set healthCheck.interval=30s to tighten the loop for the demo; the default is 60s. Metrics-server integration is on by default and adds container-level usage signals.
To install the operator without the detection loop, add
--set healthCheck.enabled=false. It will still validate personas and discover the cluster, but it will never open an incident or propose a fix.Phase 2a/2b health check and remediation enabled with the interval. If it is absent, detection is off.
Step 1b: Optional, turn on AI
Everything in this guide works without it. When you do want it, AI setup is the walkthrough, and the startup lines to look for areAI diagnosis enabled and AI remediation planning enabled.
Step 2: Confirm cluster discovery
The operator auto-creates a cluster-scopedClusterPersona named dorgu-cluster on startup and populates it with nodes, add-ons, capacity, Kubernetes version, and platform type — typically within seconds.
dorgu health should report your nodes Ready, resource saturation, control-plane health, and Active Incidents: 0.
On a managed cluster (EKS, GKE, AKS) the control plane is not visible as pods, so it is reported as
Healthy (external/managed). That is expected.dorgu health also prints an Unmonitored section naming the Deployments no persona covers. That is the blind spot, and step 3b closes it.
To manage the ClusterPersona yourself via GitOps instead, install with --set operator.autoCreateClusterPersona=false and apply your own.
Step 3: Give the operator something to watch
The operator only acts on applications that have anApplicationPersona, and it correlates a signal to a persona by name prefix, so the persona name must be a prefix of the pod names.
.status.validation.issues — missing resource requests, an unenforced runAsNonRoot, a missing liveness probe. The phase reflects the overall state (Active, Degraded, or Failed), and it re-validates every 60 seconds or immediately when the Deployment changes. See validation.
No label is required on the Deployment. The operator resolves a persona to its workload through an ordered chain: the
app.kubernetes.io/name label on the Deployment, then the app label, then metadata.name, then spec.selector.matchLabels. Helm, kustomize, and most hand-written YAML label the pod template only, and that resolves through the last rung. When it cannot resolve, the persona reports NoDeployment or AmbiguousDeployment and names what it tried: see troubleshooting.Step 3b: Brownfield, a cluster that already has apps
Writing a persona by hand is fine for one app. It is not how you onboard a cluster that already runs twenty of them, and until a persona exists the operator sees nothing: no persona means no incident, which means no proposed fix.dorgu persona import reads the live Deployments and synthesizes a persona for each from what is already in the spec: resources, probes, replicas, ports, image, and ownership labels. No local source, no Dockerfile, no relabelling.
Step 4: Break something and watch the loop
The fastest way to see self-healing work is a deliberate OOM. The CLI quickstart has a ready-to-apply manifest: a container that allocates ~90 MB against a 64Mi limit, plus a matching persona. Once it is crash-looping, one detection cycle later:Step 5: Approve, and watch it heal
ApplicationPersona spec — the app’s desired state — and moves the remediation to Applying. The CLI then patches the Deployment with your credentials so the pod actually recovers.
That last part happens only when nothing else owns the Deployment. The demo app is applied with kubectl apply, so it is unmanaged and Dorgu heals it. On an owned workload the CLI declines instead, exits 4, and hands you the owner-shaped steps: nothing is approved and nothing changes. See the ownership model.
The remediation stays in
Applying / Verifying for the verification window — 10 minutes by default — before it reports Completed. At the end of that window the operator re-runs detection and rolls the change back automatically if health regressed. Your pod is healthy well before then.Next steps
Ownership model
Which workloads Dorgu patches, which it only recommends for, and why
Security and permissions
The operator’s ClusterRole, and the verbs it deliberately lacks
Self-healing in depth
Signals, guardrails, verification, and troubleshooting
AI setup
Key handling, verification, and turning AI off
Validation details
What the operator validates and how severity works
Configuration
Flags for webhooks, ArgoCD, Prometheus, and WebSocket
dorgu persona import
Create personas from the Deployments you already run
Uninstall
Complete teardown.
helm uninstall on its own leaves your API key behind