Prerequisites:
- A Kubernetes cluster you can install into, and
kubectlpointed at it. Every Dorgu cluster command shells out tokubectl. - Helm 3
- No API key. This walkthrough runs entirely on Dorgu’s deterministic path, which is the recommended default. AI is an opt-in addition covered in step 2b.
1. Install the CLI
2. Install the operator
Detection and rule-based remediation are on out of the box, so this is the whole install. No key, no flags beyond a tighter loop for the demo.healthCheck.interval=30s only tightens the demo loop; the default is 60s. Everything below heals on this install.
Failure detection is enabled by default as of chart 0.8.0, so there is no flag to remember. To run the operator without it, install with
--set healthCheck.enabled=false.ClusterPersona named dorgu-cluster on startup and discovers your nodes, add-ons, and capacity:
2b. Optional: turn on AI
Skip this on a first run. The rest of this walkthrough does not need it, and the deterministic path is the one to see working first. If you do want the AI root-cause prose while you are here:--set aiRemediation.enabled=true hands the plan to the model as well, and needs operator v0.11.0 or newer to produce a plan that can actually be applied. AI setup covers both, including how to confirm which of the two started.
3. Deploy an app that will break
Save this asdemo-oom.yaml. It is a container that allocates ~90 MB against a 64Mi limit, so the kernel OOM-kills it within seconds.
OOMKilled, then CrashLoopBackOff. Left alone, Kubernetes would restart it forever.
This demo heals end to end because nothing reconciles it. You applied it with
kubectl apply, so Dorgu records managedBy: unmanaged and is free to patch the Deployment.On your real apps that is often not the case. For a Deployment that Helm, ArgoCD, or Flux owns, Dorgu detects and diagnoses exactly the same but will not patch it: it names the owner and tells you what to change in your source of truth instead. Read the ownership model before you point this at anything you deploy with Helm.A kustomize overlay is the exception, and Dorgu says so before it writes: it leaves no marker on what it renders, so kubectl apply -k output reads as unmanaged too. Dorgu will patch it, and your next kubectl apply -k will revert the patch rather than fail on it. See the kustomize limitation.4. Watch Dorgu detect and diagnose it
Wait one detection cycle — 30 seconds with the settings above — then:OOMKilled and CrashLoopBackOff. Two symptoms, one root cause. Read the diagnosis:
rule-based on this install, ai-enhanced if you enabled a provider in step 2b.
5. Review the proposed fix
PLAN: rule-based, or ai-anthropic if you enabled the planner in step 2b. Now read the actual plan:
dorgu remediation for a full annotated example.
6. Approve — and watch it heal
ApplicationPersona spec — the app’s desired-state record. Then the CLI shows you exactly which Deployment, container, and fields it is about to change, asks for confirmation, and patches the Deployment with your credentials. Add --yes to skip the prompt.
1/1 Running with no new restarts. That is the aha.
7. It is remembered
The remediation will read
Applying or Verifying for the next ~10 minutes, and the incident stays open until then. That is the verification window: the operator re-runs detection at the end of it and automatically rolls the change back if health regressed. Your pod is already healthy — the wait is what makes the rollback guarantee real.Clean up
The demo app and its persona:helm uninstall, which leaves the five CRDs, every record in every namespace, and any API key Secret you created. The complete teardown, in order, is Uninstall.
Brownfield: a cluster that already has apps
The walkthrough above starts from a persona you wrote by hand. Real clusters do not work that way: the apps are already running, and nobody is going to hand-write a persona for each one. Dorgu only watches workloads that have anApplicationPersona. Without one, a broken app raises no incident and gets no proposed fix. dorgu persona import closes that gap by reading the Deployments you already have and synthesizing a persona for each from what is already in the spec. No local source, no Dockerfile, no relabelling.
1
Install onto the cluster as it is
Follow steps 1 and 2 above. Nothing about an existing cluster needs special handling at install time.
2
Ask Dorgu what it cannot see
Unmonitored section names every Deployment with no matching persona, and prints one import command per namespace:kube-system are left out unless you ask for one with -n. A cluster with nothing to report prints no section at all.3
Read the personas before you apply them
Import prints YAML to stdout and sends every diagnostic to stderr, so redirecting is safe:Read the warnings on your terminal. The one that matters most is inferred resource limits: the remediation proposer skips any persona without limits, so where a container declared none, Dorgu derives them and says so. A persona healing against numbers nobody chose is worse than no persona.
4
Apply them
Active means the persona found its workload. Phase Pending with reason NoDeployment or AmbiguousDeployment means it did not: see troubleshooting.5
Break something and watch the loop
Pick any imported app and give it a limit it cannot live within:One detection cycle later the loop runs exactly as it does in step 4 above:Read the
Owner: line in that diff before you approve. On a real cluster most of these apps are deployed by something, and what happens next depends on it.6
Approve, or apply it where the workload lives
If the diff says If it names a Helm release, an ArgoCD application, or a Flux resource, Dorgu will not patch it, and the diff offers Running
Owner: unmanaged, approve and Dorgu heals the Deployment:--no-heal and reject rather than approve. Patching an owned Deployment claims those fields away from its owner, and the next helm upgrade or sync then fails outright rather than quietly reverting. So make the change at the source the plan points you to, then record the decision:approve without --no-heal on an owned workload is safe: it writes nothing, prints the owner and the owner-shaped steps, and exits 4. See the ownership model.No label is required on your Deployments. Dorgu resolves a persona to its Deployment by an ordered chain: the
app.kubernetes.io/name label, then the app label, then metadata.name, then spec.selector.matchLabels. Helm, kustomize, and most hand-written YAML label the pod template only, and that is fine. persona import picks a spec.name that resolves back to the Deployment it came from, and tells you when it cannot.dorgu persona import requires CLI v0.9.0 or newer. See dorgu persona import for every flag.
Also: generate manifests
Manifest generation is Dorgu’s other half. Point the CLI at any app with aDockerfile or docker-compose.yml:
--dry-run to print instead of write, and --llm-provider openai (or anthropic, gemini, ollama) for LLM-enhanced analysis. See the manifest generation guide.
Next steps
Self-healing in depth
Detection signals, guardrails, verification, and troubleshooting
dorgu remediation
Every flag for list, diff, approve, reject, and heal
AI setup
Key handling, verification, and how to turn AI off
Working with personas
Write, generate, or import the personas the loop depends on
dorgu persona import
Onboard a cluster that already has apps running
dorgu health
The unmonitored section, exit codes, and JSON output
Ownership model
Which workloads Dorgu patches, which it only recommends for, and why
Security and permissions
The operator’s ClusterRole, published in full