Skip to main content
This is the fastest path to seeing Dorgu work: install it, deliberately deploy an app with too little memory, and watch the loop detect the OOM, diagnose it, propose a fix, and heal the workload once you approve.
Prerequisites:
  • A Kubernetes cluster you can install into, and kubectl pointed at it. Every Dorgu cluster command shells out to kubectl.
  • Helm 3
  • No API key. This walkthrough runs entirely on Dorgu’s deterministic path, which is the recommended default. AI is an opt-in addition covered in step 2b.
Use a development cluster. dorgu remediation heal refuses to run if your kube-context name contains prod, but the safest habit is to try this somewhere disposable.

1. Install the CLI

See installation for pinned versions, checksums, and Windows.

2. Install the operator

Detection and rule-based remediation are on out of the box, so this is the whole install. No key, no flags beyond a tighter loop for the demo.
healthCheck.interval=30s only tightens the demo loop; the default is 60s. Everything below heals on this install.
Failure detection is enabled by default as of chart 0.8.0, so there is no flag to remember. To run the operator without it, install with --set healthCheck.enabled=false.
The operator auto-creates a ClusterPersona named dorgu-cluster on startup and discovers your nodes, add-ons, and capacity:
Installing onto a cluster that already has apps? Steps 3 to 7 deploy a toy app to demonstrate the loop. If you want Dorgu working on the workloads you already run, read Brownfield: a cluster that already has apps instead. dorgu health has already told you which Deployments it cannot see.

2b. Optional: turn on AI

Skip this on a first run. The rest of this walkthrough does not need it, and the deterministic path is the one to see working first.
The AI planner is off by default, and for now that is also the recommendation. aiRemediation.enabled defaults to false, llm.provider defaults to empty, and the install commands in these docs leave both alone. A default install spends nothing on inference and sends nothing out of your cluster.Rules always run. Detection, rule-based diagnosis, and rule-based remediation are local, free, deterministic, and are the floor the whole product stands on. They are what the clean-room runs have actually healed with.The AI planner is an addition on top of that floor, and across three consecutive clean-room evaluations, meaning a stranger installing from these published docs and nothing else, it has been the weaker path: it invented specifics in the first, wrote a worse plan than the rules in the second, and in the third produced nine remediations of which none could change a workload. Operator v0.11.0 is the release that fixed that last one, and it is too recent to have an evaluation behind it yet.This is not a retreat from AI. It is where the two halves have earned their place: deterministic detection and remediation, with AI for explanation. The setting that says exactly that is a provider with no planner, which gives you AI root-cause prose while the numbers that reach your cluster stay the rule engine’s:
AI setup has the full walkthrough for both, and what each one costs you.
If you do want the AI root-cause prose while you are here:
That gives you AI diagnosis with the rule engine still writing the plan. Adding --set aiRemediation.enabled=true hands the plan to the model as well, and needs operator v0.11.0 or newer to produce a plan that can actually be applied. AI setup covers both, including how to confirm which of the two started.

3. Deploy an app that will break

Save this as demo-oom.yaml. It is a container that allocates ~90 MB against a 64Mi limit, so the kernel OOM-kills it within seconds.
Within a few seconds the pod goes OOMKilled, then CrashLoopBackOff. Left alone, Kubernetes would restart it forever.
Four details in that manifest are load-bearing:
  1. The persona is named memhog, exactly like the Deployment. The operator correlates a signal to a persona by name prefix, and pods are memhog-<hash>, so the persona must be memhog. Rename one and no incident gets a persona, and no fix is ever proposed. (Resolving a persona back to its Deployment is separate and needs no label at all: see discovery.)
  2. The persona declares resources.limits. Without a current limit the proposer has nothing to compute a bounded change from, and skips.
  3. The allocation fits after the fix. Increases are capped at 2× the live limit, so 64Mi becomes at most 128Mi, and ~90 MB plus the interpreter fits in 128Mi. One approval heals it.
  4. The image is a manifest list. python:3.12-alpine publishes linux/amd64 and linux/arm64, so this manifest runs unchanged on an AWS Graviton nodegroup and on kind or k3d on an Apple Silicon Mac. A single-platform amd64 image here does not fail cleanly: on containerd the pull succeeds and the container exits 255 with exec format error, which reads as CrashLoopBackOff and gets diagnosed as something else entirely.
This demo heals end to end because nothing reconciles it. You applied it with kubectl apply, so Dorgu records managedBy: unmanaged and is free to patch the Deployment.On your real apps that is often not the case. For a Deployment that Helm, ArgoCD, or Flux owns, Dorgu detects and diagnoses exactly the same but will not patch it: it names the owner and tells you what to change in your source of truth instead. Read the ownership model before you point this at anything you deploy with Helm.A kustomize overlay is the exception, and Dorgu says so before it writes: it leaves no marker on what it renders, so kubectl apply -k output reads as unmanaged too. Dorgu will patch it, and your next kubectl apply -k will revert the patch rather than fail on it. See the kustomize limitation.

4. Watch Dorgu detect and diagnose it

Wait one detection cycle — 30 seconds with the settings above — then:
You will see incidents for OOMKilled and CrashLoopBackOff. Two symptoms, one root cause. Read the diagnosis:
The Root Cause section carries the summary, the confidence score, and the provider: rule-based on this install, ai-enhanced if you enabled a provider in step 2b.

5. Review the proposed fix

One row, reading PLAN: rule-based, or ai-anthropic if you enabled the planner in step 2b. Now read the actual plan:
This is the review surface: the plan summary, then each step in execution order with its rationale, risk level, whether it is auto-applied or advisory, and a YAML diff of the change. The first step raises the memory limit; the follow-up steps are advisory. See dorgu remediation for a full annotated example.

6. Approve — and watch it heal

Two things happen. The operator patches the ApplicationPersona spec — the app’s desired-state record. Then the CLI shows you exactly which Deployment, container, and fields it is about to change, asks for confirmation, and patches the Deployment with your credentials. Add --yes to skip the prompt.
The pod comes back 1/1 Running with no new restarts. That is the aha.

7. It is remembered

The records persist in your cluster. Next time a similar problem appears, the planner has this incident — and whether this fix actually worked — as context.
The remediation will read Applying or Verifying for the next ~10 minutes, and the incident stays open until then. That is the verification window: the operator re-runs detection at the end of it and automatically rolls the change back if health regressed. Your pod is already healthy — the wait is what makes the rollback guarantee real.

Clean up

The demo app and its persona:
Removing Dorgu itself takes more than helm uninstall, which leaves the five CRDs, every record in every namespace, and any API key Secret you created. The complete teardown, in order, is Uninstall.

Brownfield: a cluster that already has apps

The walkthrough above starts from a persona you wrote by hand. Real clusters do not work that way: the apps are already running, and nobody is going to hand-write a persona for each one. Dorgu only watches workloads that have an ApplicationPersona. Without one, a broken app raises no incident and gets no proposed fix. dorgu persona import closes that gap by reading the Deployments you already have and synthesizing a persona for each from what is already in the spec. No local source, no Dockerfile, no relabelling.
1

Install onto the cluster as it is

Follow steps 1 and 2 above. Nothing about an existing cluster needs special handling at install time.
2

Ask Dorgu what it cannot see

The Unmonitored section names every Deployment with no matching persona, and prints one import command per namespace:
Cluster add-on namespaces such as kube-system are left out unless you ask for one with -n. A cluster with nothing to report prints no section at all.
3

Read the personas before you apply them

Import prints YAML to stdout and sends every diagnostic to stderr, so redirecting is safe:
Read the warnings on your terminal. The one that matters most is inferred resource limits: the remediation proposer skips any persona without limits, so where a container declared none, Dorgu derives them and says so. A persona healing against numbers nobody chose is worse than no persona.
4

Apply them

Or skip the file and let Dorgu apply directly:
Then confirm the operator resolved each one to its Deployment:
Phase Active means the persona found its workload. Phase Pending with reason NoDeployment or AmbiguousDeployment means it did not: see troubleshooting.
5

Break something and watch the loop

Pick any imported app and give it a limit it cannot live within:
One detection cycle later the loop runs exactly as it does in step 4 above:
Read the Owner: line in that diff before you approve. On a real cluster most of these apps are deployed by something, and what happens next depends on it.
6

Approve, or apply it where the workload lives

If the diff says Owner: unmanaged, approve and Dorgu heals the Deployment:
If it names a Helm release, an ArgoCD application, or a Flux resource, Dorgu will not patch it, and the diff offers --no-heal and reject rather than approve. Patching an owned Deployment claims those fields away from its owner, and the next helm upgrade or sync then fails outright rather than quietly reverting. So make the change at the source the plan points you to, then record the decision:
Running approve without --no-heal on an owned workload is safe: it writes nothing, prints the owner and the owner-shaped steps, and exits 4. See the ownership model.
No label is required on your Deployments. Dorgu resolves a persona to its Deployment by an ordered chain: the app.kubernetes.io/name label, then the app label, then metadata.name, then spec.selector.matchLabels. Helm, kustomize, and most hand-written YAML label the pod template only, and that is fine. persona import picks a spec.name that resolves back to the Deployment it came from, and tells you when it cannot.
dorgu persona import requires CLI v0.9.0 or newer. See dorgu persona import for every flag.
The ownership guard needs CLI v0.10.0 or newer. The operator records who owns each Deployment in spec.workloadRef from v0.9.0 onward; the CLI is what refuses to patch an owned one. An older CLI ignores the record and patches anyway, which is exactly the helm upgrade breakage the guard exists to prevent, and it is the only version mismatch between the two halves that is genuinely unsafe.Check both before your first approval on a real cluster:
Which versions to be on, and what every other mismatch does: versions.

Also: generate manifests

Manifest generation is Dorgu’s other half. Point the CLI at any app with a Dockerfile or docker-compose.yml:
Dorgu analyzes the app and writes a Deployment, Service, Ingress, HPA, ArgoCD Application, a GitHub Actions workflow, and an ApplicationPersona — the same kind of persona the self-healing loop above operates on. Use --dry-run to print instead of write, and --llm-provider openai (or anthropic, gemini, ollama) for LLM-enhanced analysis. See the manifest generation guide.

Next steps

Self-healing in depth

Detection signals, guardrails, verification, and troubleshooting

dorgu remediation

Every flag for list, diff, approve, reject, and heal

AI setup

Key handling, verification, and how to turn AI off

Working with personas

Write, generate, or import the personas the loop depends on

dorgu persona import

Onboard a cluster that already has apps running

dorgu health

The unmonitored section, exit codes, and JSON output

Ownership model

Which workloads Dorgu patches, which it only recommends for, and why

Security and permissions

The operator’s ClusterRole, published in full