Skip to main content

Prerequisites

  • Kubernetes cluster 1.24+
  • kubectl configured for your cluster
  • Helm 3.x (for Helm installation)
  • Chart 0.11.1 or newer on arm64 nodes, including AWS Graviton and Apple Silicon kind or k3d clusters. Every image up to and including 0.11.0 was amd64-only. See Versions.
Install the operator from the OCI registry:
That is the canonical install command, and it is the same one used in the CLI quickstart, the operator quickstart, and AI setup.
Failure detection is on out of the box as of chart 0.8.0. The health-check reconciler runs every 60 seconds, opens IncidentMemory records, and proposes RemediationAction fixes for what it finds. Nothing is ever applied without your approval.Earlier charts defaulted healthCheck.enabled to false, which meant following this page produced an operator that detected nothing. If you are reading an older copy of these instructions, that is the difference.

AI

The AI planner is off by default, and for now that is also the recommendation. aiRemediation.enabled defaults to false, llm.provider defaults to empty, and the install commands in these docs leave both alone. A default install spends nothing on inference and sends nothing out of your cluster.Rules always run. Detection, rule-based diagnosis, and rule-based remediation are local, free, deterministic, and are the floor the whole product stands on. They are what the clean-room runs have actually healed with.The AI planner is an addition on top of that floor, and across three consecutive clean-room evaluations, meaning a stranger installing from these published docs and nothing else, it has been the weaker path: it invented specifics in the first, wrote a worse plan than the rules in the second, and in the third produced nine remediations of which none could change a workload. Operator v0.11.0 is the release that fixed that last one, and it is too recent to have an evaluation behind it yet.This is not a retreat from AI. It is where the two halves have earned their place: deterministic detection and remediation, with AI for explanation. The setting that says exactly that is a provider with no planner, which gives you AI root-cause prose while the numbers that reach your cluster stay the rule engine’s:
AI setup has the full walkthrough for both, and what each one costs you.

Versions

Current versions: CLI v0.12.0 with operator v0.11.1 (Helm chart 0.11.1; the chart version, the chart appVersion, and the operator version are always the same number). Newest of each is always the right answer, and this note is the only place in these docs that states which that is.On arm64 nodes, 0.11.1 is a floor and not a preference. Every operator image published before it, up to and including 0.11.0, was pushed as a single-platform amd64 manifest rather than a manifest list. That covers AWS Graviton nodegroups and any kind or k3d cluster on an Apple Silicon Mac. On containerd there is no manifest list to select from, so the pull succeeds and the container dies immediately with exec /manager: exec format error and exit code 255. The pod reads CrashLoopBackOff, not ImagePullBackOff. 0.11.1 is that image and nothing else: no API, CRD, controller, or chart-template change, which is why everything below still names v0.11.0 as the release each behaviour landed in.There is no version pinning between the CLI and the operator. Every field either side has added is optional and additive, so a half that does not know about a field simply sees it absent and renders exactly as it did before. Nothing has to be upgraded in lockstep.Three mismatches change behaviour. Only the first is unsafe, and it is the old CLI, not the old operator:One operator upgrade is worth doing on its own merits, with no CLI implication at all: with aiRemediation.enabled=true, operator v0.11.0 is the release that made AI-planned remediations appliable. On v0.10.0 and older, a plan that diagnosed a resource change could be persisted with nothing to apply. See AI setup.Both mismatch directions in full, with the ownership reasoning behind them: version coupling. Newest published releases: CLI and operator.
Read the ownership model before you upgrade a cluster that runs Helm or ArgoCD-managed apps. Persona writes are unaffected by any of it.

Turn detection off

To run the operator for validation and cluster discovery only:

Pin a chart version

The command above resolves the newest published chart. To pin an exact version for reproducibility, add --version:
The chart version and the operator version are the same number, so --version 0.11.1 installs operator v0.11.1. If the literal above has fallen behind, Versions is the statement to trust on this page, and operator releases is the statement to trust over both.

Install with a values file

To customize the installation, create a values.yaml file:
Then install with your values:
See the Helm values reference for all available options.

Kustomize

Install CRDs and deploy the operator using kustomize:
To reverse it:
make uninstall deletes the CRDs, so it deletes every ApplicationPersona, IncidentMemory, RemediationAction and DorguEvent in the cluster along with them. Read Uninstall before you run it, particularly the part about your API key, which neither command touches.

Build from source

Running locally connects to whatever cluster your kubeconfig points to. Make sure you are targeting the correct cluster before starting the operator.
Running the manager directly does not read the chart’s defaults. Add --enable-health-check to get the detection loop that a Helm install gives you.

Build and push Docker image

Multi-platform builds (arm64, amd64) are supported:

Verify installation

The chart prints a post-install summary that reports whether detection is on, whether AI is on, and what to run next. Re-read it at any time with helm get notes dorgu-operator -n dorgu-system. Check that the operator is running:
Expected output:
Verify all five CRDs are installed:
Expected output:
Fewer than five means an incomplete install. Detection cannot record anything without incidentmemories and cannot propose anything without remediationactions. Installing from a clone of an old chart was the usual cause; the published OCI chart has always carried all five. See Self-healing.
Confirm detection actually started:
Look for Phase 2a/2b health check and remediation enabled with the interval. If it is absent, detection is off.

Uninstall

Start with your Anthropic API key. helm uninstall does not remove it.The AI setup walkthrough has you create the key as a Secret out of band, with kubectl create secret, precisely so it never passes through Helm values. The consequence is that Helm never owned it and never deletes it, so an operator you removed an hour ago leaves your key sitting in the cluster:
Run that first, before anything else on this page, and rotate the key at Anthropic if the cluster was shared or is one you no longer control. The only case where Helm does clean it up is llm.createSecret=true with llm.apiKey set, which the docs recommend against for exactly the reasons in AI setup.

1. Remove the release

That removes everything Helm actually owns: the operator Deployment, the ServiceAccount, the dorgu-operator-manager ClusterRole and ClusterRoleBinding, and the webhook and WebSocket Services and ValidatingWebhookConfiguration where you enabled them. The operator stops within seconds and nothing further is detected, diagnosed, or proposed.

2. Know what is still there

Quite a lot, and none of it announces itself: The volume is larger than it sounds. A clean-room evaluation, meaning a stranger installing from these published docs and nothing else, measured 176 custom resources across 12 namespaces, kube-system included, on a five-application test cluster after a single hour:
DorguEvent records reach kube-system because they are written into the namespace of the object they describe, and Dorgu watches the whole cluster. That is worth knowing before you install, not after.

3. Keep the history, if you want it

The incident and remediation records are the only account of what Dorgu saw. Step 4 destroys them, so export first if that matters:

4. Delete the CRDs

This is the step that completes the teardown. Deleting a CRD garbage-collects every custom resource of that kind, in every namespace, so five commands’ worth of cleanup is one command and no label selector is needed:
Irreversible, and wider than the namespace you installed into. This removes Dorgu’s records from the entire cluster, including any you wrote by hand and any under GitOps. Re-installing the chart brings the CRDs back empty.

5. Delete the namespace

Only if Dorgu was the only thing in it:
This also removes the API key Secret, but do not rely on it for that. Someone who runs the obvious helm uninstall and stops there keeps the key, which is why it leads this section.

6. Verify

Four empty results mean Dorgu is gone. Anything printed is a leftover, and the section above says which step removes it.

Removing the operator but keeping the record

A common middle case: you are done evaluating, but the incident history is worth keeping, or the personas are under GitOps and you intend to reinstall. Run step 1 only, and stop:
The CRDs and every record stay, nothing reconciles them, and a later helm install picks the same objects straight back up. The two things to be deliberate about are that DorguEvent records will not be pruned while no operator is running, and that the API key Secret is still there. If you only want the records gone and the CRDs kept, delete the records directly instead of the CRDs:
ClusterPersona is cluster-scoped, which is why it is a separate command.

Next steps

Quickstart

Apply your first persona and see the operator in action

Already have apps running?

Import personas from your live Deployments

Turn on AI

BYO Anthropic key for AI diagnosis and ordered plans

Security and permissions

The operator’s full ClusterRole and the verbs it does not have

Configuration

Customize flags and Helm values

Uninstall

Complete teardown, starting with your API key