Prerequisites
- Kubernetes cluster 1.24+
kubectlconfigured for your cluster- Helm 3.x (for Helm installation)
- Chart
0.11.1or newer onarm64nodes, including AWS Graviton and Apple Silicon kind or k3d clusters. Every image up to and including0.11.0wasamd64-only. See Versions.
Helm (recommended)
Install the operator from the OCI registry:Failure detection is on out of the box as of chart 0.8.0. The health-check reconciler runs every 60 seconds, opens
IncidentMemory records, and proposes RemediationAction fixes for what it finds. Nothing is ever applied without your approval.Earlier charts defaulted healthCheck.enabled to false, which meant following this page produced an operator that detected nothing. If you are reading an older copy of these instructions, that is the difference.AI
Versions
Current versions: CLI
v0.12.0 with operator v0.11.1 (Helm chart 0.11.1; the chart version, the chart appVersion, and the operator version are always the same number). Newest of each is always the right answer, and this note is the only place in these docs that states which that is.On arm64 nodes, 0.11.1 is a floor and not a preference. Every operator image published before it, up to and including 0.11.0, was pushed as a single-platform amd64 manifest rather than a manifest list. That covers AWS Graviton nodegroups and any kind or k3d cluster on an Apple Silicon Mac. On containerd there is no manifest list to select from, so the pull succeeds and the container dies immediately with exec /manager: exec format error and exit code 255. The pod reads CrashLoopBackOff, not ImagePullBackOff. 0.11.1 is that image and nothing else: no API, CRD, controller, or chart-template change, which is why everything below still names v0.11.0 as the release each behaviour landed in.There is no version pinning between the CLI and the operator. Every field either side has added is optional and additive, so a half that does not know about a field simply sees it absent and renders exactly as it did before. Nothing has to be upgraded in lockstep.Three mismatches change behaviour. Only the first is unsafe, and it is the old CLI, not the old operator:One operator upgrade is worth doing on its own merits, with no CLI implication at all: with
aiRemediation.enabled=true, operator v0.11.0 is the release that made AI-planned remediations appliable. On v0.10.0 and older, a plan that diagnosed a resource change could be persisted with nothing to apply. See AI setup.Both mismatch directions in full, with the ownership reasoning behind them: version coupling. Newest published releases: CLI and operator.Turn detection off
To run the operator for validation and cluster discovery only:Pin a chart version
The command above resolves the newest published chart. To pin an exact version for reproducibility, add--version:
The chart version and the operator version are the same number, so
--version 0.11.1 installs operator v0.11.1. If the literal above has fallen behind, Versions is the statement to trust on this page, and operator releases is the statement to trust over both.Install with a values file
To customize the installation, create avalues.yaml file:
Kustomize
Install CRDs and deploy the operator using kustomize:make uninstall deletes the CRDs, so it deletes every ApplicationPersona, IncidentMemory, RemediationAction and DorguEvent in the cluster along with them. Read Uninstall before you run it, particularly the part about your API key, which neither command touches.
Build from source
--enable-health-check to get the detection loop that a Helm install gives you.
Build and push Docker image
Verify installation
The chart prints a post-install summary that reports whether detection is on, whether AI is on, and what to run next. Re-read it at any time withhelm get notes dorgu-operator -n dorgu-system.
Check that the operator is running:
Confirm detection actually started:
Phase 2a/2b health check and remediation enabled with the interval. If it is absent, detection is off.
Uninstall
1. Remove the release
dorgu-operator-manager ClusterRole and ClusterRoleBinding, and the webhook and WebSocket Services and ValidatingWebhookConfiguration where you enabled them. The operator stops within seconds and nothing further is detected, diagnosed, or proposed.
2. Know what is still there
Quite a lot, and none of it announces itself:
The volume is larger than it sounds. A clean-room evaluation, meaning a stranger installing from these published docs and nothing else, measured 176 custom resources across 12 namespaces,
kube-system included, on a five-application test cluster after a single hour:
DorguEvent records reach kube-system because they are written into the namespace of the object they describe, and Dorgu watches the whole cluster. That is worth knowing before you install, not after.
3. Keep the history, if you want it
The incident and remediation records are the only account of what Dorgu saw. Step 4 destroys them, so export first if that matters:4. Delete the CRDs
This is the step that completes the teardown. Deleting a CRD garbage-collects every custom resource of that kind, in every namespace, so five commands’ worth of cleanup is one command and no label selector is needed:5. Delete the namespace
Only if Dorgu was the only thing in it:helm uninstall and stops there keeps the key, which is why it leads this section.
6. Verify
Removing the operator but keeping the record
A common middle case: you are done evaluating, but the incident history is worth keeping, or the personas are under GitOps and you intend to reinstall. Run step 1 only, and stop:helm install picks the same objects straight back up. The two things to be deliberate about are that DorguEvent records will not be pruned while no operator is running, and that the API key Secret is still there. If you only want the records gone and the CRDs kept, delete the records directly instead of the CRDs:
ClusterPersona is cluster-scoped, which is why it is a separate command.
Next steps
Quickstart
Apply your first persona and see the operator in action
Already have apps running?
Import personas from your live Deployments
Turn on AI
BYO Anthropic key for AI diagnosis and ordered plans
Security and permissions
The operator’s full ClusterRole and the verbs it does not have
Configuration
Customize flags and Helm values
Uninstall
Complete teardown, starting with your API key