Skip to main content
The Helm chart provides a values.yaml that maps to the operator’s command-line flags and Kubernetes resource configuration.

Default values

Value reference

Image

Replicas and scheduling

Webhook

The webhook requires TLS certificates. In production, use cert-manager to provision certificates automatically. The Helm chart includes the necessary annotations when cert-manager is available.

Self-healing

Rendered as --enable-health-check=<value> on the operator container either way, so the setting is always visible in the pod spec:
When healthCheck.enabled is false, the operator still validates personas and discovers the cluster; it just never opens incidents or proposes fixes. Setting metricsServer.enabled=false passes --enable-metrics-server=false; OOM and crash-loop detection read pod status and are unaffected, but usage and saturation signals go dark. See Self-healing.
Upgrading from a chart older than 0.8.0 turns detection on. A release that never set healthCheck.enabled will start detecting after the upgrade: expect IncidentMemory records to appear, and RemediationAction proposals wherever the ClusterPersona self-healing mode is propose (the default). Nothing is applied without approval. To keep the old behaviour, upgrade with --set healthCheck.enabled=false.

DorguEvent retention

DorguEvent records are Dorgu’s own log of what it saw, readable with kubectl get dorguevent -A. They live in etcd, so they are bounded twice over: by age, and by count. The count bound is the one that holds on a large cluster. Records arrive in proportion to cluster size, so a time window alone bounds nothing: a five-app clean-room cluster produced 187 records in 100 minutes, roughly 2,700 a day, every one of them inside the 24-hour window and all of them in etcd. The cleaner runs every 30 minutes, deletes oldest-first, and runs one pass at startup rather than waiting a full interval, because a restarted operator inherits the largest backlog it will ever see. Both flags are rendered on the operator container unconditionally, so the pod spec always states what bounds the record count, maxRecords: 0 included:
New in operator v0.10.0. Before it, the cleaner had no delete permission on dorguevents.dorgu.io at all: it ran every cycle, failed, and removed nothing. The ClusterRole now grants delete on Dorgu’s own CRD and on no workload kind. See why that verb is there.

AI

Both AI values default off, so a chart installed with no --set flags makes no API calls and spends nothing. Rules run either way. The key is always delivered to the pod as the ANTHROPIC_API_KEY environment variable via secretKeyRef, and the chart never passes it as a container argument. See AI setup for the full walkthrough and the caveat about llm.apiKey being visible in helm get values.
A Secret you created yourself with kubectl create secret, which is the recommended path, is not part of the Helm release and survives helm uninstall. Delete it explicitly: Uninstall leads with it.

Bootstrap

Set autoCreateClusterPersona: false if you manage ClusterPersonas through GitOps and do not want the operator creating one.

Validation

Metrics and probes

ArgoCD

The watcher auto-detects whether the ArgoCD CRD exists. If enabled: true but the CRD is absent, the watcher silently skips.

Prometheus

Example: http://prometheus-server.monitoring:9090

WebSocket

When enabled, the chart creates a Service on the specified port.

Resources

Service account

Nothing to set. Detection, rule-based diagnosis, and rule-based remediation are all on, and no key is involved:

Example: opting in to AI

Create the Secret out of band, then:

Example: Production deployment

Configuration overview

Feature matrix and CLI flags

AI setup

Key handling, verification, and turning AI off

Self-healing

What the health-check reconciler does once enabled

Installation

Install the operator with these values