The loop
Understanding first: who owns the workload
Detection and diagnosis run on every workload with a persona. Healing does not. As of operator v0.9.0 the operator reads the live Deployment at proposal time and records who owns it onRemediationAction.spec.workloadRef.managedBy, one of helm, argocd, flux, kustomize, unmanaged, or unknown. For anything except unmanaged, Dorgu is understanding and recommendation only: it detects, diagnoses, and tells you exactly what to change in your source of truth, and it will not patch the Deployment.
The reason is concrete. Under server-side apply, patching a Helm-owned Deployment claims those fields away from Helm’s field manager, and the next helm upgrade then fails outright with a field-manager conflict rather than quietly reverting the fix. A fix that breaks your next deploy is not a fix.
The same logic runs the other way on the workloads Dorgu does patch. A heal leaves no ownership footprint: the CLI patches under its own dorgu field manager and then strips that entry, so the fields end up owned by nobody and a later server-side apply succeeds. Because an Update takes fields away from whoever held them, a heal even clears a pre-existing kubectl-set claim rather than adding to it. New in CLI v0.11.0: Dorgu leaves no field manager behind.
kubectl logs and kubectl get events matter most on exactly the workloads Dorgu will not touch.
kubectl apply -k is indistinguishable from one applied by hand. The published guarantee therefore names Helm, ArgoCD and Flux, the three that run controllers and stamp what they reconcile. Full detail, and what the CLI prints before it writes: the kustomize limitation.persona-update step patches the ApplicationPersona, the operator does that itself, and it is always safe whoever owns the workload. autoExecutable semantics are unchanged. Ownership governs one thing only: the CLI patching your Deployment.4, and what the refusal looks like: the ownership model.
Detect
The health-check reconciler runs on a fixed interval — 60 seconds by default, and 30 seconds is a reasonable setting for a tight loop. It collects signals from every detector, correlates them to anApplicationPersona, and opens or updates an IncidentMemory.
OOMKilled is read straight from the container’s current or last termination state in the pod status, so the most valuable signal works on a bare cluster. Metrics-server only adds the container-level usage and saturation signals.One application at a time
Before anything is diagnosed, signals are partitioned by application and namespace. A rule handed only one application’s signals can only ever describe one application. That ordering is load-bearing. Diagnosis used to run once across every signal in the cluster, so a rule that takes “all the OOMKilled signals” as a single finding named the first persona it saw as the owner and listed every OOM-killed pod as affected. On a clean-room cluster with four unrelated apps failing at once it produced oneIncidentMemory holding pods from three namespaces, and the planner, handed a bundle spanning three namespaces, concluded the nodes were under memory pressure. The nodes were at 23%.
Attribution is correspondingly strict:
- Namespace-scoped. Only personas in the signal’s own namespace are considered. An app can never claim another namespace’s pod.
- Exactly one owner, most specific claim wins. Personas
apiandapi-serverboth prefix-match podapi-server-7f9d-x2q; the longer name is the real claim and the shorter one is a coincidence of prefixes. A genuine tie at the same specificity is left unattributed rather than handed to whichever persona the API server listed first.
spec.attribution: unattributed and the label dorgu.io/attribution=unattributed. It is a real outage Dorgu can see but cannot diagnose against a persona, so no remediation is proposed for it and the log says why. Once you import a persona and an attributed incident starts tracking the same workload, the unattributed one closes as superseded, so onboarding an app mid-outage does not count it twice.An honest “something here is broken and I do not know whose it is” beats a confident attribution to the wrong app. List them with:Resolution needs evidence, not silence
An incident closes only when recovery can be observed. Two conditions have to hold together:- The signal has been absent for the 5-minute grace period, and
- the workload’s pods are present,
Ready, and have stayedReadywith no container restarts and no container in a waiting state for a 6-minute stability window.
spec.resolution.action as auto-resolved: <what was observed>, so the reason lives on the object rather than only in a log line that has since rotated away.
Diagnose
Diagnosis turns raw signals into a root cause. The deterministic rule-based provider is always on and always the floor. It produces a summary, a category, a suggested action, and a confidence score computed from the rule’s base confidence, how many correlated signals were seen, how unambiguous those signals are (anOOMKilled is clearer than a PodPending), and how tightly they cluster in time.
With an Anthropic key configured, an AI provider runs afterwards and enhances those results with cluster context. Its diagnoses are recorded with provider: ai-enhanced versus rule-based. When two diagnoses describe the same category, action, and resources, the higher-confidence one wins, and a tie goes to the later provider, the one enhancing the earlier result. That matters because an AI-enhanced diagnosis re-runs the rule-based logic and often carries the identical confidence to the digit, so a strict comparison used to drop every AI result silently. The losing diagnosis is now logged at INFO with the reason.
Any LLM failure degrades to the rules. It never blocks the loop.
Propose
Every proposal starts from the live workload, not the persona. TheApplicationPersona is a point-in-time record, and on a brownfield cluster it drifts from the running Deployment. Dorgu used to compute sizes, state numbers, and diff as though the persona were authoritative, which is how it once quoted a 96Mi memory limit for a container that had been running with 32Mi for weeks. Since operator v0.9.0:
- The operator reads the live Deployment at proposal time and records what it found on
spec.workloadRef: the workload’s real name, the container, the observed image, and the observedrequestsandlimits. - Every stated fact comes from the live container, and names the Deployment and container it was read from. Where the workload is unreadable, Dorgu says so and warns that the persona figure may have drifted.
- The blast-radius cap is measured against the live value, not against whichever of persona and live is larger. Capping against the larger one permitted a 192Mi proposal on a 32Mi container whenever the persona was the stale, bigger number.
- A remediation may only change a resource key the workload already sets.
observedResourcesdistinguishes absent from zero, so approving a memory fix can no longer silently introduce a CPU limit the Deployment never had. The rule-based path refuses and says why; the AI path drops the offending patch leaf, records the omission in the step rationale, and demotes a step whose patch is emptied by the pruning. - The planner may only name versions Dorgu has actually read, from the
dorgu.io/imported-imageannotation andstatus.deployments. It can no longer invent a prior-good image tag.
The default path is deterministic
By default, and on the current recommendation, the rule-based proposer writes the plan and emitsplanSource: rule-based. It handles the resource-adjustment path, which is the OOM and saturation cases, computing a bounded increase against the live container value. It needs no key, makes no API call, and is what the clean-room runs have actually healed with.
The AI planner is opt-in on top of that, behind llm.provider and aiRemediation.enabled. Why AI is off by default is the honest version of that trade, including the three runs where the rules came out ahead.
workload-apply steps, which the CRD forbids from ever being auto-executable, and give its one persona-update step no patch. The resulting object read like a fix and applied nothing.Now, after the guardrails run, such a plan either carries a patch that can be applied or Dorgu supplies the value itself, using the same conservative arithmetic the rule-based path uses, measured against the live workload. Where Dorgu cannot make it appliable the plan is refused outright and the rule-based proposal takes over, recorded as planSource: rule-based so it is auditable as what it is.The plan keeps the model’s root cause; the number that reaches your cluster is the rule engine’s. Every substitution is recorded on spec.steps[].safety with the verdict derived, so it never has to be inferred from prose.One exemption is deliberate. Where Dorgu’s own rule engine declines to size a change, the advisory plan stands: a container that sets no CPU limit today means raising one would introduce a field the workload has never had, so the rule-based path would produce nothing either. A plan that says exactly that is worth more than no plan.With the AI planner enabled
The planner assembles the context it sends to Claude:- A
## Live workload (ground truth, read from the cluster just now)section, ahead of the persona: the observed container, its resources, and its image - The affected ApplicationPersona, spec and status, including the learned resource baselines, explicitly labelled a stale snapshot
- The singleton ClusterPersona — platform, capacity, and the self-healing policy
- Recent IncidentMemory records for the same application
- The RemediationActions those incidents referenced, with their phase and verification result — so the model can see which past fixes actually worked and which were rolled back
autoExecutable flag, and optionally a ready-to-run command; persona-update steps carry the JSON merge patch plus a prePatchState snapshot for rollback. The result lands on RemediationAction.spec.steps[] with planSource: ai-anthropic.
For an owned workload the plan is then reshaped for its owner: see what changes about the plan.
persona-update steps may be auto-executable. This is enforced at the Kubernetes API server by a CEL validation rule on the CRD, not merely in operator code. workload-apply, restart, scale, config-change, and manual steps are advisory — recorded for a human, the CLI, or your pipeline to act on. No plan can escalate itself.dorgu remediation diff.
Guardrails
Every proposal passes the safety checker before it is created.Guardrail verdicts as data
New in operator v0.11.0: when a guardrail rules on a field, what it decided is recorded as structured data onspec.steps[].safety rather than written into a sentence.
Every value in that field is Dorgu’s own arithmetic. None of it comes from a model, and the model is not permitted to comment on a cap at all. That is the entire point of the field, and it exists because of a specific failure. The verdict used to be spliced onto the front of the step’s rationale as [safety:blast-radius] …, which put Dorgu’s measurement one line below the model’s own claim that the same 16x change was “well within a 2x ceiling”, with nothing on the screen to tell a reader which of the two had been computed. A person deciding whether to approve a change to a running workload should not have to work that out from the phrasing.
Each entry names one field and carries the rule, the verdict, the baseline it measured against, what the plan asked for, what will actually be applied, the ratio, the ceiling, and a message Dorgu wrote:
- A refused field is gone from the patch, so no diff and no owner instruction can advertise a change that will not happen. The plan used to keep the refused value while the step was demoted to advisory, and the diff went on offering
8Mi -> 128Mito a reader who would have got nothing. - The owner instruction is built after the guardrails, not before. It quotes concrete values out of the patches, so building it first told a Helm user to put the refused value into their values file.
- Every field that breaches the cap is reported, in a stable order. The check used to return the first breach it happened across while ranging over a map, so a plan moving two fields too far named one at random and silently ignored the other.
dorgu remediation diff, in the approve and heal preview, and in the owned-workload refusal. See what the CLI shows and the field reference.
safety is absent on every object written by an operator older than v0.11.0, and absent means no guardrail ruled. A CLI that does not know the field renders exactly as it did before, and no migration is needed. See guardrail verdicts.Rejected action used to count as terminal and therefore non-blocking, so the next health-check cycle proposed the same fix again, including another billable planning call. The remediation controller now stamps a Rejected condition once, on first sight of the phase, and the health-check reconciler consults that history before calling the proposer.Suppression lifts after an hour, or sooner when the signal materially changes: the live diagnosis outranks the severity the declined incident was opened at, so a warning someone waved off that has since gone critical is treated as a different question. An unreadable rejection history fails closed, and a rejection with no timestamp still suppresses, because an un-timestamped no is still a no.IncidentMemory used to drop the whole diagnosis with no retry and no record: one clean-room run lost 176 of them across 4 hours 20 minutes, and the dropped ones were the better ones. The spec write now retries with a re-fetch. If it still fails, the loss is logged at ERROR, recorded as a DorguEvent, and emitted as a Kubernetes Warning under the reason DorguDiagnosisDiscarded, so it shows up in kubectl get events and in log-based alerting. Every cycle ends with a tally.Verify and remember
Approval starts a timed verification, not an instant success.Apply
Applying, recording appliedAt.Wait
spec.rollback.healthCheckAfter — 10 minutes by default — then moves to Verifying.Verify
Settle
Completed, and the incident is marked Resolved with the outcome and duration. Degraded → the prePatchState is restored and the action becomes RolledBack. Unknown → retried twice, one minute apart, then Failed.IncidentMemory keeps the signal, the root cause, the confidence, the occurrence count, and the resolution outcome; RemediationAction keeps the plan and the verification result. Both become context for the next proposal.
Trust model
TheClusterPersona carries the self-healing policy:
mode gates the loop:
dorgu-cluster persona with mode: propose. Set mode: observe to watch the loop diagnose without proposing anything.
See the trust model for how the levels are meant to progress.
Prerequisites
How a persona finds its Deployment
Two separate lookups run in the loop, and confusing them is the usual reason nothing happens. Persona to Deployment. The operator lists every Deployment in the persona’s namespace and walks an ordered chain, taking the first rung that matches exactly one Deployment:spec.name.
Pending forever. The dorgu CLI walks the identical chain when it picks the Deployment to patch, so the CLI and the operator never disagree about which workload a persona describes.AmbiguousDeployment and names the candidates rather than patching one at random. If no rung matches anything, it reports NoDeployment and lists every rung it tried.
Signal to persona. Detection goes the other way and matches by name prefix: a signal about resource X is attributed to a persona when X equals the persona’s metadata.name or spec.name, or starts with either followed by a hyphen. That is what handles pod suffixes, so a persona named memhog picks up memhog-7d9f-x2k.
Two rules bound it. Only personas in the signal’s own namespace are considered, and the resource must be claimed by exactly one persona: where several match, the longest matching name wins, and a tie at the same length leaves the signal unattributed rather than filed against a guess. So a persona named api does pick up api-server-7f9d-x2q when it is the only persona in the namespace, and loses it to api-server when that persona also exists.
Troubleshooting
Nothing happens at all: no incidents, no remediations, on a cluster full of running apps
Nothing happens at all: no incidents, no remediations, on a cluster full of running apps
ApplicationPersona. Dorgu only watches workloads that have one, so a broken app with no persona raises no incident, which means no diagnosis and no proposed fix. Nothing is misconfigured; there is simply nothing registered to watch.Ask Dorgu what it cannot see:Unmonitored section names every Deployment with no matching persona and prints an import command per namespace. Then create them from the live Deployments:dorgu persona import and the brownfield walkthrough.A persona sits in Pending with reason NoDeployment
A persona sits in Pending with reason NoDeployment
- The persona is in the wrong namespace. Resolution never leaves the persona’s own namespace. The Deployment being one namespace over looks identical to it not existing.
spec.namedoes not match anything. The chain matchesspec.name, notmetadata.name. Setspec.nameto the Deployment’s name, or addapp.kubernetes.io/name: <spec.name>to the Deployment object.- The Deployment genuinely is not there yet. The persona re-resolves on the next reconcile, so this clears itself once the workload lands.
dorgu persona import picks a spec.name that resolves back to the Deployment it read, so imported personas do not hit this.A persona sits in Pending with reason AmbiguousDeployment
A persona sits in Pending with reason AmbiguousDeployment
app.kubernetes.io/name is rung 1, so setting it on the one Deployment you mean outranks the collision on a later rung. Alternatively, split them into one persona per Deployment with distinct spec.name values.No incident is created, even though the pod is crash-looping
No incident is created, even though the pod is crash-looping
metadata.name or spec.name, or starts with either followed by a hyphen, which is how pod suffixes are handled. A persona named checkout never picks up pods named payments-..., and it never picks up pods in another namespace, because correlation never leaves the signal’s own namespace.Note that this is a different lookup from persona-to-Deployment resolution, which needs no label and no prefix. A persona can be Active, having resolved its Deployment cleanly, and still collect no incidents because the pod names do not prefix-match. See how a persona finds its Deployment.If the pods do prefix-match more than one persona at the same length, the signal is deliberately left unattributed rather than filed against whichever one happened to be listed first. The incident still exists; it is filed against the workload:The operator is installed but detects nothing
The operator is installed but detects nothing
--enable-health-check=true. If it reads false, either the values set it or the release predates 0.8.0:incidentmemories.dorgu.io. See installation.An incident exists, but no remediation is proposed
An incident exists, but no remediation is proposed
persona has no resource limits configured. Add spec.resources.limits to the persona.Other logged skip reasons worth checking: safety check failed: [rate-limit ...] (5 per persona per hour), [concurrent ...] (one active remediation at a time), [deny-list ...] (the namespace is excluded), [blast-radius ...] (the change exceeded 2×), and the dedup skip when an active remediation already covers the same target.Two incidents for one problem
Two incidents for one problem
OOMKilled and CrashLoopBackOff. An incident is keyed by persona, category, and primary signal, so one application can hold several at once. Dedup runs at the remediation layer, not the incident layer, so you see two incidents and one fix. Approving that one fix resolves both.What will not happen any more is the reverse: two applications sharing one incident. Signals are grouped per application per namespace before anything is diagnosed. See one application at a time.An incident stays open although the pod looks healthy
An incident stays open although the pod looks healthy
Ready, restart-free, and out of any waiting state for a further 6 minutes. A pod that came back 90 seconds ago has not cleared that bar yet.The operator logs the specific reason each cycle at V(1), naming the pod that is holding it open:Failed (an evicted pod can outlive the failure that produced it, so delete it), or a Deployment that wants replicas and has none running. Anything Dorgu could not read leaves the incident open on purpose: under-reporting an outage is the one failure a reliability tool cannot afford. See resolution needs evidence.Approve prints "Dorgu will not patch this workload" and exits 4
Approve prints "Dorgu will not patch this workload" and exits 4
ExitDeclined) exists to say “declined by design” rather than “broke”. Nothing was approved and nothing in the cluster changed.The refusal names the owner and prints the owner-shaped steps. Apply the change there, then record the decision:unknown on a workload you know nothing reconciles, ownership detection found a server-side-apply field manager it does not recognise, and managedByDetail names it. unknown is treated as owned on purpose: nothing is patched on a guess. Full detail in the ownership model.Every heal is declined, even on workloads nothing manages
Every heal is declined, even on workloads nothing manages
spec.workloadRef is written by operator v0.9.0 and newer, and CLI v0.10.0 treats an absent record as owned, because it means either an operator too old to know or a workload that could not be read. Neither is evidence that patching is safe.--no-heal records decisions and you apply changes by hand. See version coupling.The numbers Dorgu quotes do not match my pod
The numbers Dorgu quotes do not match my pod
observedResources means the container does not set that key, which is different from a value of zero, and is what stops a memory fix from adding a CPU limit. A workloadRef with no name means the operator could not resolve the Deployment at all: fix that first, since an unresolved workload loses its grounding and its advisory commands together.heal printed a green checkmark and then exited 1, but the pod is fixed
heal printed a green checkmark and then exited 1, but the pod is fixed
RemediationAction, and exiting 0 there would reproduce the defect the recording exists to fix: a cluster and a record that disagree.Re-run the heal. It is idempotent and records the change on the way through:kubectl patch is offered: when the record cannot be written.The remediation is stuck in Applying or Verifying
The remediation is stuck in Applying or Verifying
1/1 Running with no new restarts, the heal worked and the phase will settle on its own.Plans come back rule-based when AI is enabled
Plans come back rule-based when AI is enabled
AI remediation planning failed, falling back to rules with the reason. Check that the operator can reach the Anthropic API (egress and NAT), that the key is valid, and that the startup log includes AI remediation planning enabled. See AI setup.A plan Dorgu refused to persist. Since operator v0.11.0, a plan that diagnoses a resource change but carries nothing appliable is rejected in favour of the rule-based proposal rather than written as a fix that applies nothing. planSource: rule-based is then the correct and auditable record of what actually happened. This is working as intended, and it is strictly better than the alternative it replaced. See the default path.Approve prints "No resource change to apply" on a plan that reads like a fix
Approve prints "No resource change to apply" on a plan that reads like a fix
workload-apply steps, which the CRD’s CEL rule forbids from ever being auto-executable, and gave its one persona-update step no patch at all. spec.action.type came out notification, approve had nothing to do, and the pod was still in CrashLoopBackOff 42 minutes later. Nine of nine AI-planned remediations were unappliable.A step's diff is smaller than the plan's own summary implied
A step's diff is smaller than the plan's own summary implied