Skip to main content
Vanilla Kubernetes restarts a crash-looping pod forever. It never asks why. Dorgu’s self-healing loop closes that gap: it detects the failure, diagnoses the root cause, proposes a reviewable fix, and — once you approve — heals the workload and remembers the outcome. Code detects. AI explains. A human approves.

The loop

The invariant. The operator never creates or modifies Deployments, Services, or any other workload resource. It reads cluster state, updates the Persona / Incident / Remediation CRDs, and recommends. This is enforced by RBAC, not by convention: its ClusterRole grants only get, list, and watch on Deployments, and no access to Secrets at all. The role is published in full on security and permissions.Applying workload changes stays with ArgoCD, Flux, kubectl, or the Dorgu CLI, which uses your credentials, not the operator’s.

Understanding first: who owns the workload

Detection and diagnosis run on every workload with a persona. Healing does not. As of operator v0.9.0 the operator reads the live Deployment at proposal time and records who owns it on RemediationAction.spec.workloadRef.managedBy, one of helm, argocd, flux, kustomize, unmanaged, or unknown. For anything except unmanaged, Dorgu is understanding and recommendation only: it detects, diagnoses, and tells you exactly what to change in your source of truth, and it will not patch the Deployment. The reason is concrete. Under server-side apply, patching a Helm-owned Deployment claims those fields away from Helm’s field manager, and the next helm upgrade then fails outright with a field-manager conflict rather than quietly reverting the fix. A fix that breaks your next deploy is not a fix. The same logic runs the other way on the workloads Dorgu does patch. A heal leaves no ownership footprint: the CLI patches under its own dorgu field manager and then strips that entry, so the fields end up owned by nobody and a later server-side apply succeeds. Because an Update takes fields away from whoever held them, a heal even clears a pre-existing kubectl-set claim rather than adding to it. New in CLI v0.11.0: Dorgu leaves no field manager behind. For an owned workload the operator rewrites every step whose command would write to the cluster: the command is dropped, and the description becomes the chart values to set for a Helm release, the Git path for an ArgoCD application, or the overlay for kustomize. Read-only commands survive, because kubectl logs and kubectl get events matter most on exactly the workloads Dorgu will not touch.
kustomize is named in the enum but not in the claim. It is a client-side renderer with no controller, and it stamps nothing on its output unless the kustomization opts in, so a Deployment from kubectl apply -k is indistinguishable from one applied by hand. The published guarantee therefore names Helm, ArgoCD and Flux, the three that run controllers and stamp what they reconcile. Full detail, and what the CLI prints before it writes: the kustomize limitation.
Persona writes are not gated on ownership. A persona-update step patches the ApplicationPersona, the operator does that itself, and it is always safe whoever owns the workload. autoExecutable semantics are unchanged. Ownership governs one thing only: the CLI patching your Deployment.
The guard is enforced by the CLI, so upgrade both. Operator v0.9.0 supplies the facts and strips the workload-writing commands; declining the patch is dorgu remediation approve / heal’s job, and that landed in CLI v0.10.0. An older CLI against a v0.9.0 operator will still patch an owned Deployment. See version coupling for both mismatch directions.
Full detail, including how ownership is detected, exit code 4, and what the refusal looks like: the ownership model.

Detect

The health-check reconciler runs on a fixed interval — 60 seconds by default, and 30 seconds is a reasonable setting for a tight loop. It collects signals from every detector, correlates them to an ApplicationPersona, and opens or updates an IncidentMemory.
OOM detection needs no metrics-server. OOMKilled is read straight from the container’s current or last termination state in the pod status, so the most valuable signal works on a bare cluster. Metrics-server only adds the container-level usage and saturation signals.

One application at a time

Before anything is diagnosed, signals are partitioned by application and namespace. A rule handed only one application’s signals can only ever describe one application. That ordering is load-bearing. Diagnosis used to run once across every signal in the cluster, so a rule that takes “all the OOMKilled signals” as a single finding named the first persona it saw as the owner and listed every OOM-killed pod as affected. On a clean-room cluster with four unrelated apps failing at once it produced one IncidentMemory holding pods from three namespaces, and the planner, handed a bundle spanning three namespaces, concluded the nodes were under memory pressure. The nodes were at 23%. Attribution is correspondingly strict:
  • Namespace-scoped. Only personas in the signal’s own namespace are considered. An app can never claim another namespace’s pod.
  • Exactly one owner, most specific claim wins. Personas api and api-server both prefix-match pod api-server-7f9d-x2q; the longer name is the real claim and the shorter one is a coincidence of prefixes. A genuine tie at the same specificity is left unattributed rather than handed to whichever persona the API server listed first.
Unattributed incidents. A workload no single persona claims still gets an incident, filed against the workload with spec.attribution: unattributed and the label dorgu.io/attribution=unattributed. It is a real outage Dorgu can see but cannot diagnose against a persona, so no remediation is proposed for it and the log says why. Once you import a persona and an attributed incident starts tracking the same workload, the unattributed one closes as superseded, so onboarding an app mid-outage does not count it twice.An honest “something here is broken and I do not know whose it is” beats a confident attribution to the wrong app. List them with:
Cluster-scoped findings (nodes, control plane) belong to no application. They are diagnosed and logged, and deliberately not filed against a persona, because giving them an owner they do not have is the mistake above.

Resolution needs evidence, not silence

An incident closes only when recovery can be observed. Two conditions have to hold together:
  1. The signal has been absent for the 5-minute grace period, and
  2. the workload’s pods are present, Ready, and have stayed Ready with no container restarts and no container in a waiting state for a 6-minute stability window.
If the cluster cannot be read, if the pods cannot be found, or if they are there but not yet stable, the incident stays open and the reason is logged. When one does close, the evidence is written to spec.resolution.action as auto-resolved: <what was observed>, so the reason lives on the object rather than only in a log line that has since rotated away.
Absence of a signal is not recovery. The old rule needed only two absences: no matching signal this cycle, plus the grace period. A crash loop backs off in lengthening intervals up to five minutes, so a completely dead pod falls silent inside the grace period and used to read as fixed. In clean-room run #3 an incident reached 51 occurrences, went Resolved, and stayed in CrashLoopBackOff, while dorgu health reported one active incident with three applications down. The stability window is longer than the longest backoff interval for exactly that reason.The one case where no pods still resolves: the workload has no pods and no Deployment (you deleted it), or its Deployment is scaled to zero. A Deployment that wants replicas and has none running is down, not recovered.
New in operator v0.10.0. Grouping, strict attribution, unattributed incidents, and evidence-based resolution ship in v0.10.0. None of them are in v0.9.0.

Diagnose

Diagnosis turns raw signals into a root cause. The deterministic rule-based provider is always on and always the floor. It produces a summary, a category, a suggested action, and a confidence score computed from the rule’s base confidence, how many correlated signals were seen, how unambiguous those signals are (an OOMKilled is clearer than a PodPending), and how tightly they cluster in time. With an Anthropic key configured, an AI provider runs afterwards and enhances those results with cluster context. Its diagnoses are recorded with provider: ai-enhanced versus rule-based. When two diagnoses describe the same category, action, and resources, the higher-confidence one wins, and a tie goes to the later provider, the one enhancing the earlier result. That matters because an AI-enhanced diagnosis re-runs the rule-based logic and often carries the identical confidence to the digit, so a strict comparison used to drop every AI result silently. The losing diagnosis is now logged at INFO with the reason. Any LLM failure degrades to the rules. It never blocks the loop.

Propose

Every proposal starts from the live workload, not the persona. The ApplicationPersona is a point-in-time record, and on a brownfield cluster it drifts from the running Deployment. Dorgu used to compute sizes, state numbers, and diff as though the persona were authoritative, which is how it once quoted a 96Mi memory limit for a container that had been running with 32Mi for weeks. Since operator v0.9.0:
  • The operator reads the live Deployment at proposal time and records what it found on spec.workloadRef: the workload’s real name, the container, the observed image, and the observed requests and limits.
  • Every stated fact comes from the live container, and names the Deployment and container it was read from. Where the workload is unreadable, Dorgu says so and warns that the persona figure may have drifted.
  • The blast-radius cap is measured against the live value, not against whichever of persona and live is larger. Capping against the larger one permitted a 192Mi proposal on a 32Mi container whenever the persona was the stale, bigger number.
  • A remediation may only change a resource key the workload already sets. observedResources distinguishes absent from zero, so approving a memory fix can no longer silently introduce a CPU limit the Deployment never had. The rule-based path refuses and says why; the AI path drops the offending patch leaf, records the omission in the step rationale, and demotes a step whose patch is emptied by the pruning.
  • The planner may only name versions Dorgu has actually read, from the dorgu.io/imported-image annotation and status.deployments. It can no longer invent a prior-good image tag.

The default path is deterministic

By default, and on the current recommendation, the rule-based proposer writes the plan and emits planSource: rule-based. It handles the resource-adjustment path, which is the OOM and saturation cases, computing a bounded increase against the live container value. It needs no key, makes no API call, and is what the clean-room runs have actually healed with. The AI planner is opt-in on top of that, behind llm.provider and aiRemediation.enabled. Why AI is off by default is the honest version of that trade, including the three runs where the rules came out ahead.
Since operator v0.11.0, an AI plan that diagnoses a resource change is appliable or it is not written. Before it, the planner could describe a memory fix as workload-apply steps, which the CRD forbids from ever being auto-executable, and give its one persona-update step no patch. The resulting object read like a fix and applied nothing.Now, after the guardrails run, such a plan either carries a patch that can be applied or Dorgu supplies the value itself, using the same conservative arithmetic the rule-based path uses, measured against the live workload. Where Dorgu cannot make it appliable the plan is refused outright and the rule-based proposal takes over, recorded as planSource: rule-based so it is auditable as what it is.The plan keeps the model’s root cause; the number that reaches your cluster is the rule engine’s. Every substitution is recorded on spec.steps[].safety with the verdict derived, so it never has to be inferred from prose.One exemption is deliberate. Where Dorgu’s own rule engine declines to size a change, the advisory plan stands: a container that sets no CPU limit today means raising one would introduce a field the workload has never had, so the rule-based path would produce nothing either. A plan that says exactly that is worth more than no plan.

With the AI planner enabled

The planner assembles the context it sends to Claude:
  • A ## Live workload (ground truth, read from the cluster just now) section, ahead of the persona: the observed container, its resources, and its image
  • The affected ApplicationPersona, spec and status, including the learned resource baselines, explicitly labelled a stale snapshot
  • The singleton ClusterPersona — platform, capacity, and the self-healing policy
  • Recent IncidentMemory records for the same application
  • The RemediationActions those incidents referenced, with their phase and verification result — so the model can see which past fixes actually worked and which were rolled back
Claude returns an ordered plan. Each step carries an order, type, description, rationale, risk level, an autoExecutable flag, and optionally a ready-to-run command; persona-update steps carry the JSON merge patch plus a prePatchState snapshot for rollback. The result lands on RemediationAction.spec.steps[] with planSource: ai-anthropic. For an owned workload the plan is then reshaped for its owner: see what changes about the plan.
Only persona-update steps may be auto-executable. This is enforced at the Kubernetes API server by a CEL validation rule on the CRD, not merely in operator code. workload-apply, restart, scale, config-change, and manual steps are advisory — recorded for a human, the CLI, or your pipeline to act on. No plan can escalate itself.
Read the resulting plan with dorgu remediation diff.

Guardrails

Every proposal passes the safety checker before it is created.

Guardrail verdicts as data

New in operator v0.11.0: when a guardrail rules on a field, what it decided is recorded as structured data on spec.steps[].safety rather than written into a sentence. Every value in that field is Dorgu’s own arithmetic. None of it comes from a model, and the model is not permitted to comment on a cap at all. That is the entire point of the field, and it exists because of a specific failure. The verdict used to be spliced onto the front of the step’s rationale as [safety:blast-radius] …, which put Dorgu’s measurement one line below the model’s own claim that the same 16x change was “well within a 2x ceiling”, with nothing on the screen to tell a reader which of the two had been computed. A person deciding whether to approve a change to a running workload should not have to work that out from the phrasing. Each entry names one field and carries the rule, the verdict, the baseline it measured against, what the plan asked for, what will actually be applied, the ratio, the ceiling, and a message Dorgu wrote:
Three consequences follow, and each of them was a defect before:
  • A refused field is gone from the patch, so no diff and no owner instruction can advertise a change that will not happen. The plan used to keep the refused value while the step was demoted to advisory, and the diff went on offering 8Mi -> 128Mi to a reader who would have got nothing.
  • The owner instruction is built after the guardrails, not before. It quotes concrete values out of the patches, so building it first told a Helm user to put the refused value into their values file.
  • Every field that breaches the cap is reported, in a stable order. The check used to return the first breach it happened across while ranging over a map, so a plan moving two fields too far named one at random and silently ignored the other.
The CLI prints all of this under its own heading, away from any sentence a model may have authored, in dorgu remediation diff, in the approve and heal preview, and in the owned-workload refusal. See what the CLI shows and the field reference.
The 2x cap anchors to the live value, so an app that has drifted far below its recorded intent cannot be restored in one step. After a drift down to 8Mi the largest permitted step is 16Mi, which may still be too little. Dorgu says so in the verdict and in the plan summary rather than implying the fix is complete. Moving the anchor to the persona’s recorded intent would weaken the guardrail in exactly the case where the persona is the thing that is stale, so the anchor is deliberately unchanged. Approve the capped step, let it verify, and the next proposal measures from the new live value.
Optional and additive. safety is absent on every object written by an operator older than v0.11.0, and absent means no guardrail ruled. A CLI that does not know the field renders exactly as it did before, and no migration is needed. See guardrail verdicts.
Rejecting no longer costs you money 30 seconds later. A Rejected action used to count as terminal and therefore non-blocking, so the next health-check cycle proposed the same fix again, including another billable planning call. The remediation controller now stamps a Rejected condition once, on first sight of the phase, and the health-check reconciler consults that history before calling the proposer.Suppression lifts after an hour, or sooner when the signal materially changes: the live diagnosis outranks the severity the declined incident was opened at, so a warning someone waved off that has since gone critical is treated as a different question. An unreadable rejection history fails closed, and a rejection with no timestamp still suppresses, because an un-timestamped no is still a no.
A discarded AI diagnosis is now visible. Concurrent writes to an IncidentMemory used to drop the whole diagnosis with no retry and no record: one clean-room run lost 176 of them across 4 hours 20 minutes, and the dropped ones were the better ones. The spec write now retries with a re-fetch. If it still fails, the loss is logged at ERROR, recorded as a DorguEvent, and emitted as a Kubernetes Warning under the reason DorguDiagnosisDiscarded, so it shows up in kubectl get events and in log-based alerting. Every cycle ends with a tally.

Verify and remember

Approval starts a timed verification, not an instant success.
1

Apply

The operator patches the ApplicationPersona spec and moves to Applying, recording appliedAt.
2

Wait

It waits out spec.rollback.healthCheckAfter — 10 minutes by default — then moves to Verifying.
3

Verify

It re-runs the whole detection engine and asks two questions: is the original signal still present for this persona, and are there new critical signals?
4

Settle

Clean → Completed, and the incident is marked Resolved with the outcome and duration. Degraded → the prePatchState is restored and the action becomes RolledBack. Unknown → retried twice, one minute apart, then Failed.
Your pod recovers within seconds of approve, but the remediation stays in Applying / Verifying for the full window and the incident stays open until then. That is the design, not a hang — the wait is what makes automatic rollback meaningful.
Either way the record persists. IncidentMemory keeps the signal, the root cause, the confidence, the occurrence count, and the resolution outcome; RemediationAction keeps the plan and the verification result. Both become context for the next proposal.

Trust model

The ClusterPersona carries the self-healing policy:
mode gates the loop: The operator auto-creates its default dorgu-cluster persona with mode: propose. Set mode: observe to watch the loop diagnose without proposing anything.
trustLevel, enabled, and autoApproveRule are descriptive, not enforced. trustLevel is only fed to the AI planner as context; detection, diagnosis, and proposal run regardless of enabled (use mode: observe to stop at diagnosis); and spec.approval.autoApproveRule exists in the RemediationAction CRD but no controller reads it. Every remediation requires human approval. maxRemediationsPerHour and excludeNamespaces are enforced.
See the trust model for how the levels are meant to progress.

Prerequisites

How a persona finds its Deployment

Two separate lookups run in the loop, and confusing them is the usual reason nothing happens. Persona to Deployment. The operator lists every Deployment in the persona’s namespace and walks an ordered chain, taking the first rung that matches exactly one Deployment: Earlier rungs are more explicit statements of intent, so they win. The value matched against every rung is the persona’s spec.name.
No label is required. Helm, kustomize, and most hand-written YAML label the pod template only and leave the Deployment object bare; rung 3 or rung 4 resolves those. Requiring a label on the Deployment object is what used to leave personas for pre-existing apps Pending forever. The dorgu CLI walks the identical chain when it picks the Deployment to patch, so the CLI and the operator never disagree about which workload a persona describes.
If a rung matches several Deployments, the persona reports AmbiguousDeployment and names the candidates rather than patching one at random. If no rung matches anything, it reports NoDeployment and lists every rung it tried. Signal to persona. Detection goes the other way and matches by name prefix: a signal about resource X is attributed to a persona when X equals the persona’s metadata.name or spec.name, or starts with either followed by a hyphen. That is what handles pod suffixes, so a persona named memhog picks up memhog-7d9f-x2k. Two rules bound it. Only personas in the signal’s own namespace are considered, and the resource must be claimed by exactly one persona: where several match, the longest matching name wins, and a tie at the same length leaves the signal unattributed rather than filed against a guess. So a persona named api does pick up api-server-7f9d-x2q when it is the only persona in the namespace, and loses it to api-server when that persona also exists.

Troubleshooting

Almost always the same cause: those apps have no ApplicationPersona. Dorgu only watches workloads that have one, so a broken app with no persona raises no incident, which means no diagnosis and no proposed fix. Nothing is misconfigured; there is simply nothing registered to watch.Ask Dorgu what it cannot see:
The Unmonitored section names every Deployment with no matching persona and prints an import command per namespace. Then create them from the live Deployments:
Requires CLI v0.9.0 or newer. See dorgu persona import and the brownfield walkthrough.
The persona exists but the operator could not find the workload it describes. The condition message lists every rung it tried:
Check it with:
Common causes, in the order worth checking:
  • The persona is in the wrong namespace. Resolution never leaves the persona’s own namespace. The Deployment being one namespace over looks identical to it not existing.
  • spec.name does not match anything. The chain matches spec.name, not metadata.name. Set spec.name to the Deployment’s name, or add app.kubernetes.io/name: <spec.name> to the Deployment object.
  • The Deployment genuinely is not there yet. The persona re-resolves on the next reconcile, so this clears itself once the workload lands.
dorgu persona import picks a spec.name that resolves back to the Deployment it read, so imported personas do not hit this.
Several Deployments matched the same rung, and picking one arbitrarily is how the wrong workload gets patched. The message names the candidates and the fix:
Do exactly that: app.kubernetes.io/name is rung 1, so setting it on the one Deployment you mean outranks the collision on a later rung. Alternatively, split them into one persona per Deployment with distinct spec.name values.
--workload is no longer an escape hatch here. As of CLI v0.10.0 the Deployment you name must match the one the operator recorded in spec.workloadRef, because an ownership decision made about one workload does not carry to another. When the operator could not resolve the persona at all, managedBy is unknown and the heal is declined regardless. Fix the labels; that is the only path through an AmbiguousDeployment.
Assuming the app does have a persona, check the name prefix and the namespace. Signals are correlated to a persona by name: the persona matches when the resource name equals the persona’s metadata.name or spec.name, or starts with either followed by a hyphen, which is how pod suffixes are handled. A persona named checkout never picks up pods named payments-..., and it never picks up pods in another namespace, because correlation never leaves the signal’s own namespace.Note that this is a different lookup from persona-to-Deployment resolution, which needs no label and no prefix. A persona can be Active, having resolved its Deployment cleanly, and still collect no incidents because the pod names do not prefix-match. See how a persona finds its Deployment.If the pods do prefix-match more than one persona at the same length, the signal is deliberately left unattributed rather than filed against whichever one happened to be listed first. The incident still exists; it is filed against the workload:
Since chart 0.8.0 detection is on by default, so this should not happen on a fresh install. Confirm what the pod is actually running:
You want --enable-health-check=true. If it reads false, either the values set it or the release predates 0.8.0:
Also check that all five CRDs are present. Detection cannot record anything without incidentmemories.dorgu.io. See installation.
The most common cause is a persona with no resource limits. The proposer needs a current value to compute a bounded change, so it skips with a logged persona has no resource limits configured. Add spec.resources.limits to the persona.Other logged skip reasons worth checking: safety check failed: [rate-limit ...] (5 per persona per hour), [concurrent ...] (one active remediation at a time), [deny-list ...] (the namespace is excluded), [blast-radius ...] (the change exceeded 2×), and the dedup skip when an active remediation already covers the same target.
Expected. A single OOM usually raises both OOMKilled and CrashLoopBackOff. An incident is keyed by persona, category, and primary signal, so one application can hold several at once. Dedup runs at the remediation layer, not the incident layer, so you see two incidents and one fix. Approving that one fix resolves both.What will not happen any more is the reverse: two applications sharing one incident. Signals are grouped per application per namespace before anything is diagnosed. See one application at a time.
Give it the stability window. Closing an incident needs the signal absent for the 5-minute grace period and the pods observed Ready, restart-free, and out of any waiting state for a further 6 minutes. A pod that came back 90 seconds ago has not cleared that bar yet.The operator logs the specific reason each cycle at V(1), naming the pod that is holding it open:
Typical reasons are a container that restarted inside the window, a pod in phase Failed (an evicted pod can outlive the failure that produced it, so delete it), or a Deployment that wants replicas and has none running. Anything Dorgu could not read leaves the incident open on purpose: under-reporting an outage is the one failure a reliability tool cannot afford. See resolution needs evidence.
Working as designed. Something else owns that Deployment, so changing it belongs to that owner, and Dorgu declined rather than breaking their next apply. Exit code 4 (ExitDeclined) exists to say “declined by design” rather than “broke”. Nothing was approved and nothing in the cluster changed.The refusal names the owner and prints the owner-shaped steps. Apply the change there, then record the decision:
To see what Dorgu thinks owns it:
If that reads unknown on a workload you know nothing reconciles, ownership detection found a server-side-apply field manager it does not recognise, and managedByDetail names it. unknown is treated as owned on purpose: nothing is patched on a guess. Full detail in the ownership model.
Check the operator version. spec.workloadRef is written by operator v0.9.0 and newer, and CLI v0.10.0 treats an absent record as owned, because it means either an operator too old to know or a workload that could not be read. Neither is evidence that patching is safe.
Empty output against a v0.10.0 CLI means an operator upgrade is what you need. Until then, --no-heal records decisions and you apply changes by hand. See version coupling.
They should, as of operator v0.9.0: every stated fact and every cap comes from the live container, and the explanation names the Deployment and container it read. If you are seeing persona figures quoted as current reality, you are on an older operator.Compare what Dorgu observed against what is running:
An empty string in observedResources means the container does not set that key, which is different from a value of zero, and is what stops a memory fix from adding a CPU limit. A workloadRef with no name means the operator could not resolve the Deployment at all: fix that first, since an unresolved workload loses its grounding and its advisory commands together.
Both halves of that are true, and this is CLI v0.12.0 behaviour rather than a fault. The workload patch landed, so the pod really is fixed and nothing was undone. What failed is the write that records the patch back onto the RemediationAction, and exiting 0 there would reproduce the defect the recording exists to fix: a cluster and a record that disagree.Re-run the heal. It is idempotent and records the change on the way through:
Then confirm the record caught up:
Until it does, treat the remediation record as incomplete rather than as evidence of what is in the cluster. Full detail, including why no paste-ready kubectl patch is offered: when the record cannot be written.
It is not stuck. The operator waits out the verification window — 10 minutes by default — before it will call the fix good. Check the pod itself: if it is 1/1 Running with no new restarts, the heal worked and the phase will settle on its own.
Two causes now, and the second is not a fault.A planner failure. The AI planner degrades to the deterministic rules on any error and logs AI remediation planning failed, falling back to rules with the reason. Check that the operator can reach the Anthropic API (egress and NAT), that the key is valid, and that the startup log includes AI remediation planning enabled. See AI setup.A plan Dorgu refused to persist. Since operator v0.11.0, a plan that diagnoses a resource change but carries nothing appliable is rejected in favour of the rule-based proposal rather than written as a fix that applies nothing. planSource: rule-based is then the correct and auditable record of what actually happened. This is working as intended, and it is strictly better than the alternative it replaced. See the default path.
Upgrade the operator to v0.11.0 or newer. This is the exact shape of the blocker clean-room run #4 found: with AI enabled, the planner described the memory fix as workload-apply steps, which the CRD’s CEL rule forbids from ever being auto-executable, and gave its one persona-update step no patch at all. spec.action.type came out notification, approve had nothing to do, and the pod was still in CrashLoopBackOff 42 minutes later. Nine of nine AI-planned remediations were unappliable.
Until you can upgrade, run rules only. They were healing the identical app on the first try throughout:
An advisory plan that legitimately has nothing to apply is a different thing and says so plainly. See advisory plans.
A guardrail ruled on one of its fields. Since operator v0.11.0 the field a guardrail refused is removed from the step’s patch, precisely so no diff can advertise a change that will not happen, which means the diff is the smaller and more honest of the two numbers.Read what was decided and why:
Every value there is Dorgu’s own arithmetic. See guardrail verdicts as data. On an operator older than v0.11.0 the field is absent and the verdict reaches you only as prose in the step’s rationale.

Next steps

Ownership model

Which workloads Dorgu patches, which it only recommends for, and why

Security and permissions

The operator’s ClusterRole, and the verbs it deliberately lacks

Turn on AI

BYO Anthropic key, injected from a Secret in your own cluster

dorgu remediation

Review, approve, reject, and heal from the CLI

Operator quickstart

Watch the loop run end to end on a real cluster

CRD reference

IncidentMemory, RemediationAction, and DorguEvent schemas