Overview
Thedorgu remediation command group is how you work with RemediationAction resources — the fixes the Dorgu Operator proposes after it detects and diagnoses an incident. You list them, read the proposed plan as a diff, then approve or reject.
Approving does two things: it records your decision on the RemediationAction, and where Dorgu is allowed to, it heals the running workload. The operator patches the ApplicationPersona (desired state); the CLI patches the Deployment with your credentials. See Self-healing for the full loop.
All subcommands shell out to
kubectl, so kubectl must be on your PATH and your current context must point at the target cluster. The alias dorgu rem works everywhere dorgu remediation does.Exit codes
Lifecycle
Two condition reasons are worth knowing, both on the
Applied condition. AdvisoryOnly accompanies Acknowledged. PreconditionRejected marks a plan the executor refused before touching the cluster, and is deliberately excluded from the 30-minute failure cooldown, since nothing in the cluster actually went wrong.
remediation list
Synopsis
RemediationAction resources. By default it shows only active ones — Pending, Approved, Applying, Verifying, and Failed. Use --all to include Completed, RolledBack, Rejected, and Expired.
Flags
Example output
PLAN: rule-based is what a default install produces. It reads ai-anthropic where you have turned the planner on.
Output columns
The GUARDRAIL column
It appears only when there is something to put in it, sitting besideSTEPS. On a cluster where no guardrail has ruled on anything, the column is absent and the list prints exactly what it printed before.
rejected (a field refused outright) over clamped (a value substituted) over derived (a value Dorgu sized itself), with - when no guardrail ruled.
It is a pointer, not the record. One word cannot carry a field-by-field account, and it does not try to. Read dorgu remediation diff for what was actually decided.
Requires operator v0.11.0 with CLI v0.12.0. Against an older operator the underlying spec.steps[].safety field is absent, the column never appears, and --json gains no safety key.
There is no severity here.
RemediationAction carries no severity field, so this command cannot filter or sort by it — the blank SEVERITY column and the --severity flag that matched nothing were both removed. Filter by --phase, and read severity from the linked incident with dorgu incidents describe, where --severity does work.Examples
remediation diff
Synopsis
Flags
Example output
[order] type (risk; mode): description, with the rationale indented underneath, an optional Run: command, and a unified diff of prePatchState → patch where the step carries one.
There is no
Severity: row. RemediationAction carries no severity field. Read severity from the linked incident with dorgu incidents describe.Explanation: and Plan summary: say different things: the summary is the root cause (why it broke), the explanation is the shape of the response (what the plan will and will not do for you). They used to render the same paragraph twice; where an older operator wrote both fields identically, the CLI now prints it once.The Workload: and Owner: header, and the Deployment change block
Two blocks in that output are about the thing that is actually failing, rather than about the persona.
Workload: and Owner: name the live Deployment the operator observed and who reconciles it. The persona is what Dorgu records; the workload is what runs, and on a brownfield cluster the two rarely share a name (persona frontend over Deployment frontend-podinfo). Owner: reads unmanaged (nothing reconciles it, so Dorgu may patch it) when Dorgu may write, and otherwise names the owner: Helm release "frontend" in namespace apps. Both lines are omitted when the operator resolved no workload at all.
Deployment change is a workload-to-workload diff, built from workloadRef.observedResources. The old diff read the persona on both sides, which is why a fix that silently introduced a CPU limit the container never had showed nothing at all:
Adding a key the workload does not set is now refused at proposal time, so the
(adds a key this workload does not set) marker should not appear on a v0.9.0 operator. It stays in the renderer because objects proposed by an older operator are still readable, and an added key is exactly the thing that must not be invisible.On an owned workload
The header names the owner, and the suggested actions change:approve is not offered, because that command would be declined.
Pending, that action block is gone, so the Deployment change block carries the fact instead: Dorgu did not apply this: <owner> owns this Deployment.
Guardrail verdicts
When one of Dorgu’s guardrails has ruled on a field,diff prints what it decided under its own heading, per step, between the rationale and the diff that reflects it.
applying nothing is stated rather than left to inference, because a missing word is not a fact a reader should have to notice.
The verdicts appear in three places, each of them the last screen before a decision:
Printed once, not twice. The operator writes a guarded step’s
description as its own sentence followed by each safety message verbatim, so that a client which knows nothing of the field still delivers the verdict. The CLI takes those messages back out of the description, so the verdict appears once, under the heading, with the numbers beside it, rather than twice, three lines apart.spec.steps[].safety is absent, no block is printed, and the step renders exactly as it did before. Full schema: step safety fields. How the verdicts are produced: guardrail verdicts as data.
Ready-to-run commands on advisory steps
An advisory step is one Dorgu will not carry out for you. Where a single command does the job, the step carries it anddiff prints it under the rationale:
kubectl set image yourself.
Requires operator v0.8.0 and CLI v0.9.0 or newer. Older objects simply carry no command, and no Run: line is printed.
On an owned workload, only read-only commands are printed
A command that writes to a Deployment that Helm or ArgoCD owns takes field ownership from them and breaks their next apply, so it is not something to hand over. On an owned workload the CLI prints only commands it can positively classify as read-only:kubectl logs and kubectl get events matter most on the workloads Dorgu will not patch, because reading is the whole of what is left to hand over. Anything on an unmanaged workload is printed unchanged.
Classification is positive, not inferred from the absence of a write verb: a command whose verb matches nothing at all is refused rather than assumed harmless. The verb is found by scanning for the first token that is a known kubectl subcommand, so kubectl -n apps patch ... cannot hide patch behind a flag argument.
The operator already strips workload-writing commands before it stores the object. The CLI establishes read-only-ness again anyway, because the command is model-authored and the CLI reads
RemediationAction objects straight out of the cluster, where an older operator or anything with permission to create the CRD could have written a kubectl patch. The shell-metacharacter check still runs first, on owned and unmanaged workloads alike.Advisory plans
Some plans have nothing the CLI can apply, typically anotification action or a plan whose every step is advisory. Those say so plainly and offer only reject:
Acknowledged. It used to be offered approve as its suggested next action, and following that advice failed the remediation and put the app into a 30-minute cooldown.
Only
persona-update steps can ever be auto. The Kubernetes API server enforces this with a CEL validation rule on the CRD, so no plan — AI-written or otherwise — can mark a workload change as auto-executable.RemediationAction objects that carry a single spec.action instead of an ordered plan render a single Proposed change: diff rather than a Plan (n steps): block.
remediation approve
Synopsis
Pending remediation and, when Dorgu owns the write, heals the workload.
What approval actually does
1
The CLI preflights the workload change
Before anything is written: the kube-context guard, the remediation plan, the ownership check, the target Deployment, the container, and the patch. If any of it fails, nothing is approved. On an owned workload this is where the command stops: it prints the refusal and exits
4, having written neither a Deployment patch nor a status patch.2
The CLI records your decision
It patches the RemediationAction status subresource to
phase: Approved, with approvedBy: cli-user and the current timestamp.3
The operator patches the ApplicationPersona
Seeing
Approved, the operator applies the JSON merge patch to the persona’s spec, the app’s desired-state record, then moves the action to Applying. It never touches your Deployment, and it cannot: its ClusterRole has no write verbs on apps/deployments. See security and permissions.4
The CLI heals the workload
The CLI translates the approved resource change into an equivalent strategic-merge patch on the matching Deployment and applies it with your credentials, so the pod actually restarts with the new limits. Use
--no-heal to skip this.5
The CLI records the patch
A
WorkloadPatched condition is stamped on the RemediationAction naming the Deployment, container, and fields it set, so the record and the cluster agree. New in CLI v0.12.0. The phase is already Approved by this point from step 2, so only the condition is written. See what the heal records.restart, scale, config-change, manual, and any persona-update that is not a resource change) are printed as numbered manual instructions. They are never executed.
On an owned workload
Flags
--heal and --no-heal are mutually exclusive.
--next takes the oldest pending remediation, so the longest-waiting incident goes first; ties break on namespace and name, making the pick reproducible. It previously ranked by severity, which RemediationAction does not carry — every candidate tied, so the winner was effectively arbitrary.Examples
Recording a decision without a patch
--no-heal is how you say “I have read this, I agree, and I will apply it myself”. It records phase: Approved, lets the operator patch the persona spec, skips the Deployment patch, and exits 0.
It is the intended path in three situations: an owned workload, a GitOps-managed persona, and any time you would rather apply the change through your own pipeline.
The CLI warns that the persona and the running workload will disagree until you do apply it. Heed that: the operator’s verification window will run against a workload that has not changed yet.
The target Deployment is resolved before the approval is recorded. If the workload cannot be found, nothing is approved and nothing is changed, and the error lists the Deployments that are present. Approving first and failing to resolve afterwards used to leave the persona at the new limits, the workload at the old ones, and a 10-minute verification window running over a change that never landed.
the remediation is Approved but <ns>/<deployment> was NOT patched; the persona and the workload now disagree. --no-heal warns about the same divergence up front.
That is the divergence where the workload did not change. Since CLI v0.12.0 there is a second, opposite one: the workload changed but the record could not be written. Both exit non-zero, and both name which way round they run, because they call for different responses. See when the record cannot be written.
Only Pending remediations can be approved. Anything else exits with an error naming the current phase.
remediation reject
Synopsis
Rejected. Works from Pending or Approved — any other phase is refused.
Flags
dorgu remediation heal refuses to run against a rejected remediation.
remediation heal
Synopsis
approve runs this for you unless you passed --no-heal, so reach for heal when you deferred the apply, or when a previous heal failed and you want to retry it.
It is idempotent, which is what makes re-running it the fix for a heal that patched but could not record.
Flags
How the workload is found
1
Ownership
The first gate.
spec.workloadRef.managedBy must be unmanaged. Everything else, including unknown and an absent workloadRef, is declined with the refusal and exit code 4. Nothing after this step runs.2
Namespace
The persona’s namespace (
spec.personaRef.namespace), falling back to the remediation’s own namespace. Never anywhere else.3
Deployment
Every Deployment in that namespace is a candidate, and the CLI walks the same ordered chain the operator uses, taking the first rung that matches exactly one Deployment: the
app.kubernetes.io/name label, then the app label, then metadata.name, then spec.selector.matchLabels. All against the persona’s spec.name.No label is required on the Deployment object. Helm, kustomize, and most hand-written YAML label the pod template only; that resolves on the last rung. Zero matches and an ambiguous rung are both errors, and both list the Deployments actually present in the namespace and point at --workload. See discovery.Whatever this resolves to, and however it was resolved, it must agree with spec.workloadRef.name. A mismatch is refused rather than patched.4
Container
The container the operator observed, unless
--container names another. Failing that, the only container if there is one, otherwise the container whose name matches the app. Anything else asks for --container.5
Patch
A strategic-merge patch that sets exactly the
resources.limits and resources.requests fields the remediation changed (cpu and memory only), and nothing broader than the proposal. It runs under the field manager dorgu rather than kubectl’s default kubectl-patch, so the entry it creates is distinguishable from a kubectl patch you ran yourself.6
Release
The
dorgu entry is then removed from metadata.managedFields and the object is read back to confirm it is gone, so a heal leaves no ownership footprint and a later server-side apply does not conflict with it. If the removal fails, the heal still succeeded and the CLI says loudly what is left behind and how to clear it. New in CLI v0.11.0. See Dorgu leaves no field manager behind.7
Record
The
RemediationAction is updated to say what just happened, so the record and the cluster agree. New in CLI v0.12.0. See what the heal records.What the heal records
New in CLI v0.12.0. Two writes follow a patch the cluster accepted, and only ever follow one:- The marker is a condition, not
status.appliedAt.appliedAtis the operator’s: it is stamped when the persona patch lands and read duringApplyingto time the verification window, so writing to it would move a clock the CLI does not drive. It is also needed on top of the phase transition, because a re-heal after--no-healstarts from a phase already pastApproved, where a phase change would be both wrong and invisible. - The CLI does not move the phase of an action already
Approved,Applying,Verifying, or terminal. That lifecycle belongs to the operator, and a CLI writing into the middle of a state machine it does not drive is a different bug from the one being fixed. Those heals stamp the condition and leave the phase alone.
The write is a read-modify-write, because conditions are keyed by type and adding one means sending the whole list. Every other writer’s condition goes back as raw JSON byte for byte, the CLI’s own prior marker is replaced rather than appended, the
resourceVersion is sent as a precondition, and a conflict is retried against fresh state three times before it is reported. Round-tripping the operator’s conditions through a struct would silently drop fields the CLI does not know about, and quietly rewriting the operator’s record while fixing a record bug is a poor trade.When the record cannot be written
The workload is patched by then, so nothing is undone. But the command exits1 and says which way round the disagreement runs, because exiting 0 there would reproduce the defect above:
No paste-ready
kubectl patch is offered, on purpose. The only patch that would work replaces the whole conditions list, so handing one over would have you overwrite the operator’s own conditions to fix a record bug. Re-running the heal is both shorter and safe: it is idempotent, and it records the change on the way through.Safety gates
The ownership gate runs before the plan is resolved against the cluster, and an advisory-only plan never reaches it: there is no workload change to refuse, so an advisory remediation behaves identically on owned and unmanaged workloads.
Next steps
Ownership model
Why a Helm or ArgoCD-owned workload is declined, and what to do instead
Self-healing
How detection, diagnosis, planning, and verification fit together
AI setup
Turn on the AI planner with your own Anthropic key
dorgu incidents
Inspect the IncidentMemory a remediation was proposed for
CRD reference
Full RemediationAction schema, including
steps[] and rollback