Skip to main content

Overview

The dorgu remediation command group is how you work with RemediationAction resources — the fixes the Dorgu Operator proposes after it detects and diagnoses an incident. You list them, read the proposed plan as a diff, then approve or reject. Approving does two things: it records your decision on the RemediationAction, and where Dorgu is allowed to, it heals the running workload. The operator patches the ApplicationPersona (desired state); the CLI patches the Deployment with your credentials. See Self-healing for the full loop.
Dorgu will not patch a Deployment that Helm, ArgoCD, or Flux owns. approve and heal read spec.workloadRef and decline for every managedBy except unmanaged, exiting 4 (ExitDeclined). Nothing is approved and nothing changes; instead you get the owner named and the fix restated as what to change in your source of truth.There is no --force and no override flag, on purpose. Read the ownership model for why. The guard needs CLI v0.10.0 or newer, because an older CLI ignores the record and patches anyway: see versions.
All subcommands shell out to kubectl, so kubectl must be on your PATH and your current context must point at the target cluster. The alias dorgu rem works everywhere dorgu remediation does.

Exit codes

4 is not a failure, and a wrapper that treats every non-zero exit as breakage will now misreport it. Update the wrapper to accept 4, or use dorgu remediation approve <name> -n <ns> --no-heal, which records the decision without a workload patch and exits 0.
New in CLI v0.12.0: a non-zero heal no longer means “nothing happened”. heal and approve can now exit 1 after successfully patching your workload, when the patch landed but the record of it could not be written back to the RemediationAction. The workload is patched by then and nothing is undone.A script that treats non-zero as “the change did not apply, retry or roll forward” will be wrong about the cluster in exactly that case. Read the message: it names which way round the disagreement runs. See when the record cannot be written.

Lifecycle

Two condition reasons are worth knowing, both on the Applied condition. AdvisoryOnly accompanies Acknowledged. PreconditionRejected marks a plan the executor refused before touching the cluster, and is deliberately excluded from the 30-minute failure cooldown, since nothing in the cluster actually went wrong.
Completed only arrives after the verification window — 10 minutes by default (spec.rollback.healthCheckAfter). Your pod recovers within seconds of approve, but the remediation sits in Applying and then Verifying for the rest of that window. This is not stuck. It is the operator waiting long enough to be able to tell a real fix from a temporary one, so it can auto-roll-back if the health regresses.

remediation list

Synopsis

Lists RemediationAction resources. By default it shows only active ones — Pending, Approved, Applying, Verifying, and Failed. Use --all to include Completed, RolledBack, Rejected, and Expired.

Flags

Example output

PLAN: rule-based is what a default install produces. It reads ai-anthropic where you have turned the planner on.

Output columns

The GUARDRAIL column

It appears only when there is something to put in it, sitting beside STEPS. On a cluster where no guardrail has ruled on anything, the column is absent and the list prints exactly what it printed before.
Ranked by how much a reader needs to know: rejected (a field refused outright) over clamped (a value substituted) over derived (a value Dorgu sized itself), with - when no guardrail ruled. It is a pointer, not the record. One word cannot carry a field-by-field account, and it does not try to. Read dorgu remediation diff for what was actually decided. Requires operator v0.11.0 with CLI v0.12.0. Against an older operator the underlying spec.steps[].safety field is absent, the column never appears, and --json gains no safety key.
There is no severity here. RemediationAction carries no severity field, so this command cannot filter or sort by it — the blank SEVERITY column and the --severity flag that matched nothing were both removed. Filter by --phase, and read severity from the linked incident with dorgu incidents describe, where --severity does work.

Examples


remediation diff

Synopsis

Shows the full proposal: the target, the workload and its owner, the confidence, the AI’s plan summary, every step in execution order with its own YAML diff, and what the change does to the running Deployment. This is the review surface. Read it before you approve anything.

Flags

Example output

Each step line reads [order] type (risk; mode): description, with the rationale indented underneath, an optional Run: command, and a unified diff of prePatchState → patch where the step carries one.
There is no Severity: row. RemediationAction carries no severity field. Read severity from the linked incident with dorgu incidents describe.Explanation: and Plan summary: say different things: the summary is the root cause (why it broke), the explanation is the shape of the response (what the plan will and will not do for you). They used to render the same paragraph twice; where an older operator wrote both fields identically, the CLI now prints it once.

The Workload: and Owner: header, and the Deployment change block

Two blocks in that output are about the thing that is actually failing, rather than about the persona. Workload: and Owner: name the live Deployment the operator observed and who reconciles it. The persona is what Dorgu records; the workload is what runs, and on a brownfield cluster the two rarely share a name (persona frontend over Deployment frontend-podinfo). Owner: reads unmanaged (nothing reconciles it, so Dorgu may patch it) when Dorgu may write, and otherwise names the owner: Helm release "frontend" in namespace apps. Both lines are omitted when the operator resolved no workload at all. Deployment change is a workload-to-workload diff, built from workloadRef.observedResources. The old diff read the persona on both sides, which is why a fix that silently introduced a CPU limit the container never had showed nothing at all:
Adding a key the workload does not set is now refused at proposal time, so the (adds a key this workload does not set) marker should not appear on a v0.9.0 operator. It stays in the renderer because objects proposed by an older operator are still readable, and an added key is exactly the thing that must not be invisible.

On an owned workload

The header names the owner, and the suggested actions change: approve is not offered, because that command would be declined.
Once the action leaves Pending, that action block is gone, so the Deployment change block carries the fact instead: Dorgu did not apply this: <owner> owns this Deployment.

Guardrail verdicts

When one of Dorgu’s guardrails has ruled on a field, diff prints what it decided under its own heading, per step, between the rationale and the diff that reflects it.
Every number in that block is Dorgu’s own arithmetic. None of it is the model’s, and the model is not permitted to comment on a cap at all. The heading says so out loud because the lines immediately above it, the description and the rationale, may be a model’s prose, and the reader is about to approve a change to a running workload.It used to arrive as a [safety:blast-radius] … prefix on the model’s rationale. In clean-room run #4 that put Dorgu’s measurement one line below the model’s claim that the same 16x change was “well within a 2x ceiling”, with nothing to tell the reader which of the two had been computed.
Read the three columns of each entry as: what the guardrail decided, on which field, under which rule. Then the facts line, which answers what was asked for, what it was measured against, and what will actually happen. applying nothing is stated rather than left to inference, because a missing word is not a fact a reader should have to notice. The verdicts appear in three places, each of them the last screen before a decision:
Printed once, not twice. The operator writes a guarded step’s description as its own sentence followed by each safety message verbatim, so that a client which knows nothing of the field still delivers the verdict. The CLI takes those messages back out of the description, so the verdict appears once, under the heading, with the numbers beside it, rather than twice, three lines apart.
Requires operator v0.11.0 with CLI v0.12.0. Against an older operator spec.steps[].safety is absent, no block is printed, and the step renders exactly as it did before. Full schema: step safety fields. How the verdicts are produced: guardrail verdicts as data.

Ready-to-run commands on advisory steps

An advisory step is one Dorgu will not carry out for you. Where a single command does the job, the step carries it and diff prints it under the rationale:
This is the difference between a correct diagnosis and an actual fix. Before this existed, Dorgu would identify a mistyped image tag at 0.91 confidence, name the correct tag, link Docker Hub, and then leave you to write the kubectl set image yourself.
The command is only ever printed. Nothing executes it, not the operator and not the CLI. Read it before you run it.Because the field can be written by an AI planner, the operator sanitizes it before storing it and the CLI re-checks it before displaying it: it must be a single line, start with kubectl , contain none of ; & | < > $`, and stay under 1024 characters. A command that fails any of those checks is not shown at all rather than shown with a warning.
Requires operator v0.8.0 and CLI v0.9.0 or newer. Older objects simply carry no command, and no Run: line is printed.

On an owned workload, only read-only commands are printed

A command that writes to a Deployment that Helm or ArgoCD owns takes field ownership from them and breaks their next apply, so it is not something to hand over. On an owned workload the CLI prints only commands it can positively classify as read-only: kubectl logs and kubectl get events matter most on the workloads Dorgu will not patch, because reading is the whole of what is left to hand over. Anything on an unmanaged workload is printed unchanged. Classification is positive, not inferred from the absence of a write verb: a command whose verb matches nothing at all is refused rather than assumed harmless. The verb is found by scanning for the first token that is a known kubectl subcommand, so kubectl -n apps patch ... cannot hide patch behind a flag argument.
The operator already strips workload-writing commands before it stores the object. The CLI establishes read-only-ness again anyway, because the command is model-authored and the CLI reads RemediationAction objects straight out of the cluster, where an older operator or anything with permission to create the CRD could have written a kubectl patch. The shell-metacharacter check still runs first, on owned and unmanaged workloads alike.

Advisory plans

Some plans have nothing the CLI can apply, typically a notification action or a plan whose every step is advisory. Those say so plainly and offer only reject:
Approving one anyway is safe: it records the approval, changes nothing, and settles the action as Acknowledged. It used to be offered approve as its suggested next action, and following that advice failed the remediation and put the app into a 30-minute cooldown.
Only persona-update steps can ever be auto. The Kubernetes API server enforces this with a CEL validation rule on the CRD, so no plan — AI-written or otherwise — can mark a workload change as auto-executable.
Older RemediationAction objects that carry a single spec.action instead of an ordered plan render a single Proposed change: diff rather than a Plan (n steps): block.

remediation approve

Synopsis

Approves a Pending remediation and, when Dorgu owns the write, heals the workload.

What approval actually does

1

The CLI preflights the workload change

Before anything is written: the kube-context guard, the remediation plan, the ownership check, the target Deployment, the container, and the patch. If any of it fails, nothing is approved. On an owned workload this is where the command stops: it prints the refusal and exits 4, having written neither a Deployment patch nor a status patch.
2

The CLI records your decision

It patches the RemediationAction status subresource to phase: Approved, with approvedBy: cli-user and the current timestamp.
3

The operator patches the ApplicationPersona

Seeing Approved, the operator applies the JSON merge patch to the persona’s spec, the app’s desired-state record, then moves the action to Applying. It never touches your Deployment, and it cannot: its ClusterRole has no write verbs on apps/deployments. See security and permissions.
4

The CLI heals the workload

The CLI translates the approved resource change into an equivalent strategic-merge patch on the matching Deployment and applies it with your credentials, so the pod actually restarts with the new limits. Use --no-heal to skip this.
5

The CLI records the patch

A WorkloadPatched condition is stamped on the RemediationAction naming the Deployment, container, and fields it set, so the record and the cluster agree. New in CLI v0.12.0. The phase is already Approved by this point from step 2, so only the condition is written. See what the heal records.
Advisory steps (restart, scale, config-change, manual, and any persona-update that is not a resource change) are printed as numbered manual instructions. They are never executed.
Approval is withheld along with the patch. On an owned workload the gate sits in the preflight, ahead of any write, so approve records nothing at all.That is deliberate rather than incidental. Approving is what tells the operator to patch the persona and start the verification clock, so approving a change the CLI will not apply would leave the persona at 128Mi, the workload at 32Mi, and a ten-minute verification window running over a fix that was never coming. Use --no-heal when you want the decision recorded and intend to apply the change at its source yourself.

On an owned workload

The owner-shaped steps come from the operator, which generated them at proposal time, so the refusal and the plan send you to the same place. Where the plan carries no owner-shaped step, the CLI builds one instruction from the concrete resource change, so a refusal is never a dead end. Read the ownership model for how ownership is detected and what each owner’s failure mode is.

Flags

--heal and --no-heal are mutually exclusive.
--next takes the oldest pending remediation, so the longest-waiting incident goes first; ties break on namespace and name, making the pick reproducible. It previously ranked by severity, which RemediationAction does not carry — every candidate tied, so the winner was effectively arbitrary.
--workload can no longer aim the patch at a different Deployment. It still resolves the workload, but the name must agree with the one in spec.workloadRef, and a mismatch is refused:
Ownership is a fact about one specific object, so a flag that redirects the patch to a Deployment the operator never observed is the guard with a hole in it: Dorgu would clear frontend as unmanaged and then write to frontend-canary, which Helm owns.--container still overrides freely, because ownership is per-Deployment, not per-container. Omit it and the CLI uses the container the operator actually read, so the patch targets the same container whose values were the diff’s before-state.

Examples

Recording a decision without a patch

--no-heal is how you say “I have read this, I agree, and I will apply it myself”. It records phase: Approved, lets the operator patch the persona spec, skips the Deployment patch, and exits 0. It is the intended path in three situations: an owned workload, a GitOps-managed persona, and any time you would rather apply the change through your own pipeline. The CLI warns that the persona and the running workload will disagree until you do apply it. Heed that: the operator’s verification window will run against a workload that has not changed yet.
The target Deployment is resolved before the approval is recorded. If the workload cannot be found, nothing is approved and nothing is changed, and the error lists the Deployments that are present. Approving first and failing to resolve afterwards used to leave the persona at the new limits, the workload at the old ones, and a 10-minute verification window running over a change that never landed.
If a heal does fail after approval, the CLI names the divergence rather than reporting success: the remediation is Approved but <ns>/<deployment> was NOT patched; the persona and the workload now disagree. --no-heal warns about the same divergence up front. That is the divergence where the workload did not change. Since CLI v0.12.0 there is a second, opposite one: the workload changed but the record could not be written. Both exit non-zero, and both name which way round they run, because they call for different responses. See when the record cannot be written. Only Pending remediations can be approved. Anything else exits with an error naming the current phase.

remediation reject

Synopsis

Moves a remediation to Rejected. Works from Pending or Approved — any other phase is refused.

Flags

Rejecting is also a safety gate: dorgu remediation heal refuses to run against a rejected remediation.

remediation heal

Synopsis

Applies an approved remediation’s resource change to the workload on its own, then records that it did. approve runs this for you unless you passed --no-heal, so reach for heal when you deferred the apply, or when a previous heal failed and you want to retry it. It is idempotent, which is what makes re-running it the fix for a heal that patched but could not record.

Flags

How the workload is found

1

Ownership

The first gate. spec.workloadRef.managedBy must be unmanaged. Everything else, including unknown and an absent workloadRef, is declined with the refusal and exit code 4. Nothing after this step runs.
2

Namespace

The persona’s namespace (spec.personaRef.namespace), falling back to the remediation’s own namespace. Never anywhere else.
3

Deployment

Every Deployment in that namespace is a candidate, and the CLI walks the same ordered chain the operator uses, taking the first rung that matches exactly one Deployment: the app.kubernetes.io/name label, then the app label, then metadata.name, then spec.selector.matchLabels. All against the persona’s spec.name.No label is required on the Deployment object. Helm, kustomize, and most hand-written YAML label the pod template only; that resolves on the last rung. Zero matches and an ambiguous rung are both errors, and both list the Deployments actually present in the namespace and point at --workload. See discovery.Whatever this resolves to, and however it was resolved, it must agree with spec.workloadRef.name. A mismatch is refused rather than patched.
4

Container

The container the operator observed, unless --container names another. Failing that, the only container if there is one, otherwise the container whose name matches the app. Anything else asks for --container.
5

Patch

A strategic-merge patch that sets exactly the resources.limits and resources.requests fields the remediation changed (cpu and memory only), and nothing broader than the proposal. It runs under the field manager dorgu rather than kubectl’s default kubectl-patch, so the entry it creates is distinguishable from a kubectl patch you ran yourself.
6

Release

The dorgu entry is then removed from metadata.managedFields and the object is read back to confirm it is gone, so a heal leaves no ownership footprint and a later server-side apply does not conflict with it. If the removal fails, the heal still succeeded and the CLI says loudly what is left behind and how to clear it. New in CLI v0.11.0. See Dorgu leaves no field manager behind.
7

Record

The RemediationAction is updated to say what just happened, so the record and the cluster agree. New in CLI v0.12.0. See what the heal records.

What the heal records

New in CLI v0.12.0. Two writes follow a patch the cluster accepted, and only ever follow one:
This fixes a heal that reported success and left the record wrong. heal patched a workload from 64Mi to 128Mi, printed ✓ Healed, exited 0, and left the RemediationAction on Pending. The two things Dorgu keeps then disagreed: the cluster carried the new limits, and the record said nobody had approved anything and nothing had happened. For a product sold on organizational memory, memory that contradicts the cluster is worse than none.The phase transition also closed a half nobody had noticed. A Pending action is a no-op to the operator, so the persona never learned the new limits either, which left the app’s desired-state record out of step with the cluster as well.
Two design points, because both are easy to assume otherwise:
  • The marker is a condition, not status.appliedAt. appliedAt is the operator’s: it is stamped when the persona patch lands and read during Applying to time the verification window, so writing to it would move a clock the CLI does not drive. It is also needed on top of the phase transition, because a re-heal after --no-heal starts from a phase already past Approved, where a phase change would be both wrong and invisible.
  • The CLI does not move the phase of an action already Approved, Applying, Verifying, or terminal. That lifecycle belongs to the operator, and a CLI writing into the middle of a state machine it does not drive is a different bug from the one being fixed. Those heals stamp the condition and leave the phase alone.
The write is a read-modify-write, because conditions are keyed by type and adding one means sending the whole list. Every other writer’s condition goes back as raw JSON byte for byte, the CLI’s own prior marker is replaced rather than appended, the resourceVersion is sent as a precondition, and a conflict is retried against fresh state three times before it is reported. Round-tripping the operator’s conditions through a struct would silently drop fields the CLI does not know about, and quietly rewriting the operator’s record while fixing a record bug is a poor trade.

When the record cannot be written

The workload is patched by then, so nothing is undone. But the command exits 1 and says which way round the disagreement runs, because exiting 0 there would reproduce the defect above:
Note that the success line still prints, because the heal genuinely succeeded. It is the record that failed, and the two are reported separately.
No paste-ready kubectl patch is offered, on purpose. The only patch that would work replaces the whole conditions list, so handing one over would have you overwrite the operator’s own conditions to fix a record bug. Re-running the heal is both shorter and safe: it is idempotent, and it records the change on the way through.
This is the same class as the field-manager footprint warning, with one deliberate difference. A failed footprint strip leaves the heal green, because it is a future apply conflict rather than a wrong record. A failed record write is not green, because a record that contradicts the cluster is the thing being fixed.

Safety gates

The ownership gate runs before the plan is resolved against the cluster, and an advisory-only plan never reaches it: there is no workload change to refuse, so an advisory remediation behaves identically on owned and unmanaged workloads.
heal only auto-applies resource limits and requests (cpu, memory) — the OOM and saturation path. A remediation with no resource change reports that there is nothing to heal automatically and prints its advisory steps instead.

Next steps

Ownership model

Why a Helm or ArgoCD-owned workload is declined, and what to do instead

Self-healing

How detection, diagnosis, planning, and verification fit together

AI setup

Turn on the AI planner with your own Anthropic key

dorgu incidents

Inspect the IncidentMemory a remediation was proposed for

CRD reference

Full RemediationAction schema, including steps[] and rollback