Skip to main content
Dorgu extends Kubernetes with five Custom Resource Definitions (CRDs) that capture application identity, cluster context, and the self-healing record. These CRDs are the central data model that connects the CLI, Operator, and integrations — and they mean your incident history and remediation record live in your own cluster, not a vendor’s database.

CRD Overview

The primary CRD is ApplicationPersona, which describes a single application’s desired operational profile and the Operator’s observed state. ClusterPersona captures cluster-level context. The remaining three are the self-healing loop’s memory — see Self-healing.

Ownership Model

The spec and status of an ApplicationPersona are owned by different actors. This separation is fundamental to Dorgu’s design. The CLI and GitOps tools define the desired state in the spec. The Operator observes the cluster and writes its findings to the status. The one exception is remediation: when you approve a RemediationAction, the Operator applies that action’s JSON merge patch to the ApplicationPersona spec. That is a change you explicitly authorized, scoped to the fields in the proposal you reviewed. Nothing else lets the Operator write a spec.
The Operator never creates or modifies Deployments, Services, or any other workload resource, with or without approval. Its ClusterRole grants get, list, watch on apps/deployments and no write verb at all, so this is enforced by the API server. See security and permissions and the trust model.
If you manage personas through GitOps, an approved remediation’s patch will show as drift on the next sync — mirror the change back into your repo so your pipeline does not revert the fix. dorgu remediation approve --no-heal does not avoid this: it skips the Deployment patch only, and the persona is still updated by the Operator once the action is Approved.The one case where the persona is not updated is a refusal: on a workload something else owns, approve writes nothing at all, neither the Deployment patch nor the status patch that would trigger the persona write. See the ownership model.

API Group and Scope

All CRDs belong to the dorgu.io/v1 API group. ApplicationPersonas live in the same namespace as the workloads they describe. ClusterPersona is cluster-scoped since it represents the entire cluster.

ApplicationPersona Spec Fields

The spec defines the desired operational profile for an application.

ApplicationPersona Status Fields

The status is populated entirely by the Operator and reflects the observed state of the application.

Full ApplicationPersona Example

A complete ApplicationPersona for a critical API service:

ClusterPersona

ClusterPersona is a cluster-scoped resource that captures the overall cluster context. The Operator’s ClusterPersona controller automatically discovers and populates its status.

ClusterPersona Spec Fields

ClusterPersona Status Fields

The ClusterPersona controller reconciles every 5 minutes, scanning nodes, namespaces, and well-known add-on namespaces to keep the status current.
resourceSummary counted pods no node had accepted, until operator v0.11.0. cpuUtilization read 1689% on a cluster where 25% was requested, because every non-terminal pod had its requests summed against node allocatable and a pod sitting in the scheduling queue is non-terminal. That is not an over-estimate to be tightened: an unscheduled pod holds no allocation on any node, and because it can request more than the cluster owns, the error had no upper bound. Only pods a node has accepted are counted now.The schema was right and the code did not match it: this field’s own description already said “the CPU claimed by the resource requests of scheduled pods”.Two consequences worth knowing. This field feeds the dashboard’s cluster view, where a wrong number reads as authoritative in a way terminal output does not, so upgrade the operator if anything renders it. And dorgu health no longer reads it at all: as of CLI v0.12.0 the CLI computes saturation itself from nodes and pods, so the two are corrected from both ends and the terminal figure does not depend on which operator version happens to be installed.

Self-healing policy

spec.policies.selfHealing configures how the cluster heals.
mode is enforced by the proposer: observe records the incident and proposes nothing, propose proposes with approval required, and auto-approve is not implemented — it is accepted but degraded to propose with a warning. maxRemediationsPerHour and excludeNamespaces are enforced by the safety checker. enabled and trustLevel are not enforced: trustLevel is only fed to the AI planner as context, and detection/diagnosis/proposal run regardless of enabled — use mode: observe to stop at diagnosis. The operator’s auto-created persona uses mode: propose.

IncidentMemory

IncidentMemory is the record of a detected problem: what was seen, what caused it, and how it ended. It is created and maintained by the operator’s health-check reconciler. Read them with dorgu incidents.

Spec Fields

rootCause.provider is rule-based or ai-enhanced. resolution.outcome is resolved, partial, failed, or rollback — written by the remediation controller when the loop finishes.

Status Fields

Attribution

spec.attribution records how confidently the incident was tied to an application, and is mirrored to the label dorgu.io/attribution so it is one query: An unattributed incident closes once an attributed one is tracking the same workload, and its resolution.action says handover rather than recovery, because nothing observed the workload.

Resolution

An incident auto-resolves only on positive evidence of recovery: its signal absent for a 5-minute grace period, and its pods observed Ready, restart-free, and out of any waiting state for a 6-minute stability window. Anything else, including any failure to read the cluster, leaves it open. spec.resolution.action records the evidence as auto-resolved: <what was observed>.
attribution and evidence-based resolution are new in operator v0.10.0. In v0.9.0 an incident resolved on the absence of a signal alone, which a crash loop in backoff could produce while still completely dead. See self-healing.

RemediationAction

RemediationAction is a proposed fix. The operator creates it, you approve or reject it, and the operator records what happened. Work with them via dorgu remediation.

Spec Fields

action.type is one of persona-update, notification, or git-pr. action.patch is a JSON merge patch applied to the Persona spec — never to a workload.

workloadRef Fields

The operator populates this from the live Deployment at proposal time. It exists because the persona is a point-in-time import that drifts from the running workload, so every stated fact and every blast-radius cap is grounded here rather than in the persona.
managedBy governs one thing: whether the CLI may patch the Deployment. Only unmanaged permits it. Everything else, including unknown and an entirely absent workloadRef, means Dorgu recommends and does not write.Absent is treated as owned on purpose: it means either an operator older than v0.9.0 or a workload that could not be read, and neither is evidence that patching is safe. Enforced by CLI v0.10.0. See the ownership model.It has no bearing on persona writes. A persona-update step patches the ApplicationPersona, the operator does that itself, and its autoExecutable semantics are unchanged whoever owns the workload.

Step Fields

Each entry in steps[]:
autoExecutable may only be true on a persona-update step. This is enforced at the Kubernetes API server by a CEL validation rule on the CRD (!self.autoExecutable || self.type == 'persona-update'), not just in operator code. Every other step type is advisory, which is what preserves the operator’s guarantee that it never writes workloads.
On an owned workload, steps are reshaped before they are persisted. Where workloadRef.managedBy is anything but unmanaged, the operator drops the command from any step whose command writes to the cluster, rewrites description as what to change at the source (chart values for a Helm release, the Git manifests for an ArgoCD application), and appends one line to rationale on what a direct patch would have broken. Read-only commands such as kubectl logs survive. persona-update steps are never reshaped. See what changes about the plan.

Step Safety Fields

New in operator v0.11.0. Each entry in steps[].safety is one guardrail’s verdict on one field of that step.
Every value in this field is Dorgu’s own arithmetic. No part of it comes from a model, and the model is not permitted to characterise a guardrail’s verdict at all. That is why the field exists rather than being a sentence.It replaces a [safety:blast-radius] … prefix that used to be spliced onto the front of rationale, a string the AI planner authors. A computed refusal therefore read as part of the model’s reasoning, one line below the model’s own claim that the same 16x change was “well within a 2x ceiling”, and nothing on the screen distinguished the measurement from the assertion. A verdict is not the plan’s to give, so it no longer arrives inside the plan’s prose.

rule

verdict

Optional and additive, so nothing has to migrate. safety is absent on every object an operator older than v0.11.0 wrote, and absent means no guardrail ruled. A client that does not know the field renders exactly as it did before and gains no safety key in JSON output. There is no version pinning between the CLI and the operator on account of it: see guardrail verdicts.
A step with no patch is removed from the plan; a step whose patch a guardrail emptied is kept. The difference is what the object is for. A persona-update step carrying no patch applies nothing and instructs nobody, since updating the ApplicationPersona is Dorgu’s own job, and it is what used to render as (no changes) underneath a plan that read like a fix. A step a guardrail emptied carries the record of which field was refused and why, which is the difference between a step that explains an absence and a step that is one.
How the CLI prints these: guardrail verdicts. How they are produced: guardrail verdicts as data.

Status Fields

steps[] is populated, validated, and rendered by the CLI, and currentStep / stepStatuses[] exist in the schema — but the controller currently executes the single spec.action patch rather than walking the plan step by step. autoApproveRule is likewise present in the CRD and ignored by controllers: auto-approve graduation is not implemented.

DorguEvent

DorguEvent is a write-once, classified Kubernetes event record. The event pipeline watches core Kubernetes events, classifies them by severity and category, correlates them to a persona and incident where it can, and stores them. There is no status subresource, so the record is immutable. Records are bounded by age and by count: dorguEvents.retention (default 24h) and dorguEvents.maxRecords (default 2000, oldest pruned first). A per-record spec.ttl overrides the age bound for that record. See DorguEvent retention. Stream them live with dorgu watch events.