Dorgu
Dorgu is an open-source AI SRE for Kubernetes — built for teams that don’t have a platform team. It runs on your own cluster, watches your apps, and when something breaks it doesn’t just alert. It diagnoses the root cause, proposes a reviewable fix, and heals the workload with your approval. Code detects. AI explains. A human approves.See it in three minutes
A 3-minute screencast recorded on a real EKS cluster: an app is deployed with too little memory, gets OOMKilled, and crash-loops. Dorgu detects it, Claude diagnoses the root cause and proposes an ordered fix, one command approves it, and the pod recovers. No audio — the terminal is the whole story.
The self-healing loop
1
Detect
The operator’s health-check reconciler watches for OOMKills, crash loops, image-pull failures, long-pending pods, resource saturation, node pressure, and control-plane trouble — and opens an
IncidentMemory.2
Diagnose
Deterministic rules produce a root cause and a confidence score. They are always on. With your own Anthropic key, Claude enhances the diagnosis with cluster context.
3
Propose
Deterministic rules compute a bounded fix against the live container and return an ordered plan with a per-step YAML diff. That is the default path and needs no key. With the AI planner turned on, Claude writes the plan instead, reading your app persona, cluster persona, past incidents, and which past fixes actually worked. Either way the guardrails are the same, only persona updates are ever auto-executable, and every verdict a guardrail reaches is Dorgu’s own arithmetic.
4
Approve
You review the plan with
dorgu remediation diff and approve with one command. Approval is required by default.5
Heal, or hand over
The operator updates the app’s desired state. Where nothing else reconciles the Deployment, the CLI applies the change with your credentials and the pod recovers. Where Helm, ArgoCD, or Flux owns it, Dorgu names the owner and tells you what to change in your source of truth instead. It does not fight your deployment tool.
6
Remember
The operator verifies the fix, rolls it back automatically if health regresses, and keeps the incident and its outcome as organizational memory for next time.
The operator never writes workloads. It reads cluster state and updates its own CRDs. Deployments, Services, and everything else stay under the control of ArgoCD, Flux,
kubectl, or the Dorgu CLI. This is an architectural invariant, not a setting, and it is enforced by RBAC rather than by convention: the operator’s ClusterRole is published in full, with no create, update, or patch on Deployments and no access to Secrets.And the CLI will not write workloads you have not handed it. It detects who owns each Deployment and refuses to patch one that Helm, ArgoCD, or Flux reconciles, because that patch would break the owner’s next apply. Those three run controllers and stamp what they reconcile; kustomize does neither, and Dorgu says so rather than claiming otherwise. See the ownership model.Products
Dorgu Operator
The AI SRE: detects incidents, diagnoses root cause, and proposes reviewable fixes with guardrails.
Dorgu CLI
Review and approve remediations, inspect cluster health, and generate K8s manifests and personas.
Key capabilities
- AI self-healing loop — detect → diagnose → propose → approve → heal → remember, with human approval required by default
- Respects your deployment tool: Dorgu detects who owns each workload and refuses to patch a Helm, ArgoCD, or Flux-managed Deployment. It tells you what to change in your source of truth instead, so your next
helm upgradestill works. And where it does patch, it removes its own field-manager entry afterwards, so it leaves no ownership footprint for your next apply to fight - Grounded in the live workload: every number Dorgu states, and every cap it computes, is read from the running container rather than from a stale record
- Ordered, reviewable fixes — a plan with per-step diffs, rationale, and risk levels, not an opaque action
- Real guardrails: a 2× blast-radius cap measured against the live value, no resource keys added that the workload does not set, rate limits, one remediation per incident,
kube-systemdeny-listed, automatic rollback on regression - A published permission model: the operator’s ClusterRole in full, with the denied verbs called out, so “it never writes workloads” is checkable rather than asserted
- Kubernetes-native memory —
IncidentMemoryandRemediationActionCRDs, so the record lives in your cluster, not a vendor’s database - BYO key, and AI off by default: Anthropic only, operator-side. The deterministic path is the default and the current recommendation, and it heals on its own. Why
- ApplicationPersona CRDs — living identity documents for your apps, and the desired-state record a fix is applied to
- Manifest generation — analyze a Dockerfile, docker-compose file, or source tree and produce Deployment, Service, Ingress, HPA, ArgoCD Application, and CI/CD workflow
- Cluster bootstrapping — an interactive wizard installs cert-manager, ingress-nginx, CloudNativePG, OpenObserve, Argo CD, and External Secrets
Next steps
Self-healing quickstart
Break an app and watch Dorgu fix it, in under 15 minutes
Ownership model
Which workloads Dorgu changes, which it only recommends for, and why
Security and permissions
The operator’s ClusterRole, published in full
Turn on AI
Enable the AI planner with your own Anthropic key
Installation
Install the Dorgu CLI via Go, binary release, or from source
Manifest generation
Generate Kubernetes manifests from your app