Skip to main content

Synopsis

Description

Display a comprehensive cluster health summary by querying the Kubernetes API directly. The output includes node status, resource saturation (CPU and memory), control plane component health, active incidents, and pending remediations. When the Dorgu Operator is installed, the command adds IncidentMemory and RemediationAction data. Without the operator it degrades to nodes, resource saturation, and control plane, all of which come from the Kubernetes API and need nothing installed. It also names what it cannot see: any Deployment with no matching ApplicationPersona is listed under Unmonitored, because a workload Dorgu is not watching can never appear as an incident.
Breaking in CLI v0.12.0: --json reports saturation in a new shape, and the old used key never meant used.resourceSaturation.{cpu,memory} carried {percentage, used, allocatable}, where used actually held the requests the scheduler had committed. On the cluster that surfaced this, requested was 25% and used was 1%, which is the difference between nearly full and nearly idle. Anything reading the old used key was reading the wrong number, not a stale one.The shape is now {allocatable, requested, requestedPercent, used, usedPercent}, with nodes, scheduledPods, unscheduledPods and usedUnavailable added alongside. Renaming in place was rejected on purpose: a key that silently changes meaning is the defect, not the fix.If you script against dorgu health --json, see JSON output before upgrading. The terminal output changed too, and resource saturation covers both.

Flags

Output sections

Resource saturation

New in CLI v0.12.0. Saturation is computed by the CLI from the node list it already fetched plus a pod list, and reports requested and used as two separate columns:
They are two columns because they call for different actions and used to share one line. The header states which is which, so 25% can never be read as usage.
This replaces a figure that could be wrong without limit. dorgu health printed 1689% CPU on a cluster where 25% was requested and 1% was in use. The number was not the CLI’s: it was ClusterPersona.status.resourceSummary.cpuUtilization rendered verbatim, and the operator computed it as the requests of every non-terminal pod over node allocatable. A pod no node has accepted is non-terminal, so a backlog of unschedulable pods inflated the figure with no upper bound, because such a pod can request more than the cluster owns.Computing it in the CLI is deliberate rather than duplicated work. dorgu health documents itself as querying the Kubernetes API directly, saturation was the one section that did not, and a figure read from a five-minute-stale field on whatever operator version happens to be installed cannot be relied on. The same defect is fixed at the source in operator v0.11.0, because the dashboard reads that field.

Unscheduled pods are excluded, and counted

A pod with no spec.nodeName holds no allocation on any node, so counting its requests against allocatable is a category error rather than an approximation. Those pods are excluded from the figures and reported as the real problem the old percentage was burying:
A Pending pod that has been bound to a node still counts, because a reservation held while an image pulls is a real reservation. Succeeded and Failed pods are excluded too, for the different reason that they ran to completion and released what they held.

A resource close to booked out is named

At 90% of allocatable requested or above, the command says so rather than leaving you to do the division it had previously got wrong:
No saturation incident is raised at any level. Detection is the operator’s job, and this section only reports. A cluster at 95% requested prints the warning above and opens no IncidentMemory.

A missing used figure says why

metrics-server is not installed by default, so its absence renders as a reason rather than as zero:
Requested still reports without it. An absent figure is never printed as 0, and never as a blank operand.
The whole section is bounded by a 15-second deadline of its own, because listing every pod is the expensive call in dorgu health and one slow pod list should not take the node table, the incident list, and the exit code down with it. On a timeout the section is omitted and the reason is printed. Saturation also needs the node list to have succeeded, since allocatable comes from it.
dorgu health --watch labels the operator’s figure cpu-requested= and mem-requested=. The stream carries ClusterPersona values rather than the CLI’s own computation, and a bare cpu=25% reads as usage when it is requests.

Unmonitored Deployments

Dorgu only watches workloads that have an ApplicationPersona. A Deployment without one raises no incident no matter what it does, so reporting Active Incidents: 0 for a cluster full of unmonitored apps would present a blind spot as health. The Unmonitored section names them:
Coverage is decided with the same discovery chain the operator uses, so a Deployment labelled only on its pod template counts as monitored. See dorgu persona import.

Exit codes

Without --exit-code the command exits 0 whenever it managed to read the cluster, so interactive use is unchanged.
An unreachable cluster always exits 1, with no flag to opt out. The command probes the API server before it renders anything, and no summary is produced from a failed call. It used to exit 0 with an empty node table and Active Incidents: 0, which meant a monitoring script could not tell “healthy” from “on fire” from “cannot see the cluster”.If you have a job that alerts on any non-zero exit, it will now alert when the cluster is unreachable. That is intended, but check the job before upgrading if a transient VPN or kubeconfig failure would page someone.
The API-server probe reads /version, which any authenticated user can read, so a namespace-scoped role is not mistaken for an unreachable cluster.

Examples

JSON output

resourceSaturation changed shape in CLI v0.12.0, and the change is breaking. The old used key held requests, so a consumer that trusted the key name was reading the wrong number rather than an out-of-date one. Nothing was renamed in place, precisely so that a script reading used now gets used instead of quietly continuing to get requests under a name that had started meaning something else.Migrating: a script that read resourceSaturation.cpu.used as “how full is the cluster” wants requested or requestedPercent. One that genuinely wanted consumption wants the new used, and must handle its absence, since metrics-server is not installed by default. Read unscheduledPods too: a non-zero value means pods were excluded from requested, which is the case the old figure got catastrophically wrong.
When --json is set, the output follows this structure:
unmonitored.items is the full, uncapped list even when the human-readable output truncated it. unmonitored.personaCRDMissing is present and true when the ApplicationPersona CRD is not installed. Where metrics-server did not answer, the used half is omitted rather than zeroed, and the reason is stated once at the top level:
resourceSaturation itself is omitted entirely when the node list could not be read or the pod list timed out, since allocatable and requests both come from those. An absent object means “not measured”, which is why it is absent rather than a set of zeroes. When the incident records cannot be read at all, activeIncidents carries "unavailable": true and a "reason" string instead of a count of 0, so a reader can tell “none” from “unknown”:
Requires kubectl in your PATH. Nodes, resource saturation, and control plane come from the Kubernetes API and need no operator installed; saturation in particular used to require one and no longer does. The used half of saturation needs metrics-server, and says so when it is missing. Incident and remediation data requires the IncidentMemory and RemediationAction CRDs; when they are absent that is reported as unavailable rather than as a count of zero.