Platform · Live Monitoring

Your whole stack, on one pane of glass.

CPU, memory, node health, pod status and storage, streaming live across every environment you run, on one screen. No console hopping, no CLIs.

  • Live health across every environment
  • One screen, no console hopping
  • 7 health states with instant drill-down
kubentic · live-monitoring

Overview

Why fragmented monitoring is the real outage

Most teams don't lack monitoring, they have too much of it, scattered across the EKS console, GCP Monitoring, the AKS blade, and a Grafana per region, each with its own idea of what "healthy" means. The cost isn't missing data; it's the minutes spent stitching four dashboards together at 2am to answer one question. Live Monitoring collapses that into one place where CPU, memory, node health, pod status, and storage stream live across every cloud, in one shared health vocabulary.

Under the hood it maps every node, pod, and workload to one of 7 discrete health states, richer than a green/red light, and lets you drill from a fleet-wide grid to a single pod's restart timeline through region, cluster, namespace, and node. Connected and imported clusters report at full parity, so a self-managed cluster sits next to your managed EKS with the same states and the same drill path. It's the live substrate the rest of the platform reads from: the same signals feed Foresight AI's predictions and the Browser Terminal's context.

Live Monitoring

Your whole fleet, one screen.

Live everything

CPU, memory, node health, pod status and storage, streaming live, with no refresh.

One screen

Every environment across every cloud on one view, no console hopping, no CLIs.

Drill in instantly

Seven health states with one-click drill-down to the node, pod or namespace.

Capabilities

Everything inside Live Monitoring.

Six-level drill path to a single pod

Navigate fleet to region to cluster to namespace to node to pod without leaving the dashboard, and every layer carries the same 7-state health model. A red tile at the top always resolves to a concrete offender at the bottom: one pod's restart timeline, one node's condition, one PVC nearing full.

Seven health states, not up or down

Every node, pod, and workload resolves to Healthy, Degraded, Warning, Critical, Pending, Unknown, or Terminated. Degraded surfaces the pod that's throttled but still serving; Unknown flags the node that stopped reporting before it silently drops off your view, so you triage by severity instead of chasing color changes.

Node capacity and pressure telemetry

Per node you see live CPU and memory pressure, allocatable-versus-requested headroom, disk and PID pressure conditions, and Ready/NotReady transitions as they land. When a node crosses into memory pressure, you see which pods it's about to evict before the eviction fires.

Restart, crash-loop, and OOMKilled signals

Watch pod phase, container readiness (e.g. 2/3 ready), and restart counts stream in as the Kubelet emits them. CrashLoopBackOff and OOMKilled are called out inline with their back-off state, so a restart spike surfaces on the tile instead of hiding three screens deep in the CLI.

PVC saturation with projected fill window

Track PersistentVolumeClaim usage, bound-versus-pending volume status, and per-volume fill rate across every StorageClass in the fleet. A PVC trending toward capacity shows its projected fill window, so you resize before writes start failing rather than after the first ENOSPC.

EKS, GKE, AKS, and self-managed side by side

Managed and self-managed clusters render in the same grid under one health vocabulary, no per-cloud console, no context switch, no separate tool per region. Every connected or imported environment reports into that grid the moment it comes under Kubentic, at full parity.

How it works

Visibility the second an environment connects.

01

Connect or import

New or existing environments start reporting the moment they come under Kubentic.

02

Stream live

Node, pod and storage metrics flow into one real-time view.

03

Act with context

Drill from a fleet overview to a single failing pod in a couple of clicks.

Built for real work

Where Live Monitoring earns its keep.

The platform lead running 40 clusters across three clouds

Instead of juggling the EKS console, GCP Monitoring, and the AKS blade, they open one grid of 7-state health tiles each morning and scan it in seconds. A red tile on a GCP staging cluster in Mumbai is two clicks from the exact pod that's crash-looping, no console hopping, no per-region dashboards.

The on-call engineer triaging a 2am pod alert

Paged about a namespace going sideways, they drill from the fleet view to the cluster, filter to the namespace, and see which pods flipped to Degraded, their restart counts, and the node they share. The shared node's memory-pressure condition tells the whole story in under a minute, no SSH, no CLI archaeology.

The engineer who just imported a legacy self-managed cluster

A self-managed cluster that was a monitoring blind spot starts reporting node, pod, and storage metrics the moment it's connected. Within seconds it joins the fleet grid at parity with the managed ones, and the team finally sees the PVCs that had been filling for weeks.

At a glance

The technical details.

Metrics streamed
Node CPU/memory/disk pressure, pod phase & readiness, restart counts, PVC usage, live
Health states
7, Healthy, Degraded, Warning, Critical, Pending, Unknown, Terminated
Cluster support
EKS, GKE, AKS, and self-managed in one unified view
Drill-down depth
6 levels, fleet, region, cluster, namespace, node, pod
Fleet coverage
100% of connected & imported environments, 24/7
Setup for imported clusters
Reports on connect, no separate agents to stand up
Health states
7

Health states

Live metrics
24/7

Live metrics

Unified view
1

Unified view

Fleet coverage
100%

Fleet coverage

FAQ

Questions, answered.

What exactly are the 7 health states, and why not just up/down?

Every node, pod, and workload resolves to Healthy, Degraded, Warning, Critical, Pending, Unknown, or Terminated. The extra states are what let you triage by severity: Degraded catches the pod that's throttled but still serving traffic, and Unknown flags a node that stopped reporting before it falls off your view. A binary light hides both of those until they become outages.

How fresh are the metrics, real-time or polled?

Metrics stream live 24/7 across every connected environment, so node CPU, memory, pod phase, restart counts, and PVC fill update continuously rather than on a manual refresh. When a pod flips to CrashLoopBackOff or a node crosses into memory pressure, the state change surfaces as the Kubelet emits it. You're watching the fleet as it is, not a snapshot from your last page load.

Does this work on clusters I import, or only ones I ship with Kubentic?

Both, at full parity. The moment a cluster is connected or imported it reports live, including self-managed clusters that were previously blind spots. Imported clusters join the same fleet grid, use the same 7-state model, and support the same six-level drill-down as clusters you shipped through Kubentic, with no separate setup.

How does drill-down actually work when a tile goes red?

You navigate from the fleet overview down through region, cluster, namespace, and node to a single pod, and every layer carries the same health vocabulary. A red tile at the fleet level always traces to a concrete offender at the bottom, a specific pod's restart timeline, a node's memory-pressure condition, or a PVC nearing capacity. Triage is a few clicks, not a CLI expedition.

Do I need Prometheus, Grafana, or extra agents to run this?

No separate observability stack to stand up. Live Monitoring pulls node, pod, and storage telemetry directly and renders it across AWS, GCP, Azure, and self-managed, no per-cloud console, no Grafana per region, no agents to babysit. If you also run Foresight AI, it consumes these same live signals to predict issues before they cascade.

We run three clouds and twenty clusters. Kubentic is the only tool that actually gives us a unified view without making us learn three different CLIs.

Anika Sharma

CTO · Developer Tools company

Your first environment is
15 minutes away.

No credit card. No infrastructure expertise. Just a cloud account and a browser, and you're shipping.

Free plan · No credit card · Cancel any time