Zero-config agent · Kubernetes + Docker

Monitor every container
from one dashboard

Deploy a lightweight agent in 60 seconds. No cloud credentials required. Get full visibility into your infrastructure, CPU, memory, logs, and alerts, in minutes.

app.kubewatchlabs.com / overview
Live
Cluster throughput live
48.2k-10.1%req/min
Requests Errors
Resource usage6 nodes
CPU34%
Memory61%
Disk78%
Latency by servicep95 · ms
authquerylivealertbilltenant
Active alerts2 firing
API latency p99 > 800msgateway2m
Memory pressure on node-3kubelet14m
Pod restart loop clearedbilling1h
Disk usage back under 80%postgres3h
Containers128 total
gateway-7f4c
running
auth-9b2a
running
query-3d8e
running
billing-1a5f
stopped
Cloud spendthis month
$3,2840.5%
60s
To first metric
< 0.5%
Agent CPU overhead
99.9%
Uptime SLA

Integrates with the tools you already run

PostgreSQL
PostgreSQL
MySQL
MySQL
Redis
Redis
Kafka
Kafka
Argo CD
Argo CD
Jenkins
Jenkins
Prometheus
Prometheus
Grafana
Grafana
Capabilities

Everything you need to stay on top of your containers

From single Docker hosts to multi-cluster Kubernetes environments.

Docker Monitoring

Full visibility into every container, CPU, memory, network I/O, live log streaming, and restart tracking.

  • Container list & status
  • CPU & memory metrics
  • Network I/O per container
  • Live log tail
  • Image digests & volume mounts
  • Interactive exec & volume browsing (Enterprise)

Kubernetes Monitoring

Understand your cluster at a glance. Connect a cluster in one click or bring your own kubeconfig, then track pods, nodes, services, and per-node resource usage without leaving your browser.

  • Auto-discover EKS, GKE & AKS clusters
  • Per-node CPU, memory, network & disk IOPS
  • API server request rate
  • Classic Ingress and Gateway API side by side
  • Read-write cluster tools with self-service permission toggles (Enterprise)

Cluster & Host Backups

Back up the actual infrastructure your agents watch, not just KubeWatch itself: Kubernetes object manifests by namespace or label selector, or named volumes and bind mounts on a Docker host.

  • Scheduled or on-demand, per agent
  • S3, SFTP, webhook & Azure Blob destinations
  • Per-namespace / per-mount result history
  • Bytes stream straight from the agent, never through our servers

Node Cordon & Drain

Move workloads off a node before maintenance, a resize, or decommissioning it, respecting every PodDisruptionBudget you have configured.

  • Cordon, drain & uncordon from the Nodes page
  • PDB-aware eviction with automatic backoff
  • DaemonSet pods & emptyDir data called out before you approve
  • Admin-approved before anything is applied

Vulnerability Scanning

Scan every container image for known CVEs with your choice of engine, Trivy, Grype, or Snyk, surfaced right on the image detail page next to its digest and size.

  • Trivy, Grype & Snyk support
  • Severity-ranked CVE findings
  • Scan on demand, cached per image digest
  • Reverse lookup: every affected container & pod

Real-time Alerts

Know before your users do. Set threshold rules, route to Slack or email, and silence noisy alerts intelligently.

  • Threshold-based rules
  • Slack & email channels
  • Silence detection
  • Alert history & audit log

Incident Management

Detect, respond, resolve, and learn, all in one place. Escalate a firing alert into an incident, work it with your team, and capture a postmortem when it is done.

  • On-call scheduling & dynamic routing
  • Timeline, responders & related incidents
  • Postmortems with action items
  • MTTA / MTTR insights

Auto Scaling

Close the loop from watching to acting. Scale Kubernetes pods and Docker containers on the metrics you already track, from one policy.

  • Native HPA & Karpenter on K8s
  • Orchestrated scaling on Docker
  • Dry-run, cooldowns & approvals
  • One-click rollback

Integrations

Monitor the services around your apps. Connect databases, caches, message queues, CI/CD, and observability tools with live health and deep metrics.

  • Postgres, MySQL, Redis, Kafka
  • Argo CD, Jenkins, Prometheus, Grafana
  • Latency & uptime health checks
  • Per-service deep metrics

Performance & Load Testing

Run performance, load, stress, soak, and spike tests against any endpoint and read real latency percentiles, throughput, and a full success and error breakdown.

  • 5 test types: load, stress, soak, spike & more
  • P50 / P90 / P99 / max latency
  • Breaking point & recovery time
  • Uptime & latency-drift tracking

Kafka Topics Monitoring

Watch topic size and consumer group health without a separate collector. The agent talks directly to your Kafka cluster's admin API, filterable by cluster, topic, and consumer group.

  • Topic partition offsets over time
  • Consumer group total & per-partition lag
  • Committed offset tracking
  • No consuming, no separate service

AI Agent Cost & Reliability Ops

Bring your AI agents (LangChain, CrewAI, custom OTel-instrumented systems) into KubeWatch as a first-class monitored entity, with the same alerting your infrastructure already gets.

  • Per-session cost, tokens & tool calls
  • Waterfall trace view per session
  • Cost, error-rate & runaway-session policy rules
  • Dry-run policy evaluation with a full decision log

Docker-vs-Kubernetes Advisor

A workload-level recommendation, backed by its own historical usage, of whether it needs Kubernetes or would run just as well as a KubeWatch-managed Docker replica.

  • CPU volatility & restart-pattern scoring
  • Structured, readable rationale per recommendation
  • Conservative savings estimates
  • Advisory only, nothing applied automatically

GPU & Inference Cost Control

The same autoscaler policy engine, made GPU- and cost-aware. Live GPU utilization, cost-per-1k-tokens attribution, and GPU-aware scaling policies.

  • Real-time GPU utilization, memory & power
  • Cost-per-1k-tokens attribution
  • Scale-to-zero, rightsize & spot-shift policies
  • Same dry-run & rollback conventions

Self-Hosted AI Stack Observability

Treat vLLM, Ollama, SGLang, and TGI as first-class monitored workloads, with a Model Registry and a Sovereign Fallback Health Checker that proves your fallback path actually works.

  • vLLM / Ollama / SGLang / TGI metrics
  • Model Registry by role & version
  • Sovereign Fallback Health Checker
  • Agent-side probing, read-only

Autonomous Remediation

Attach an automated restart to any alert rule you already have. Every playbook starts in forced dry-run, is rate-limited by a circuit breaker, and defaults to requiring human approval.

  • Docker container restart & K8s rolling restart
  • Forced dry-run until you promote it
  • Per-playbook circuit breaker
  • Full append-only audit trail

Edge AI Fleet Observability

Bring cameras, sensors, and industrial controllers into the same fleet view as your containers and nodes, over the same long-poll channel built for the autoscaler.

  • Store-and-forward buffering for spotty links
  • Three-state connectivity: online / expected / unexpected
  • Per-device offline windows, not one global grace period
  • OTA model updates via your own apply script

Infrastructure Automation

A GitOps control plane for OpenTofu and Ansible: sync Projects from Git, run Stacks through a plan-then-approve-then-apply workflow, and inject secrets from an encrypted vault, never plaintext, never committed to your repo.

  • Git-synced Projects & Stacks, OpenTofu or Ansible
  • Mandatory human-approved dry run before any real change
  • AES-256 encrypted secrets vault, including SSH keys
  • 50 built-in blueprints to scaffold a new Stack straight to Git
  • Project-scoped roles & a full audit trail

DataOps

Orchestrate and monitor your data pipelines and cloud infrastructure, all from one place. Trigger runs across AWS Glue/Step Functions, Airflow, Azure Data Factory, and GCP Dataflow/Composer, then start, stop, resize, or delete databases, storage, and serverless functions across AWS, Azure, and GCP, without leaving KubeWatch.

  • Trigger, cancel & monitor runs across 4 ETL tool families
  • Automatic run history, refreshed roughly every 20 seconds
  • Database, storage & serverless inventory across AWS, Azure & GCP
  • Start, stop, resize & delete resources with a confirm step
  • Admin-only, available on every plan

Analytics

Connect your own Postgres, MySQL, Redshift, or S3 data, clean it with a pipeline that runs natively inside KubeWatch, and chart the result right on the dashboard, no external BI tool needed.

  • Postgres, MySQL, Redshift & S3 connections
  • Encrypted-at-rest credentials
  • Dedupe, fill/drop missing, coerce, filter & custom-transform steps
  • Manual or recurring-interval runs
  • Line, bar & area charts from a cleaned dataset
  • Admin-only, available on every plan
Metrics

Real-time metrics with deep history

Every container and node streams CPU, memory, network, and request throughput to high-resolution time-series charts. Zoom from the last minute to months back without sampling gaps.

  • Sub-second collection interval
  • Per-service request & error rates
  • Powered by VictoriaMetrics
Cluster throughput live
48.2k+12.4%req/min
Requests Errors
Alerting

Catch incidents before your users do

Define threshold rules on any metric, route them to Slack or email, and let intelligent silencing cut the noise. A full audit trail shows every fire and resolution.

  • Threshold & anomaly rules
  • Slack and email channels
  • Acknowledge & audit history
Active alerts2 firing
API latency p99 > 800msgateway2m
Memory pressure on node-3kubelet14m
Pod restart loop clearedbilling1h
Disk usage back under 80%postgres3h
Incident Management

Detect, respond, resolve, and learn

Escalate a firing alert into an incident, and KubeWatch pages whoever is currently on call while your team works it, timeline, responders, and related past incidents all on one page. Resolve it, then capture a postmortem with tracked action items.

  • On-call scheduling & dynamic routing
  • Timeline, responders & related incidents
  • Postmortems with action items
  • MTTA / MTTR insights
Checkout API degraded critical
Declared
Acknowledged
Mitigated
Resolved
Responders2 · on-call paged
Related incidents1 match
Time to acknowledge4m
AI Observability

Monitor AI and ML workloads, by cost and quality

Track every model call: tokens, latency, error rate, and spend, with a built-in price table for OpenAI, Anthropic, and more. Watch GPU utilization and scrape vLLM, Triton, and KServe inference servers from the same agent.

  • Cost and token tracking per model
  • GPU utilization, memory, and power
  • vLLM / Triton / KServe metrics
AI observability live
Spend
$1,284
Tokens
4.2M
p95
1.3s
gpt-4o18.4k$6120.3% err
claude-sonnet-49.1k$4710.6% err
text-embedding-352k$2010.1% err
Requests & Latency

Request rates, error rates, and latency everywhere

Report application API requests to see throughput, error rate, and p95 latency per route. Synthetic probes measure latency to your services, nodes, and the agent itself, with uptime tracking.

  • Per-route request and error rates
  • p50 / p95 latency breakdowns
  • TCP and HTTP latency probes
API requestslast 24h
48.2kreq/min0.6% error rate
GET /v1/orders38ms0.2%
POST /v1/pay210ms2.1%
GET /v1/users24ms0%
Latency probesagent · node · probe
agent pushagent8 ms
node-1node1.2 ms
node-2node1.4 ms
api.svc/healthprobe42 ms
postgres:5432probe3 ms
vllm:8000probedown
OpenTelemetry

Bring your own telemetry with OTLP

Point any OpenTelemetry SDK or Collector straight at KubeWatch. We ingest traces, metrics, and logs over OTLP/HTTP, decoding both Protobuf and JSON, so your existing instrumentation works with no rewrites and no vendor lock-in.

  • Traces, metrics, and logs over OTLP/HTTP
  • Protobuf and JSON encoding
  • Drop-in for OpenTelemetry SDKs & Collector
  • Open standard, no vendor lock-in
OpenTelemetryOTLP/HTTP
Traces
2.4k
Metrics
38k
Logs
12k
distributed trace · 5 spans312 ms
GET /v1/checkout
312 ms
auth.verify
41 ms
db.query orders
96 ms
payment.charge
108 ms
cache.set
22 ms
protobufJSONgzip ingesting
Auto Scaling

From watching your workloads to scaling them

Set a per-workload policy and KubeWatch acts on the same metrics you already see. On Kubernetes it writes native HorizontalPodAutoscaler and Karpenter objects and lets the cluster execute them. On standalone Docker, where there is no HPA, KubeWatch is the orchestrator: it picks placement, scales containers, and routes traffic through a managed load balancer.

  • Pods on Kubernetes, containers on Docker, one policy
  • Dry-run first, then go live with asymmetric cooldowns
  • Approval gates and an append-only decision log
  • One-click rollback on either runtime
Auto Scaling live
Replicas
3 → 4
CPU target
70%
Bounds
2 to 10
k8scheckout-api · scale up 3→4cpu 84% > 70%
dockerweb · placed on host west-162% headroom
k8sworkers · steadywithin target
Integrations

Watch the services your apps depend on

Your containers are only half the picture. Connect the databases, caches, message queues, CI/CD, and observability tools around them, and KubeWatch tracks their health, latency, and uptime, then pulls deep per-service metrics like connection pools, cache hit rates, and replication lag.

  • Postgres, MySQL, Redis, and Kafka
  • Argo CD, Jenkins, Prometheus, Grafana, and more
  • Continuous latency and uptime health checks
  • Deep metrics, not just up or down
Integrationsdatabases · caches · CI/CD
postgres-prodPostgreSQL2 ms
redis-cacheRedis1 ms
kafka-eventsKafka8 ms
mysql-billingMySQL4 ms
argocdArgo CD41 ms
prometheusPrometheusdown
Performance Testing

Know how an endpoint holds up before your users do

Pick a test type from the dropdown, point it at any URL, and KubeWatch runs it and reports the numbers that matter for that scenario, everyday latency percentiles for a load test, the breaking point and recovery time for a stress test, latency drift and uptime for a long-running soak, or how the target handles a sudden burst in a spike test.

  • Load & Performance: latency percentiles, throughput, error rate
  • Stress: max load capacity, failure points, recovery time
  • Soak: uptime % and latency drift over long runs
  • Spike: response during bursts, error handling, recovery
Performance testStress
GET /v1/checkout
Req/sec
48.2k
Avg
34 ms
Errors
0.4%
P50
28 ms
P90
96 ms
P99
180 ms
Max
240 ms
Kafka Topics

Catch a consumer falling behind, before it pages you

Point the agent at your Kafka cluster's admin API and KubeWatch tracks topic size and consumer group lag right on the Overview page. Filter by cluster, topic, or consumer group to drill into exactly the partition that's falling behind.

  • Topic partition offsets over time
  • Consumer group total & per-partition lag
  • Committed offset tracking
  • Direct admin API connection, no separate collector
Kafka Topics live
orders.created
payments.captured
shipping.updates
user.events
checkout-service lag124
Infrastructure

Pods, nodes, containers, and networks in one view

Full visibility across Docker and Kubernetes, from a single node to multi-cluster fleets.

Pods6 namespaces
Running142
Pending3
Failed1
Nodes4 ready
node-1
node-2
node-3
node-4
CPU Memory
Networksrx / tx
1.8 GB/s 920 MB/s
bridgebridge12 attached
kube-overlayoverlay34 attached
hosthost3 attached
Containers128 total
gateway-7f4c
running
auth-9b2a
running
query-3d8e
running
billing-1a5f
stopped
Latency probesagent · node · probe
agent pushagent8 ms
node-1node1.2 ms
node-2node1.4 ms
api.svc/healthprobe42 ms
postgres:5432probe3 ms
vllm:8000probedown
Setup

Up and running in minutes

No complex setup. No cloud IAM roles. Just deploy and watch.

01

Sign up in 30 seconds

Create your account with just your email. No credit card, no sales call, no waiting.

02

Deploy the agent with one command

Run a single Docker or Helm command on your host. The agent securely streams metrics to your dashboard.

03

See your infrastructure in under 5 minutes

Within minutes you have a live view of every container and node, CPU, memory, logs, and more.

Deployment

Two ways to run KubeWatch

Choose the deployment model that fits your team.

Hosted SaaS

30-day trial, no card required. We manage the infrastructure.

  • Agent-only deployment
  • We manage the backend
  • No credit card to start
  • Scale with your team
Start trial

Self-Hosted

Pro or Enterprise. Your data stays in your infrastructure.

  • Full platform deployment
  • Data residency & compliance
  • Docker Compose or Helm install
  • Annual license with updates
Learn more
Pricing

Simple, transparent pricing

Start with a 30-day trial. Choose Pro or Enterprise before it ends.

Most popular

Pro

$69/mo or $799/yr
  • 50 agents
  • 10 users
  • 50 alert rules
Start trial

Enterprise

$299/mo or $2,990/yr
  • 200 agents
  • Unlimited users
  • Unlimited alert rules
Start trial
AI-Ops

Your AI workloads get the same rigor as everything else

Model calls, GPUs, and inference servers are just another part of the stack KubeWatch already watches, not a separate product bolted on.

  • AI Agent Cost & Reliability Ops: Policy rules catch a runaway model call by cost or error rate as it happens
  • Docker-vs-Kubernetes Advisor: A data-backed read on which workloads are paying for orchestration they don’t use
  • GPU & Inference Cost Control: Utilization, memory, and spend across vLLM, Triton, and KServe
  • Edge AI Fleet Observability: The same agent, watching inference at the edge
  • Self-Hosted AI Stack Observability: For teams running their own models instead of calling out to an API
  • Autonomous Remediation: Dry-run reviewed before it acts, and it grades its own track record: a playbook that stops actually fixing things automatically falls back to human review
  • AI Log Diagnostics: Every alert gets a plain-English root cause for free, no API key required, and links straight back to the alert that triggered it

The first four are on the Pro plan; Self-Hosted AI Stack Observability and Autonomous Remediation are Enterprise. AI Log Diagnostics ships free on every plan and every self-hosted install, using a small model bundled in, with no API key and no per-token cost; bring your own OpenAI, Anthropic, or self-hosted key instead if you want a higher-accuracy provider for harder cases.

See how each one works
30-day Enterprise trial · No credit card

See your whole stack in under 5 minutes

Join engineering teams who replaced their scattered dashboards with one place for Docker and Kubernetes.