Telemetry

Dewy provides built-in telemetry support based on OpenTelemetry (OTel) for monitoring proxy performance, deployment activity, and container health. Telemetry is particularly valuable in container mode, where Dewy operates as an independent TCP proxy rather than as part of the application process.

Architecture

In server mode, Dewy runs as part of the deployed application, so telemetry can be collected through the application's own OTel SDK (e.g., via otel-collector configured with systemd).

In container mode, Dewy operates as a standalone reverse proxy managing container lifecycle. Since it is separate from the application, Dewy needs its own telemetry pipeline. Dewy uses the OpenTelemetry SDK internally and supports two export paths:

  • Prometheus exporter: Exposes a /metrics endpoint on the Admin API for scraping
  • OTLP exporter: Sends metrics to an OpenTelemetry Collector via gRPC

Both exporters can be used simultaneously, allowing you to choose the best approach for your infrastructure.

┌──────────────────────────────────────────────────────┐
│  Dewy (container mode)                               │
│                                                      │
│  ┌──────────────┐   ┌────────────────────────────┐   │
│  │  TCP Proxy   │──▶│  OTel SDK (MeterProvider)  │   │
│  │  Deploy Mgr  │   │                            │   │
│  │  Health Check│   │  ┌──────────────────────┐  │   │
│  └──────────────┘   │  │ Prometheus Exporter  │──┼───┼──▶ GET /metrics (Admin API)
│                     │  └──────────────────────┘  │   │
│                     │  ┌──────────────────────┐  │   │
│                     │  │   OTLP Exporter      │──┼───┼──▶ OTel Collector (gRPC)
│                     │  └──────────────────────┘  │   │
│                     └────────────────────────────┘   │
└──────────────────────────────────────────────────────┘

Enabling Telemetry

Telemetry is disabled by default. Enable it with the --telemetry flag or by specifying an --otlp-endpoint.

Prometheus Only

Expose a /metrics endpoint on the Admin API server. Prometheus can then scrape this endpoint.

dewy container --telemetry \
  --registry img://ghcr.io/myorg/myapp \
  --port 8080 --health-path /health

The metrics endpoint is available at http://localhost:17539/metrics (the Admin API port).

Prometheus + OTLP

In addition to the Prometheus endpoint, send metrics to an OpenTelemetry Collector via gRPC.

dewy container --telemetry \
  --otlp-endpoint localhost:4317 \
  --registry img://ghcr.io/myorg/myapp \
  --port 8080 --health-path /health

When --otlp-endpoint is specified, telemetry is automatically enabled even without the --telemetry flag.

OTLP Only

If you only want OTLP export without the Prometheus endpoint, specify the endpoint:

dewy container --otlp-endpoint otel-collector.internal:4317 \
  --registry img://ghcr.io/myorg/myapp \
  --port 8080

Note: The Prometheus exporter is always registered when telemetry is enabled. The /metrics endpoint will be available regardless, but you can simply choose not to scrape it.

Metrics Reference

All metrics use the dewy. prefix and follow OpenTelemetry semantic conventions.

Proxy Metrics

These metrics are recorded for each TCP proxy connection and are labeled with proxy_port.

MetricTypeUnitDescription
dewy.proxy.connections.totalCounter{connection}Total number of proxy connections accepted
dewy.proxy.connections.activeUpDownCounter{connection}Number of currently active proxy connections
dewy.proxy.connection.durationHistogramsDuration of proxy connections (from accept to close)
dewy.proxy.connect.latencyHistogramsTime to establish connection to a backend
dewy.proxy.bytes.transferredCounterByTotal bytes transferred through the proxy (both directions)
dewy.proxy.errors.totalCounter{error}Total number of proxy errors (no backend, connection failures)
dewy.proxy.backendsUpDownCounter{backend}Number of active proxy backends

The dewy.proxy.connect.latency histogram uses the following bucket boundaries optimized for network latency measurement: 0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 5 seconds.

Deployment Metrics

MetricTypeUnitDescription
dewy.deployments.totalCounter{deployment}Total number of successful deployments
dewy.deployment.durationHistogramsDuration of the deployment process
dewy.deployment.errors.totalCounter{error}Total number of failed deployments

The dewy.deployment.duration histogram uses the following bucket boundaries: 1, 5, 10, 30, 60, 120, 300, 600 seconds.

Health Check Metrics

MetricTypeUnitDescription
dewy.healthchecks.totalCounter{check}Total number of health checks performed
dewy.healthchecks.failures.totalCounter{check}Total number of failed health checks

Container Metrics

In container mode, Dewy inspects the containers it manages on each scrape and reports their lifecycle state. These are the metrics that answer "did a container crash or restart?", analogous to what kube-state-metrics exposes for Kubernetes pods.

Per-container metrics — every one below except dewy.container.replicas and dewy.container.desired_replicas — are labeled with app, container (the container name), and replica (its replica index). The two replica-count metrics are app-wide and carry only the app label. dewy.container.info additionally carries image and version.

MetricTypeUnitDescription
dewy.container.replicasGauge{replica}Number of running container replicas
dewy.container.desired_replicasGauge{replica}Number of replicas Dewy is configured to run
dewy.container.restartsGauge{restart}Restart count reported by the runtime
dewy.container.statusGauge1 for the container's current state label, 0 otherwise
dewy.container.last_terminated.exit_codeGaugeExit code of a stopped container
dewy.container.oom_killedGauge1 if a stopped container was OOM-killed, else 0
dewy.container.started.timestampGaugesContainer start time (Unix seconds)
dewy.container.infoGaugeAlways 1; carries image and version labels

The state label on dewy.container.status takes one of created, running, paused, restarting, exited, dead. The exit-code and OOM metrics are reported only while a stopped container still exists; Dewy stops reporting a container that terminated more than one hour ago so old series do not accumulate.

dewy.container.restarts is a gauge, not a counter. It mirrors the runtime's own restart count, which resets to 0 when Dewy replaces the container on a deploy. Query it with changes() or delta(), not increase()/rate() (those assume a monotonic counter and would miss or misread the resets).

Enabling restarts

Dewy does not restart crashed containers itself — it observes and reports. To get automatic restarts (and a non-zero dewy.container.restarts), pass a restart policy through to the runtime after the -- separator:

dewy container --registry ... -- --restart=on-failure

Without a restart policy the container stays stopped after a crash, which shows up as dewy.container.replicas falling below dewy.container.desired_replicas.

kube-state-metrics equivalents

Dewy metrickube-state-metrics analogue
dewy.container.restartskube_pod_container_status_restarts_total
dewy.container.statuskube_pod_status_phase
dewy.container.oom_killed / last_terminated.exit_codekube_pod_container_status_last_terminated_reason
dewy.container.started.timestampkube_pod_start_time
dewy.container.replicaskube_deployment_status_replicas_available
dewy.container.desired_replicaskube_deployment_spec_replicas
dewy.container.infokube_pod_info

Integration Examples

Prometheus + Grafana

A typical setup with Prometheus scraping Dewy's metrics endpoint:

# prometheus.yml
scrape_configs:
  - job_name: 'dewy'
    scrape_interval: 15s
    static_configs:
      - targets: ['localhost:17539']

Useful PromQL queries:

# Request rate (connections per second)
rate(dewy_proxy_connections_total[5m])

# Active connections
dewy_proxy_connections_active

# P99 backend connection latency
histogram_quantile(0.99, rate(dewy_proxy_connect_latency_bucket[5m]))

# Deployment frequency (per hour)
increase(dewy_deployments_total[1h])

# Error rate
rate(dewy_proxy_errors_total[5m])

# Container restarts in the last hour (gauge: use changes, not increase)
changes(dewy_container_restarts[1h])

# Replica shortfall: containers Dewy wants minus what is actually running
dewy_container_desired_replicas - dewy_container_replicas

# Containers OOM-killed
dewy_container_oom_killed == 1

OpenTelemetry Collector

Send metrics to any OTel-compatible backend (Datadog, New Relic, Grafana Cloud, etc.):

# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

exporters:
  # Example: export to Prometheus remote write
  prometheusremotewrite:
    endpoint: "https://prometheus.example.com/api/v1/write"

  # Example: export to OTLP-compatible backend
  otlp:
    endpoint: "https://otel-ingest.example.com:4317"

service:
  pipelines:
    metrics:
      receivers: [otlp]
      exporters: [prometheusremotewrite]

systemd Integration

When running Dewy as a systemd service with telemetry:

# /etc/systemd/system/dewy.service
[Unit]
Description=Dewy Container Deployment
After=network.target docker.service

[Service]
Type=simple
ExecStart=/usr/local/bin/dewy container \
  --telemetry \
  --otlp-endpoint localhost:4317 \
  --registry img://ghcr.io/myorg/myapp \
  --port 8080 \
  --health-path /health \
  --replicas 2 \
  --log-format json \
  --log-level info
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target

CLI Options

OptionDescription
--telemetryEnable telemetry (Prometheus metrics on Admin API /metrics endpoint)
--otlp-endpointOTLP gRPC endpoint for exporting metrics (e.g., localhost:4317). Automatically enables telemetry.
--otlp-insecureUse insecure (plaintext) gRPC for OTLP export. Default is TLS. Use this for local or internal collectors without TLS.