1. The Problem
In a production environment, knowing that an application is deployed isn't enough — you also need to know whether the devices and services running at the edge are actually healthy. In my case, that meant thousands of edge devices deployed across many customer locations, each running local software that depends on network connectivity, hardware health, and correct configuration.
A device can silently go offline, lose connectivity to its local network, run low on storage, or fall out of compliance with the fleet's expected OS version — and none of that is visible unless something is actively watching for it. We needed a way to answer, at any moment, for any location: "Is this device alive, and is it healthy?"
2. Requirements
- Every device should report its liveness on a predictable cadence (a heartbeat), scoped to a specific location.
- The ingestion path must be resilient to bursts, retries, and duplicate/late messages without losing data.
- Failed or malformed messages must not be silently dropped — they need a dead-letter path for investigation.
- In addition to "is it alive," we wanted richer telemetry (CPU, memory, storage, battery, compliance) without requiring every device to push it directly.
- All ingested data needed to land in a time-series store so it could be queried, graphed, and alerted on.
- Operators needed dashboards for fleet-wide and per-location visibility, plus alerts when something goes wrong — without needing to read logs.
3. Architecture
End-to-End System Diagram

Two ingestion paths feed the same metrics backend:
- Direct heartbeats — each edge device pushes a small heartbeat payload on its own cadence.
- Scheduled telemetry pull — a Lambda polls an external device management API on a fixed schedule to enrich the picture with hardware and compliance data the devices don't push themselves.
Both paths converge on the same time-series backend, so a single set of dashboards and alerts covers "is it alive" and "is it healthy."
4. AWS Components
| Component | Role |
|---|---|
| API Gateway + Lambda authorizer | Authenticates and routes incoming heartbeat requests, scoped per location |
| Lambda (enqueue) | Validates the heartbeat payload and pushes it onto SQS |
| SQS + DLQ | Buffers and batches heartbeats; isolates malformed/failed messages after repeated delivery failures instead of dropping them |
| Lambda (dequeue) | Consumes SQS batches, converts each heartbeat into Prometheus-format metrics, and writes them to the metrics backend |
| Lambda (scheduled telemetry sync) | Runs on a fixed EventBridge schedule, pulls device and management telemetry from an external device-management API, and writes it as metrics |
| S3 | Stores a raw, timestamped backup of every metrics payload as an audit trail, with lifecycle expiry |
| SSM Parameter Store | Holds service credentials and sync-state (e.g. last successful poll timestamp) outside of source code |
| CloudWatch Alarms + SNS | Watches queue age, DLQ depth, and Lambda error rates — the "is the pipeline itself healthy" layer, separate from device health |
Keeping the ingestion pipeline serverless meant we didn't need to run or patch any always-on infrastructure just to accept heartbeats — it scales automatically with the size of the fleet and the queue absorbs bursts.
5. Heartbeat Design
The heartbeat payload is intentionally small and cheap to send. At minimum, a device only needs to identify itself and increment a counter:
{
"serialNumber": "XXXXXXXXXXXXXXXXXXXXXXXXXX",
"count": 14
}
Optionally, a device can include local network connectivity state, since a healthy device that's lost its local network connection is a different failure mode from an offline device:
{
"serialNumber": "XXXXXXXXXXXXXXXXXXXXXXXXXX",
"count": 14,
"isConnected": true,
"lastPublished": "2025-07-29T20:12:26Z",
"connectionType": "local",
"messageCount": 5
}
A few deliberate design choices here:
- The device serial number is validated against the authenticated caller, not trusted blindly — a device can't report a heartbeat under an identity it isn't authorized for.
- The schema is strict (
additionalProperties: false) — unexpected fields fail validation loudly instead of being silently ignored. - The payload has no dependency on any specific transport; it's just JSON over HTTPS, which keeps the device-side implementation trivial.
6. Metrics and Monitoring
Rather than storing heartbeats as raw events in a database, every heartbeat is immediately converted into Prometheus-format gauges at ingestion time. That decision shaped everything downstream: querying "which devices haven't reported in the last N minutes" becomes a standard PromQL query instead of a custom aggregation job.
Two families of metrics exist:
- Liveness metrics (from the heartbeat path): heartbeat counter, network connection status, last publish time, message count — labeled by device serial number.
- Deep telemetry metrics (from the scheduled sync path): CPU utilization and temperature, memory and storage capacity, battery health, network diagnostics, OS version and compliance status, enrollment state.
Because both are just labeled Prometheus gauges in the same store, a dashboard or alert doesn't need to care which pipeline produced a given metric.
7. Prometheus / VictoriaMetrics
We chose VictoriaMetrics as the metrics backend because it speaks the Prometheus remote-write/import protocol (so nothing upstream needs to know it isn't "real" Prometheus) while being cheaper to run at scale for a high cardinality of devices and locations.
The dequeue Lambda batches metrics and writes them via VictoriaMetrics' /api/v1/import/prometheus endpoint, gzip-compressed, over HTTPS with basic auth. The scheduled telemetry sync does the same, but chunks its output into sub-1.5MB batches (to stay under payload limits) with retry and exponential backoff, since a single sync run can cover the entire device fleet.
Locally, the same VictoriaMetrics image runs in Docker so the ingestion path can be tested end-to-end without touching production metrics.
8. Grafana Dashboards
VictoriaMetrics is queried by Grafana for both fleet-wide and per-location visibility:
- A fleet overview dashboard shows aggregate liveness across all locations/devices — how many are currently reporting heartbeats, how many have gone silent, and overall compliance/health trends.
- Per-location dashboards drill into a single location's devices — heartbeat recency, network connectivity, hardware telemetry (CPU/memory/storage/battery) — so an operator investigating a specific customer complaint can go straight to the relevant view.
Because everything is labeled consistently by device serial number, the same panels/queries generalize across dashboards just by changing the label filter.
9. Alerting
Alerting is defined as Grafana alert rules evaluated directly against the VictoriaMetrics data — for example, a device with no heartbeat metric update within an expected window, or a rising rate of network disconnections at a location. This keeps alerting logic close to the same PromQL used for the dashboards, rather than as a separate system with its own definition of "healthy."
This is layered with CloudWatch alarms that watch the pipeline itself rather than the devices — SQS queue age, DLQ message visibility, and Lambda error rates — because a healthy-looking dashboard doesn't help if the ingestion pipeline is actually the thing that's broken.
10. Challenges and Trade-offs
- Two ingestion paths, one data model. Direct heartbeats and polled telemetry arrive on completely different cadences and shapes, but both had to converge into the same labeled metric format so dashboards and alerts didn't need to special-case the source.
- Batching vs. payload limits. The scheduled telemetry sync can produce a large volume of metrics in one run across the whole fleet; naively posting them in one request risks hitting the metrics backend's payload limits, so batching with retry/backoff was necessary.
- Trusting device-reported data. A device claiming a given serial number can't be taken at face value — it has to be checked against what the caller is actually authorized for, otherwise one compromised or misconfigured device could pollute another device's metrics.
- Losing data safely. Anything that fails to enqueue or process shouldn't just vanish — the DLQ and the raw S3 backup of every metrics payload exist specifically so a bad batch can be replayed or inspected after the fact.
- Two definitions of "healthy." Device health (via Grafana/VictoriaMetrics) and pipeline health (via CloudWatch/SNS) are deliberately separate concerns, alerting through different channels, because conflating them makes it harder to tell "the fleet is unhealthy" from "our monitoring is broken."
11. What We Learned
- Converting data into Prometheus format as early as possible in the pipeline (at ingestion, not at query time) simplified everything downstream — dashboards, alerts, and ad-hoc debugging all speak the same query language.
- Treating the metrics pipeline itself as a first-class thing to monitor (via CloudWatch) is just as important as monitoring the devices it reports on.
- Keeping the heartbeat payload minimal paid off — it made the device-side integration trivial and kept the ingestion Lambda's job simple and fast.
- A raw payload backup (S3) turned out to be cheap insurance — being able to replay or inspect exactly what was received, even non-fatally, saved investigation time more than once.
12. Possible Improvements
- Push more of the "deep telemetry" fields into liveness-style alerting (e.g. combining "hasn't heartbeated" with "was last seen non-compliant") for higher-signal alerts.
- Explore anomaly-detection style alerting rather than fixed thresholds for fleet-wide trends.
- Reduce the polling cadence gap between direct heartbeats and scheduled telemetry sync so both liveness and deep health data are closer to real-time.
13. Conclusion
Observability at the edge isn't just about collecting data — it's about designing the smallest possible signal (a heartbeat) that reliably answers "is this thing alive," and complementing it with richer telemetry pulled on our own terms rather than depending on every device to push everything. By normalizing both into the same time-series format early, we got a single, consistent place — Grafana on top of VictoriaMetrics — to answer both "is the fleet healthy" and "is our monitoring pipeline itself healthy."

Top comments (0)