NATS JetStream configuration errors cause silent delivery failures
Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation
Networking and Internet Architecture
Summary
NATS JetStream promises to deliver messages at least once, but some common setup mistakes break this guarantee without any clear warning, leading to message loss, duplicates, or repeated redelivery floods. The authors found that existing monitoring tools cannot detect these problems. They created nats-lens, a tool that watches the system’s internal status and spots all these errors quickly and accurately, without changing how users build or run their apps. This works across multiple programming languages and presents results in several easy-to-use formats.
What this means in practice
- •For cloud infrastructure teams: Detect and fix silent message delivery errors in NATS JetStream clusters to ensure reliable service operations.
- •For distributed application developers: Use nats-lens to monitor message delivery correctness without modifying client code across Rust, Go, and Python applications.
Authors
Biplab Kumar Das
Abstract
NATS JetStream's at-least-once delivery guarantee is conditional: five common configuration mistakes silently violate it, causing duplicate message processing, data loss, or redelivery storms with no error logged anywhere. The standard Prometheus NATS exporter exposes only server-level throughput metrics and cannot detect any of these failures. We present nats-lens, a standalone monitor that reads from the JetStream management API and detects all five violation classes without requiring changes to monitored applications or client code. We formally characterize each class with a precise condition, prove that standard Prometheus NATS metrics are structurally incapable of detecting any of them, and implement five targeted detectors. In a controlled evaluation of 30 rounds per scenario on both single-node and 3-node JetStream clusters, nats-lens achieves 100% detection coverage across all five classes---versus 0% for the baseline---with zero false positives over 30 minutes of healthy operation. Detection latency ranges from 2,003 ms to 8,013 ms (within three poll cycles). We confirm language-agnostic detection empirically using consumers in Rust, Go, and Python. The tool is open source and exposes findings through four output channels: web dashboard, Prometheus metrics, REST API, and NATS health events.