Alerting baselines
Setting up alerts on key metrics helps detect issues before they impact users. The following are suggested baselines for common Akka components.
| These thresholds are starting points — tune them based on your workload characteristics and operational experience. |
| For a full list of available metrics and their attributes, see the Telemetry metrics reference. To set up metric exports to your monitoring system, see Exporting metrics, logs, and traces. |
HTTP Endpoints
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Failing requests rate |
|
> 1 req/s |
Processing time p99 |
|
> 200 ms |
| For streaming HTTP endpoints, the duration metric measures the full stream lifetime, so the p99 threshold above may not be meaningful. Focus on the failing requests rate for streaming endpoints instead. |
gRPC Endpoints
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Failing requests rate |
|
> 1 req/s |
Call duration p99 |
|
> 200 ms |
| For gRPC streaming calls, the duration metric measures the full stream lifetime, so the p99 threshold above may not be meaningful. Focus on the failing requests rate for streaming endpoints instead. |
Agent components
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Commands failed rate |
|
> 1 cmd/s |
Tools failed rate |
|
> 1 cmd/s |
Content load duration p99 |
|
> 2 s (depends on content source) |
Event Sourced Entities
Event Sourced Entity metrics use akka.component.type="event-sourced-entity" to distinguish them from Key Value Entity metrics, which share the same metric names.
|
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Persist time p99 |
|
> 100 ms (average over 60s) |
Processing time p99 |
|
> 100 ms (average over 60s) |
| For Event Sourced Entities, occasional spikes in persist or processing time are expected (for example, during snapshot creation). Using a 60-second averaging window helps filter out these transient peaks. |
Key Value Entities
Key Value Entity metrics use akka.component.type="key-value-entity" to distinguish them from Event Sourced Entity metrics, which share the same metric names.
|
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Persist time p99 |
|
> 100 ms (average over 60s) |
Processing time p99 |
|
> 100 ms (average over 60s) |
Workflows
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Commands failed rate |
|
> 1 req/s |
Steps failed rate |
|
> 1 req/s |
Active workflow count |
|
Set a reasonable max for your workload (sustained counts above the expected maximum may indicate stuck workflows) |
Recovery duration p99 |
|
> 1 s |
Views
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Query duration p99 |
|
> 500 ms |
Update duration p99 |
|
> 500 ms |
Consumers
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Events consumption lag (max) |
|
> 1 min |
Events processing time (avg) |
|
> 1 s |
Subscriptions failed |
|
> 0 |
Replication
| These metrics are only applicable to multi-region deployments. |
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Replication lag (avg) |
|
> 10 s |
Timers and Timed Actions
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Timer call failures |
|
> 0 |
Timed action failures |
|
> 0 |
JVM runtime
Every Akka service exports the standard OpenTelemetry JVM metrics. The alerts below catch resource leaks and saturation that component metrics do not show.
| Alert condition | Metric | Suggested threshold |
|---|---|---|
Thread count |
|
> 500 |
GC time |
|
> 6 s per minute (10% of the time spent in GC pauses) |
CPU utilization |
|
> 0.9 for 5 min |
| A steadily rising thread count usually means a resource is created per request and never closed, for example an HTTP client or an executor. Use the SDK-provided HTTP client instead of creating your own. |
CPU exhaustion does not need immediate action. Akka autoscales the service when CPU usage exceeds the configured cpuUsageThreshold, up to maxInstances. Alert on CPU to know when the service runs at its maximum bound. See ServiceAutoscaling.
|