← All articles

6 min read

Following a slow Worker request through Peren

A practical walk through Peren health checks, Prometheus metrics, admission signals and tail events.

Suppose an API that usually answers in 40 milliseconds begins taking half a second. The Worker bundle has not changed. The node is still running. Requests are still succeeding, but users can feel the delay.

Peren gives us three places to look: its health endpoints, Prometheus metrics and local tail events. Let’s follow them in that order and see how far they take us.

PEREN TELEMETRY SURFACEThree views, three different questions
01 · PROBES/healthz/readyz

Is the process alive, and should this listener receive traffic?

02 · METRICS/metrics

How much work is happening, where are errors rising, and is admission saturated?

03 · TAILperen tail

Which service requests completed, with what status, outcome and wall time?

FLEET CONTROLS

Logging levels, ownership hot-path logging, trace sampling, OTLP and logpush settings.

Start with the broad fleet signal, then move toward the individual service event.

First, is the node ready?

Every Peren listener exposes three operational routes:

  • /healthz says whether the process is alive;
  • /readyz says whether the listener should receive traffic; and
  • /metrics describes the work passing through the node.

Health and readiness answer different questions. A node can remain alive while it drains or while a listener is not ready for Worker traffic. In our case both checks pass, so removing the node from service would hide the symptom rather than explain it.

The metrics endpoint gives us one more shape check. peren_ready reports readiness by listener, and peren_listener_worker identifies listeners that dispatch Worker traffic. Peren also reports the number of configured listeners, services and queue consumers. Those values match the deployment we expected.

The process is alive, the listener is ready and the right services are loaded. We can move into the request path.

HTTP time gives us the first clue

Peren counts HTTP requests, HTTP errors and the total time spent serving HTTP traffic. The bundled Grafana dashboard turns the two relevant totals into a recent average:

rate(peren_http_duration_ms_total[5m])
/
clamp_min(rate(peren_http_requests_total[5m]), 1)

Request volume is steady, but average wall time rises. The error counter stays flat. We now know that the node is completing roughly the same amount of work more slowly.

This average is a clue, not a request trace. Peren exposes cumulative duration totals rather than latency histograms, so it cannot tell us the 95th percentile or point to the slowest request. We need another signal to narrow the cause.

Is the node full?

Before a request reaches a Worker, Peren admits it against the node’s concurrency budget. Three metrics describe that decision:

  • peren_admission_capacity is the available concurrency budget;
  • peren_admission_active is the top-level work using that budget; and
  • peren_admission_refused_total counts work refused because the ledger was full.

Active work remains below capacity and the refusal counter does not move. The slowdown did not begin because the node ran out of admission slots.

That matters because an admission refusal and a slow Worker need different responses. Adding capacity can help the first case. It tells us very little about the second.

Storage becomes the next suspect

The Worker writes durable state. Peren counts completed storage commits, storage rollbacks and the cumulative time spent committing. We compare commit time with commit volume using the same rate-over-rate calculation as HTTP.

Commit volume remains stable. Average commit time rises alongside HTTP wall time.

That does not prove that one particular commit delayed one particular request. These are node-level counters, and they carry no cell or request identifier. It does tell us where to look next. The slowdown follows durable storage work rather than admission pressure or a rise in HTTP failures.

Peren exposes the same kind of separation for other host operations. Queue sends, R2 calls, cache operations, AI calls, service fetches and Durable Object fetches have their own volume, error and duration counters. If the slow route read from R2 instead of committing storage, we would follow that series instead.

ILLUSTRATIVE INCIDENTA slow request, narrowed with node signals
START WITH THE NODEHTTP wall time rises

The request count is steady, but total HTTP duration is growing faster. The node is slower; the metric does not yet say why.

rate(peren_http_duration_ms_total[5m]) / rate(peren_http_requests_total[5m])
What this evidence cannot proveAn aggregate average cannot identify one request or reveal tail latency.
Move through the investigation. Each step narrows the cause without treating an aggregate as a request trace.

Tail gives us a real request

Metrics showed the pattern. Tail events make it concrete.

For each recorded service request, Peren stores the service, method, path, status, outcome and wall time. peren tail reads those events for a service and can filter them by level. A warning filter includes status codes from 400 upward; an error filter starts at 500.

The slow requests in our example are successful, so an error-only filter would miss them. Reading the service events shows repeated POST requests to the affected route with wall times close to the spike in the dashboard. We now have a route to reproduce and a time window to compare with storage and provider diagnostics.

Tail records live as JSON lines beneath the node’s data directory. They are useful for local review, but they are not a distributed trace or a long-term event store. A tail record does not join the request to each binding operation across several nodes.

Keep identities out of metrics

It would be tempting to add the route, tenant and cell ID to every metric. That would make this investigation easier for a few minutes and make Prometheus harder to operate for much longer.

A Peren fleet can execute work for many routes and a much larger number of cells. Turning those values into labels creates a new time series for every distinct combination. The cost grows with application traffic rather than with the small set of operational dimensions an operator expects to manage.

Peren keeps its metrics bounded. The endpoint does not expose Worker payloads, tenant data, provider endpoints, filesystem paths or secrets. Detailed request identity belongs in event and trace data, where high-cardinality values can be searched without reshaping the metric store.

Where tracing fits

Fleet configuration includes separate levels for internal logs and Worker console output, an opt-in ownership hot-path setting and a trace sampling ratio validated between zero and one. It also defines settings for OTLP and logpush exporters.

Peren does not yet emit the distributed request trace we would want for this incident. There is no span sequence that connects dispatch, Worker execution, a storage commit and the response under one trace ID. For now, the investigation moves from aggregate metrics to local tail events, then into storage or provider diagnostics.

That missing connection gives tracing a clear job. A useful request trace should show where time was spent, retain the service and node context needed for an investigation, and keep secrets and Worker payloads out of exported attributes.

What the signals tell us

We started with a vague report: the API feels slow. Readiness showed that the node was still eligible for traffic. HTTP metrics confirmed a latency change without an error spike. Admission metrics ruled out a full concurrency ledger. Storage totals narrowed the slowdown to commit work. Tail events gave us the service and route to reproduce.

No single signal answered the whole question. Together they reduced a fleet-wide symptom to a concrete request and a likely subsystem without claiming more precision than the data contained.

Read the observability guide for the endpoint and dashboard contract, or the peren tail reference for event review.