Tail Latency

Turkish equivalent: Üst yüzdelik gecikmeDomain: Performance Engineering

Tail Latency — Latency at the slow end of a distribution, commonly summarized with percentiles such as p95, p99 or p99.9.

Why the Tail Matters

An average can remain low while a small fraction of requests take much longer. In fan-out systems, the probability that at least one dependency is slow increases as more parallel calls are required.

Tail behavior therefore becomes a reliability concern, not merely a statistical detail.

Common Sources

Queueing near saturation, GC, storage stalls, lock contention, retries and heterogeneous work sizes can all widen the tail.

Percentiles need a stated window and measurement boundary. Averaging percentiles from independent instances is generally not equivalent to merging their underlying distributions.

See High-Performance Java Data Systems and Queueing Theory.

Related technical publications

Publications whose title or summary directly references this concept.

The Real Cost of Lock-Free Queues

A lock-free queue can still lose on tail latency when CAS contention, cache-line ownership, allocation and backpressure are ignored.