Tail Latency & Autoscaling Simulator
Simulates a service under load: requests arrive as a Poisson stream (steady, diurnal, periodic spikes or one big burst), wait in a FIFO queue, and are handled by a pool of workers with your choice of service-time distribution — exponential, lognormal, or a bimodal mix where a few requests are much slower. A utilization slider sets ρ directly and moves λ to match, so you can walk the queue up to saturation without doing arithmetic. Live charts plot p50/p95/p99 over time next to queue depth, offered vs effective arrivals vs goodput, a latency histogram, and a utilization-vs-p99 curve that accumulates the famous hockey stick. Turn on client timeouts and retries to reproduce a retry storm, add request hedging and compare p99 with and without it at the same load, or start the autoscaler with a realistic scale-up delay and watch when the new workers actually arrive. Three story presets — Black Friday spike, retry storm, goodput collapse — set everything up for you.
Runs 100% in your browser — nothing you paste leaves your device.
Read the full guide to this tool
Notes
- Queueing theory in one line: waiting time scales like ρ/(1−ρ) with utilization ρ, so the step from 80% to 95% utilization costs far more latency than 50% to 80% did. Averages hide this; percentiles don't.
- Retries during overload are gasoline on a fire: every timed-out request becomes extra arrivals, pushing utilization further past 100%. A retry budget caps the amplification — watch the effective arrival rate and the amplification factor climb while goodput falls.
- Throughput is not goodput. A saturated service can keep completing requests at full capacity while almost every answer arrives after the client already gave up — completed work nobody is waiting for. The "wasted work" readout is the difference.
- The bimodal distribution shows why p99 matters: 5% of requests being much slower barely moves the mean but dominates the tail — and a fan-out calling 20 such services hits a slow one almost every time. Hedging (a duplicate request after a short delay, first answer wins) is the standard cure, and it costs real extra capacity: the duplicates queue like everything else.
- The autoscaler reacts to utilization but every new worker has to boot, so a sudden spike still hurts for the whole scale-up delay: capacity you must boot is not capacity you have. The cyan markers show when workers actually came online.
- Runs 100% in your browser — nothing you paste leaves your device.