Stackrig

Blog

Tech guides and explanations: servers, databases and the cloud. How things work and what they cost, told plainly, with our own measurements on cloud machines where numbers decide and designs you can run in Stackrig. Every piece states its setup and what it does not tell you.

measured our run on real machines model the simulator explained a concept, with one simulated example

Newest explained

Why does fan-out make fast services slow?

A request that waits for 100 parallel calls is as slow as the slowest of them: the arithmetic of Dean and Barroso's paper, and the same effect in a design you can run.

63 % requests meet a slow call

  1. explained

    Why does p99 latency explode long before 100 % CPU?

    The queueing arithmetic for 1, 2 and 16 cores, what bursts do to it, our Postgres measurement, and the same curve in four designs you can open.

    921 ms p99 at 95 % CPU

  2. measured

    How many users can a $12 server handle? We measured it

    A $12-a-month VM with Nginx, Node.js and Postgres under virtual users: about 800 users at one second of think time, and a CPU throttle the guest does not show.

    800 users on a $12 server

  3. measured

    How many requests per second can one Postgres handle?

    One PostgreSQL on a 2-vCPU VM, measured: about 6,000 simple indexed reads per second at 100 % CPU, where p99 starts to climb, and how to estimate your own number.

    6,000 queries per second one Postgres, 2 vCPUs

  4. measured

    How big should your database connection pool be?

    A pool of 4 against a pool of 16 connections on the same Postgres, measured, and Little's law for sizing your own.

    4 connections enough for 2 vCPUs

  5. model

    Cache hit ratio: what 80 % vs 95 % means for your database

    The arithmetic of cache misses, and what a cold cache after a restart does to the database behind it.

    4× less database read load

  6. model

    Retry storms: why retries take your service down

    How retries multiply load, why a storm can outlive its trigger, and the fixes that work, shown in a design you can run.

    16× traffic from two retry layers

  7. model

    System design interview: the URL shortener, simulated live

    The classic interview design with its numbers simulated over a real day of traffic: what breaks first, at which load, and what it costs.

    3× traffic the API breaks first

How closely the live model agrees with an exact simulation is on How accurate is Stackrig?