How many users can a $12 server handle? We measured it
Short answer: it depends on what the users do. A $12/month cloud VM running Nginx, Node.js and Postgres served about 800 users who each pause 1 second between page views (796 page views/s, p99 5 ms). At 1,000 users p99 was 865 ms; at 1,300 it broke. The same box with users who never pause, like a load test without sleep, broke at 256 users.
The surprise: under load this box's CPU gets throttled, and Linux does not show it. Steal time stayed at 0 %, while the same code ran 3.5–7× slower per second of CPU time. Profiling it at low load overestimates it by about 2×.
Try it: open “Web app with database” in the playground →
"How many users can my server handle?" is the question behind a lot of capacity planning, and a popular video ("How Many Users Can a $12 Server Handle?") made us want to answer it with our own numbers. Everything below is our own measurement, with its setup, and what our live model predicted for the same box.
Users are not requests per second
A user who loads a page, reads for a while and clicks again sends far fewer requests than a load test that fires as fast as it can. The relation is Little's law for a closed loop:
users = page views per second × (think time + response time)
796 views/s × (1 s + 0.0025 s) ≈ 800 users
So the same box "handles" 800 users who pause a second, and a few hundred who never pause. Ask which users before you ask how many.
The setup
- The box: Google Cloud
e2-small(2 shared vCPU, 2 GB, us-central1), $12.23 a month on demand, Debian 12. - The stack, all on that one VM: Nginx in front of one Node.js process, PostgreSQL 15 with its default settings and a connection pool of 10.
- One page view: one primary-key read from a 1,000,000-row table (popular rows are read more often), an insert on every 10th view, about 900 bytes of HTML.
- The users: virtual users from a second VM in the same zone. Each one thinks for a random time (mean 1 s, or 0 s for the no-pause run), requests a page, waits for the full answer, and repeats. 90 s per step, the first 15 s not counted.
The results (1 s think time)
| Users | Views/s | Mean response | p99 | Box busy |
|---|---|---|---|---|
| 200 | 201 | 2.8 ms | 5 ms | 14 % |
| 400 | 398 | 2.5 ms | 5 ms | 22 % |
| 600 | 600 | 2.4 ms | 4 ms | 29 % |
| 800 | 796 | 2.5 ms | 5 ms | 37 % |
| 1,000 | 740 | 343 ms | 865 ms | 80 % |
| 1,300 | 757 | 712 ms | 1,419 ms | 96 % |
| 1,600 | 740 | 1,161 ms | 2,291 ms | 81 % |
Up to 800 users every page came back in a few milliseconds. Then the box hit its ceiling of about 760 page views per second: more users no longer meant more throughput, only longer waits. Memory was never the limit; the CPU was. Without think time the same ceiling was reached with far fewer users: from 16 to 64 users p99 sat around 200 ms, at 128 users it was near half a second, and at 256 users it passed 1.5 s. (That run started right after the saturated one, on an already throttled box; more on the throttle below.)
The limit you cannot see: a throttled shared CPU
Below the knee one page view cost about 0.8 ms of CPU for the whole box. Past the knee it cost about three times as much, while Linux showed 0 % steal, the number that usually tells you a neighbour or the hypervisor is taking your CPU. So we measured the CPU itself: a fixed loop, counting how much work it got done per second of its own CPU time.
| Box state | Loop iterations per CPU-second |
|---|---|
| Below the knee | 364–396 million |
| Saturated | 52–116 million |
The same code got 3.5–7× less done per second of "CPU time". The e2-small is sold as a shared-core machine with a fraction of its two vCPUs sustained and short bursts above it; when the burst runs out, the vCPUs are slowed down, and Linux books the lost time as busy, not as steal. Two lessons for sizing: a profile taken at low load overestimates this box by about 2×, and "steal = 0 %" proves nothing on a shared-core VM. The honest capacity is the throughput under sustained load.
What Stackrig's live model predicted
We ran the same box in Stackrig's live model with virtual users, twice: once with the capacity taken from a low-load profile, the way you would if you profiled first, and once with the capacity the box really sustains under load. These are model results, compared with the measurement:
| What | Measured | Model, capacity under load | Model, low-load profile |
|---|---|---|---|
| Knee | about 800 users | about 760 users (−5 %) | about 1,600 users (2×) |
| Views/s at 1,300 users | 757 | 759 | 1,298 |
| Mean response at 1,300 users | 712 ms | 711 ms | 3 ms |
| p99 at 1,300 users | 1,419 ms | 820 ms | 9 ms |
With the right capacity the model finds the knee within 5 %, and past it throughput within 3 % and mean response time within 10 %. Two honest limits: past the knee the model's p99 is 1.7–2× too low, because the throttle adds long pauses on top of ordinary queueing, and just below the knee it queues a little too early, because the real box still had burst headroom there. With the low-load profile the model is off by 2× on capacity, which is exactly the trap the throttle sets for any estimate. How the model compares on other setups is on How accurate is Stackrig?
Does a cache help?
We added an in-process cache of the 10,000 most recently read rows (about 56 % of the reads hit it). The ceiling rose from about 760 to about 940 page views per second, +23 %, but the break point barely moved: most of the CPU per page view went to Node.js and Nginx, not to Postgres. One caveat: the cache run started on a box that was already throttled from the previous run, so this is a measured observation, not a clean before/after, and we do not claim a user number for it.
What this measurement does not tell you
- One box, one run per configuration, one zone; shared-core VMs differ between hosts, and the burst credit depends on what the box did before.
- A tiny page. A page with several queries, templates or API calls costs more CPU per view and moves every number down.
- Other VM types, dedicated cores and managed databases behave differently; a dedicated-core VM does not have this throttle.
Try it on your own design
The playground's “Web app with database” template has the same shape (web servers in front of one Postgres). Switch its client to virtual users, set a think time, and watch where your knee is, and what it costs.
Open “Web app with database” in the playground
The playground is invite-only during the private preview: join the waitlist to get an invite.