A Local LLM on a $9 VPS: 9 Tokens a Second, No Matter How Many Users
Ollama on a CPU-only 2 vCPU server: 8.7 tokens per second, throughput flat under concurrency, 12 GB of disk, and double the response time on the five sites sharing the machine.
Ollama on a CPU-only 2 vCPU server: 8.7 tokens per second, throughput flat under concurrency, 12 GB of disk, and double the response time on the five sites sharing the machine.
Same server, same MySQL, same article injected into both. Ghost served 3.1x more requests uncached — and generated HTML at the same rate, because its page was 3.4x smaller.
Cloudways Medium is $54 a month for 4 GB and 2 vCPU; a Hostinger KVM2-class VPS is around $9 for 8 GB. The decision is labour, not hardware: the price gap buys 52 minutes of your month. With measured VPS benchmarks showing a 40x spread from caching alone, and a protocol for benchmarking any host yourself.
Both deployed on a live VPS under identical 1 vCPU limits. Cal.com idled at 1.05 GB of RAM with an 8.05 GB image; Easy!Appointments at 27 MB and 877 MB. Includes the reverse-proxy rate-limiter trap that returns 429 to every visitor, and its verified one-line fix.
↑↓ navigate ↵ open esc close