Budget VPS for Ollama & Vector Search
🏆 Best overall: Hetzner Cloud · 🥈 Runner-up: Contabo
Run quantized LLMs (7B Q4) and pgvector/Qdrant on 8–16GB RAM VPS with fast NVMe. GPU not mandatory for RAG demos.
Cheapest for this Workload (Top 6)
Our Recommendations
🏆 Best Price/GB Hetzner Cloud
CPX31 (4vCPU/8GB/160GB NVMe) €14.60/mo; scale to CCX for 16GB+ — green EU.
🥈 Most RAM per € Contabo
8–16GB VPS for €8–€16 with unmetered — cost-effective for large contexts.
🥉 US GPU Handoff Vultr
GPU marketplace when you outgrow CPU inference.
Key Takeaways
- Ollama 7B Q4 needs ~5GB + OS — 8GB floor.
- Store embeddings on NVMe — SATA tanks HNSW latency.
Gotchas
- No GPU on pure VPS — latency ~5x vs A10.
- Contabo CPU steal shows under sustained inference — monitor.
Test your own workload
Adjust vCPU, RAM and bandwidth in the calculator.