Best VPS for AI: Choose CPU or GPU by Workload
The best server for AI depends on what runs on the machine. An AI agent or web application that calls an external model API can often use a normal CPU VPS. A local language model, image generator or training job usually needs a GPU with enough VRAM for the selected model, context and concurrency.
This is a primary-source comparison and repeatable test checklist, not a hands-on cross-provider benchmark. Provider documentation and public pricing pages were checked on 2 October 2026. We did not assign ratings or declare a universal winner.
Commercial disclosure: some provider links can earn us a commission. Current provider documentation determines the factual comparison. Commercial relationships do not determine the order.
Start with the workload, not a provider ranking
| Workload | Start with | Why | Do not assume |
|---|---|---|---|
| API-backed agent or web app | CPU VPS | The remote API performs model inference. | A local GPU is included or required. |
| Local LLM inference | GPU instance sized to the exact model | GPU acceleration usually matters for interactive latency and concurrency. | Parameter count alone determines memory or speed. |
| Image, audio or video generation | GPU instance tested with the exact pipeline | Weights, intermediate tensors and batch size consume VRAM. | A cheap CPU VPS will provide useful generation speed. |
| Fine-tuning or training | GPU or multi-GPU system | Training needs memory beyond model weights alone. | An 8 GB CPU VPS is a general training server. |
| Embeddings, preprocessing or vector database | CPU or GPU based on measured throughput | These services can be separated from the model server. | Every component needs the same machine type. |
Four honest AI hosting options
RunPod and Vast.ai provide flexible GPU capacity through different supply and billing models. DigitalOcean publishes fixed GPU configurations and rates. Hostinger is a conventional CPU VPS for the application layer when inference runs elsewhere.
RunPod Pods
Pods are billed by the second for compute and storage. RunPod directs customers to its deployment console for current GPU prices. On-demand Pods use dedicated resources and cannot be displaced.
- Use when: you know the required GPU and VRAM, want a container environment and expect to start or stop workloads.
- Cost check: storage is separate. Persistent volume storage can keep accruing charges while a Pod is stopped.
- Commitment: savings plans cover GPU compute, are prepaid and are non-refundable.
Vast.ai
Vast.ai hosts set real-time prices, locations, storage rates, bandwidth rates and availability. Instances are containerized environments with exclusive GPU access and second-based billing.
- Use when: hardware choice and price flexibility matter and you can evaluate individual offers.
- Cost check: compute, storage and bandwidth are separate marketplace components.
- Risk check: review host reliability, verification, location, rental duration and interruptibility.
DigitalOcean GPU Droplets
DigitalOcean currently lists one NVIDIA RTX 4000 Ada GPU with 20 GB GPU memory, 32 GiB system memory, 8 vCPU and a 500 GiB boot disk at 0.76 USD per GPU-hour.
- Use when: a published configuration and conventional cloud workflow matter more than marketplace pricing.
- Billing: per second with a five-minute minimum.
- Cost check: a powered-off GPU Droplet remains billable until it is destroyed.
Hostinger KVM 2
KVM 2 currently lists 2 vCPU, 8 GB RAM and 100 GB NVMe storage. The public page shows an 8.99 USD monthly promotional equivalent and a 14.99 USD monthly renewal rate for a two-year term.
- Use when: the VPS runs an agent app, automation, database, queue or API gateway while inference runs elsewhere.
- Billing: plans are paid upfront and the monthly figure is an equivalent rate.
- Limit: this is not a GPU server. Measure the exact model before attempting local CPU inference.
Prices, promotions, hardware supply and regional availability can change. Recheck the exact configuration and total cost immediately before purchase.
Split the application layer from model inference
An API-backed agent can keep its web application, queue and database on a modest CPU VPS while a hosted model API performs inference. A self-hosted model replaces that remote call with a separately sized GPU service. This separation prevents buying GPU capacity for components that do not need it.
Persistent path: store application state, model outputs and backups according to the actual recovery and retention requirements. A stopped GPU instance does not always mean storage billing has stopped.
Size the model before the server
There is no honest universal RAM or VRAM minimum for AI hosting. Record these inputs before choosing a plan:
- The exact model and model file or checkpoint.
- The quantization and inference engine.
- The context length and expected output length.
- The number of simultaneous requests and loaded models.
- The target time to first token, throughput or job completion time.
- Persistent storage, download size and output retention.
- Region, data-handling and availability requirements.
Ollama reports whether a model is loaded on GPU, CPU or both through ollama ps. Its documentation also explains that new models must fit in available VRAM for concurrent GPU model loads, while parallel requests and larger context increase memory requirements. A generic label such as 8 GB for AI cannot replace this workload test.
A repeatable benchmark checklist
Use the same container, model, quantization and request set for every candidate:
- Record the exact instance, GPU, driver, region and billing mode.
- Record image pull, model load and cold-start time separately.
- Measure time to first token and tokens per second for local LLM inference.
- Measure median and 95th percentile latency at the intended concurrency.
- Confirm whether the model is fully on GPU or partially offloaded.
- Record maximum VRAM, system RAM and persistent storage use.
- Run long enough to expose throttling, eviction or interruptibility.
- Calculate compute, storage and bandwidth from the same test duration.
- Repeat the test before publishing a winner, score or savings claim.
Which option should you choose?
Choose RunPod when
You want a dedicated GPU Pod, a container workflow and the ability to match the GPU type to a known model. Include persistent storage in the cost estimate.
Choose Vast.ai when
You are comfortable evaluating marketplace offers and want broad hardware and price choice. Review each host and all three billing components.
Choose DigitalOcean when
A published configuration and conventional cloud account matter more than searching a marketplace. Destroy unused GPU Droplets instead of only powering them off.
Choose a CPU VPS when
Your local server runs the application layer and inference happens through an external API or separate GPU service. Do not market it as a GPU substitute.
Frequently Asked Questions
Can a cheap CPU VPS run a local LLM?
It can run some small or heavily quantized models, but useful speed is not guaranteed. Model size, CPU instruction support, memory bandwidth, context and concurrency all matter. Test the exact model and latency target before calling a CPU VPS suitable for production inference.
How much VRAM do I need?
Enough for the exact model, quantization, context cache and concurrent work. The model download size is not a complete capacity estimate. Load the model, inspect actual allocation and keep headroom for the inference engine and expected parallel requests.
Do AI agents need a GPU VPS?
Not when the agent calls a hosted model API. The VPS can run the agent code, database, queue and integrations on CPU. A GPU becomes relevant when model inference or another accelerated workload runs locally.
Is self-hosting automatically more private?
No. Privacy depends on the full data flow, provider, region, access controls, logs, backups, model downloads and external APIs. Self-hosting changes who operates parts of the stack, but it does not by itself establish compliance or eliminate third-party processing.
Is the lowest hourly GPU rate the cheapest option?
Not necessarily. Include storage while stopped, bandwidth, minimum billing, interruptibility, cold starts, unused reservations and the time needed to complete the workload. Compare a representative job or 24-hour operating profile instead of one headline rate.