Home AI Tools Tools VPS Finder Pricing VPS Calculator Benchmarks Migration Guide Cheap VPS Guides Blog
Compare VPS →
AI hosting decision guide

Best VPS for AI: Choose CPU or GPU by Workload

The best server for AI depends on what runs on the machine. An AI agent or web application that calls an external model API can often use a normal CPU VPS. A local language model, image generator or training job usually needs a GPU with enough VRAM for the selected model, context and concurrency.

How this comparison was checked

This is a primary-source comparison and repeatable test checklist, not a hands-on cross-provider benchmark. Provider documentation and public pricing pages were checked on 2 October 2026. We did not assign ratings or declare a universal winner.

Commercial disclosure: some provider links can earn us a commission. Current provider documentation determines the factual comparison. Commercial relationships do not determine the order.

Start with the workload, not a provider ranking

WorkloadStart withWhyDo not assume
API-backed agent or web appCPU VPSThe remote API performs model inference.A local GPU is included or required.
Local LLM inferenceGPU instance sized to the exact modelGPU acceleration usually matters for interactive latency and concurrency.Parameter count alone determines memory or speed.
Image, audio or video generationGPU instance tested with the exact pipelineWeights, intermediate tensors and batch size consume VRAM.A cheap CPU VPS will provide useful generation speed.
Fine-tuning or trainingGPU or multi-GPU systemTraining needs memory beyond model weights alone.An 8 GB CPU VPS is a general training server.
Embeddings, preprocessing or vector databaseCPU or GPU based on measured throughputThese services can be separated from the model server.Every component needs the same machine type.

Four honest AI hosting options

RunPod and Vast.ai provide flexible GPU capacity through different supply and billing models. DigitalOcean publishes fixed GPU configurations and rates. Hostinger is a conventional CPU VPS for the application layer when inference runs elsewhere.

GPU cloud

RunPod Pods

Pods are billed by the second for compute and storage. RunPod directs customers to its deployment console for current GPU prices. On-demand Pods use dedicated resources and cannot be displaced.

  • Use when: you know the required GPU and VRAM, want a container environment and expect to start or stop workloads.
  • Cost check: storage is separate. Persistent volume storage can keep accruing charges while a Pod is stopped.
  • Commitment: savings plans cover GPU compute, are prepaid and are non-refundable.
Check current RunPod GPU availability
GPU marketplace

Vast.ai

Vast.ai hosts set real-time prices, locations, storage rates, bandwidth rates and availability. Instances are containerized environments with exclusive GPU access and second-based billing.

  • Use when: hardware choice and price flexibility matter and you can evaluate individual offers.
  • Cost check: compute, storage and bandwidth are separate marketplace components.
  • Risk check: review host reliability, verification, location, rental duration and interruptibility.
Compare live Vast.ai offers
Published GPU plan

DigitalOcean GPU Droplets

DigitalOcean currently lists one NVIDIA RTX 4000 Ada GPU with 20 GB GPU memory, 32 GiB system memory, 8 vCPU and a 500 GiB boot disk at 0.76 USD per GPU-hour.

  • Use when: a published configuration and conventional cloud workflow matter more than marketplace pricing.
  • Billing: per second with a five-minute minimum.
  • Cost check: a powered-off GPU Droplet remains billable until it is destroyed.
Review DigitalOcean GPU Droplets
CPU VPS

Hostinger KVM 2

KVM 2 currently lists 2 vCPU, 8 GB RAM and 100 GB NVMe storage. The public page shows an 8.99 USD monthly promotional equivalent and a 14.99 USD monthly renewal rate for a two-year term.

  • Use when: the VPS runs an agent app, automation, database, queue or API gateway while inference runs elsewhere.
  • Billing: plans are paid upfront and the monthly figure is an equivalent rate.
  • Limit: this is not a GPU server. Measure the exact model before attempting local CPU inference.
Check Hostinger KVM 2 terms

Prices, promotions, hardware supply and regional availability can change. Recheck the exact configuration and total cost immediately before purchase.

Split the application layer from model inference

An API-backed agent can keep its web application, queue and database on a modest CPU VPS while a hosted model API performs inference. A self-hosted model replaces that remote call with a separately sized GPU service. This separation prevents buying GPU capacity for components that do not need it.

Size the model before the server

There is no honest universal RAM or VRAM minimum for AI hosting. Record these inputs before choosing a plan:

  1. The exact model and model file or checkpoint.
  2. The quantization and inference engine.
  3. The context length and expected output length.
  4. The number of simultaneous requests and loaded models.
  5. The target time to first token, throughput or job completion time.
  6. Persistent storage, download size and output retention.
  7. Region, data-handling and availability requirements.

Ollama reports whether a model is loaded on GPU, CPU or both through ollama ps. Its documentation also explains that new models must fit in available VRAM for concurrent GPU model loads, while parallel requests and larger context increase memory requirements. A generic label such as 8 GB for AI cannot replace this workload test.

A repeatable benchmark checklist

Use the same container, model, quantization and request set for every candidate:

  1. Record the exact instance, GPU, driver, region and billing mode.
  2. Record image pull, model load and cold-start time separately.
  3. Measure time to first token and tokens per second for local LLM inference.
  4. Measure median and 95th percentile latency at the intended concurrency.
  5. Confirm whether the model is fully on GPU or partially offloaded.
  6. Record maximum VRAM, system RAM and persistent storage use.
  7. Run long enough to expose throttling, eviction or interruptibility.
  8. Calculate compute, storage and bandwidth from the same test duration.
  9. Repeat the test before publishing a winner, score or savings claim.

Which option should you choose?

Choose RunPod when

You want a dedicated GPU Pod, a container workflow and the ability to match the GPU type to a known model. Include persistent storage in the cost estimate.

Choose Vast.ai when

You are comfortable evaluating marketplace offers and want broad hardware and price choice. Review each host and all three billing components.

Choose DigitalOcean when

A published configuration and conventional cloud account matter more than searching a marketplace. Destroy unused GPU Droplets instead of only powering them off.

Choose a CPU VPS when

Your local server runs the application layer and inference happens through an external API or separate GPU service. Do not market it as a GPU substitute.

Frequently Asked Questions

Can a cheap CPU VPS run a local LLM?

It can run some small or heavily quantized models, but useful speed is not guaranteed. Model size, CPU instruction support, memory bandwidth, context and concurrency all matter. Test the exact model and latency target before calling a CPU VPS suitable for production inference.

How much VRAM do I need?

Enough for the exact model, quantization, context cache and concurrent work. The model download size is not a complete capacity estimate. Load the model, inspect actual allocation and keep headroom for the inference engine and expected parallel requests.

Do AI agents need a GPU VPS?

Not when the agent calls a hosted model API. The VPS can run the agent code, database, queue and integrations on CPU. A GPU becomes relevant when model inference or another accelerated workload runs locally.

Is self-hosting automatically more private?

No. Privacy depends on the full data flow, provider, region, access controls, logs, backups, model downloads and external APIs. Self-hosting changes who operates parts of the stack, but it does not by itself establish compliance or eliminate third-party processing.

Is the lowest hourly GPU rate the cheapest option?

Not necessarily. Include storage while stopped, bandwidth, minimum billing, interruptibility, cold starts, unused reservations and the time needed to complete the workload. Compare a representative job or 24-hour operating profile instead of one headline rate.

Primary sources checked on 2 October 2026

Related decision guides