Archive / Vps

How to Self-Host an LLM on a VPS: A Practical Setup Guide

You can self-host a large language model (LLM) on a VPS by choosing a model that fits your server's RAM, using an inference server like Ollama or vLLM, and exposing it through an API. The key is matching the model's memory requirements to your VPS specs, which often means starting with a smaller, quantized model on an entry-tier KVM VPS with at least 4-8GB of RAM.

Step 1: Choose the Right LLM for Your VPS

Your most important choice is a model that fits your VPS memory. Unquantized models often need over 20GB of RAM, which is too much for most virtual servers. Quantized versions (like GGUF or GPTQ formats) use much less memory with little drop in quality.

Note

If you need a VPS with predictable, dedicated resources for consistent LLM performance, consider a KVM-based VPS from Hostinger. Their plans offer full virtualization, which is better for memory-intensive AI workloads than shared hosting.

Step 2: Prepare Your VPS for LLM Deployment

After getting a VPS, configure it to run an LLM inference server. This usually means installing a container runtime and the server software.

  1. Access Your Server: Connect via SSH. Most LLM tools assume a Linux environment like Ubuntu 22.04.
  2. Install Docker (Recommended): This simplifies dependency management. Run sudo apt update && sudo apt install docker.io.
  3. Pull an Inference Server: Ollama is the easiest for beginners. Install it with a single curl command from their docs. For higher performance and API compatibility, vLLM is a good alternative.

Make sure your server has enough swap space (at least 4-8GB) to handle memory spikes when loading models, even if your VPS plan has limited physical RAM.

Step 3: Run and Access Your Self-Hosted LLM

Once the software is installed, you can pull and run your model. Using Ollama as an example, the commands are simple.

# Pull a quantized model (e.g., Mistral 7B)
ollama pull mistral:7b
# Run the model server
ollama serve
# In a new terminal, run a query
timeout 30 ollama run mistral "Explain quantum computing"

To use the LLM from other applications, expose its API. Ollama runs on localhost:11434 by default. Use a reverse proxy like Nginx with SSL for production. For simple testing, create an SSH tunnel to your local machine. Your LLM then acts as a private API endpoint.

What to Do After Your LLM is Running

Getting the model running is the first step. For a practical setup, you'll want to integrate it.

Self-hosting changes the cost from API fees to infrastructure. The break-even point depends on your usage. For consistent, private, or high-volume queries, running on your own VPS server can be more economical and controllable over time.

Frequently asked questions

How much does it cost to self-host an LLM on a VPS?

The primary cost is the VPS itself. You can start with an entry-level KVM VPS plan costing a few dollars a month, providing 2-4 vCPUs and 4-8GB of RAM suitable for smaller, quantized models. Costs scale with RAM and CPU; running larger models like Llama 3 70B may require a high-memory tier plan costing significantly more.

Do I need a GPU on my VPS to run an LLM?

No, but it helps immensely. You can run quantized LLMs using only the CPU on any standard VPS, but inference will be slow. A VPS with a dedicated GPU (often called a Cloud GPU) can make responses 10-100x faster, but these plans are more expensive. For initial testing and light use, a CPU-only VPS is sufficient.

What's the difference between Ollama and vLLM for self-hosting?

Ollama is designed for simplicity and local use, making it easier to get started with pulling and running models. vLLM is a high-performance inference server optimized for throughput and production API serving, offering features like continuous batching. Start with Ollama for prototyping; consider vLLM if you need to serve multiple users or integrate with other tools expecting an OpenAI-compatible API.

Is self-hosting an LLM more private than using ChatGPT?

Yes. When you self-host, all prompts, model weights, and generated text remain on your virtual private server. There is no data sent to or stored by a third-party company like OpenAI. This guarantees data privacy and sovereignty, which is crucial for handling sensitive or proprietary information.

Affiliate disclosure: If you buy through our links, we may earn a commission at no extra cost to you.

Ready to get started? Check out Hostinger's plans.

Related reads