🇮🇹 This article is also available in Italian.

A year ago we published a MojaLab guide on building a self-hosted LLM stack with Ollama and LiteLLM.

A year in AI, however, feels like a geological era.

The models have changed, the engines have changed, Docker has changed, and above all the way we use these systems has changed.

This time we started from a different question:

Can I rent a GPU VPS, quickly turn it into my own OpenAI-compatible LLM API and, when I am done, destroy the VM without losing models, configuration or data?

The answer is yes.

And that is exactly what this project does:

https://github.com/doradame/GPU_VPS_LiteLLM

The core idea is very simple:

the GPU VPS is temporary. The disk is not.

The machine can be created today, destroyed tonight and replaced by another one tomorrow.

Everything that matters — models, database, configuration, certificates and Docker runtime state — lives on a separate persistent volume encrypted with LUKS.

In practice, we can rent the GPU only when we actually need it.


1. Why Build This Stack?

Before opening a terminal, it is worth asking a much less technical question:

does self-hosting an LLM actually make sense?

Commercial APIs are excellent. We do not have to install NVIDIA drivers, worry about VRAM, maintain containers, or manage the infrastructure, and most importantly we only pay for what we consume.

If what we need is occasional access to the best model currently available, a commercial API is probably the most rational choice.

There is another fairly obvious point in 2026: true frontier models are either not available for self-hosting at all, or they require much more hardware than the kind of GPU we can rent for a few euros per hour.

There are, however, at least three situations where the balance starts to shift.

Data

There are situations where we simply do not want prompts and documents to leave our perimeter.

That may be a business requirement, a regulatory constraint, or simply a design choice.

With this stack, the request reaches a machine we control directly:

flowchart TD
    C([client]) -->|HTTPS| CA[Caddy]
    CA --> L[LiteLLM]
    L --> M[Ollama / vLLM]

The model runs on a GPU we rented, while the state of the system lives on an encrypted disk for which we hold the key.

This is not an air-gapped system, and it does not pretend to be one, but we know exactly where the data flows.

Compute Volume

Commercial APIs are billed by the token.

If we need a few hundred requests per day, that is perfectly fine. If instead we are generating synthetic datasets, running thousands of evaluations, or operating pipelines that process millions of tokens per day, the economics can change significantly.

A GPU rented for a few hours and running a quantized open-weight model can become very attractive from a cost perspective.

There is no universal rule: the numbers have to be worked out case by case.

Control

Self-hosting also means choosing exactly which model to run, which quantization to use, which engine version to pin, how much context to allocate, how many concurrent requests to allow, and when to upgrade.

No rate limit changing unexpectedly, and no model version being silently replaced underneath us.

We can also serve our own fine-tuning, or a model that no commercial provider exposes.

Of course, there is another side to this: a GPU costs money even when it is idle, and updates, security and backups become our responsibility.

The stack we are about to look at is designed to reduce that operational burden as much as possible.

And that is precisely why the machine itself has been designed to be disposable.


2. The Core Idea: The VM Is Temporary, the Data Is Not

Anyone who works with infrastructure knows the principle:

cattle, not pets.

Machines should be replaceable.

With GPU VPS instances this becomes even more important, because an expensive GPU sitting idle still keeps generating a bill.

The stack therefore starts from a fairly strict decision.

The VM system disk contains only:

  • the operating system;
  • the NVIDIA driver;
  • Docker.

Everything else lives on a second disk.

flowchart TD
    NET([Internet]) -->|HTTPS| CA[Caddy]
    CA --> L[LiteLLM]
    L --> O[Ollama]
    L --> V[vLLM]
    L --> PG[(PostgreSQL)]
    subgraph VOL["🔒 persistent LUKS volume"]
        direction TB
        ST["stack/ — compose · config · certs · PostgreSQL"]
        DK["docker/ · containerd/"]
        OL["ollama/ — models"]
        VL["vllm/ — HuggingFace cache"]
    end
    O -.-> OL
    V -.-> VL
    PG -.-> ST

The volume is encrypted with LUKS2.

If we destroy the VM, that disk remains.

When we need GPU capacity again, we create a new VPS, attach the volume and rebuild only the disposable part of the system.

That is the real heart of this project.

LiteLLM, Ollama and vLLM are almost consequences of that architectural choice.


3. Choosing the Infrastructure

Before thinking about models, we need to choose the machine. There is an important distinction here that is not always obvious in the GPU cloud market, because services that are fundamentally different are often sold under very similar names.

GPU Pod or GPU VM?

On one side there are GPU pods or GPU containers: containerized environments where the provider gives us access to a GPU and a ready-to-use workspace.

They are inexpensive, extremely fast to start, and excellent for notebooks, training jobs or temporary workloads. But they are not necessarily real virtual machines: often there is no systemd, we do not truly control the operating system, and most importantly we may not have access to a real block device on which to build our encrypted volume.

For this project we need a real GPU VM: a machine where we have root access, a complete operating system and the ability to attach at least one additional persistent volume.

The minimum checklist therefore becomes:

  • a real VM with root access;
  • an NVIDIA GPU with enough VRAM for the models we want to run;
  • a second persistent block volume, or at least storage that the provider allows us to detach from one VM and attach to another;
  • a public IP reachable on at least ports 80 and 443;
  • the ability to create, with any DNS provider under our control, an A record pointing a domain or subdomain to the public IP of the VPS.

That last point is important: the domain and DNS do not need to come from the same company that rents us the GPU VPS.

They can live somewhere completely different.

What the stack needs is simply a DNS name such as:

llm.example.com

that we can point to the machine's public IP. Caddy will use that hostname to obtain and automatically renew the TLS certificate.

Persistent Storage

The second disk deserves a little more attention, because it is what makes the entire operating model possible.

Ideally, the additional storage should survive deletion of the VM and be attachable later to another machine, at least within the same provider region.

That is where we keep models, database files, configuration, certificates, Docker images and everything else that would be expensive or time-consuming to reconstruct.

The VM, by contrast, must be allowed to disappear without regret.

Which GPU Should You Choose?

When choosing a GPU for LLM inference, the first number to look at is still VRAM, because it determines whether the model can fit at all.

But it is not the only factor.

Once the model fits, memory bandwidth, GPU generation, available Tensor Cores and hardware support for the numerical formats used by the engine — FP8 or FP4, for example — all start to matter.

Two GPUs with the same amount of VRAM may therefore host roughly the same class of models while delivering very different inference performance.

A deliberately rough orientation map, valid as of August 2026, looks like this:

GPU VRAM How I Would Think About It in This Stack
NVIDIA L4 24 GB Interesting entry point for quantized 7B–14B models and moderate concurrency
GeForce RTX 5090 32 GB Strong performance per euro when available from cloud providers, with more headroom than traditional 24 GB cards
NVIDIA L40S 48 GB One of the most comfortable sizes for this project: quantized 32B models, longer contexts or multiple workloads
NVIDIA H100 80/94 GB A clearly higher tier, useful when throughput and concurrency start to matter
RTX PRO 6000 Blackwell 96 GB A lot of VRAM plus Blackwell, attractive for larger models and modern inference workloads
NVIDIA H200 141 GB Opens the door to substantially larger models without immediately moving to multi-GPU
NVIDIA B200 180 GB Blackwell datacenter class, far beyond the needs of a small lab but now part of the cloud landscape

This is not a ranking, and it certainly does not mean that an H100 is always a smarter choice than an L40S.

Hourly price matters at least as much as raw performance.

For the kind of workload described in this article, the best question is often not:

What is the fastest GPU?

but rather:

What is the cheapest GPU that comfortably fits my model and still gives me the throughput I need?

As a very rough rule of thumb, with 4-bit quantization the weights alone require a little more than half the model's parameter count expressed in gigabytes.

A 32B model therefore easily lands in the 18–20 GB range.

But we cannot fill the card with weights alone, because we also need room for the KV cache, engine overhead and concurrent requests.

So instead of thinking only about theoretical limits, I prefer to think in broad classes:

24 GB — 7B and 14B models are the natural territory; larger models may fit with aggressive quantization, but headroom becomes limited.

32 GB — more freedom around quantized 20B–30B-class models, with extra room for context.

48 GB — probably the most flexible tier for a serious lab: a 32B becomes much more comfortable, and running two smaller models on the same GPU starts to become realistic.

80–96 GB — larger models become practical, or the same models can be served with much more KV cache and concurrency.

140 GB and above — we enter a class where very large models can become manageable on a single GPU, naturally at a very different hourly cost.

These are guidelines, not a compatibility chart: model architecture, quantization, context window and inference engine can move the numbers substantially.

And this is exactly where the architecture of the stack becomes useful again: the GPU is not part of persistent state.

We can start today with an L40S, shut everything down, and tomorrow rebuild the machine around an RTX PRO 6000, an H100 or any other GPU the provider makes economically attractive, as long as it has enough VRAM and is supported by the software stack.

The encrypted volume does not need to know which GPU is on the other side.


4. Building the Stack: From Zero to /v1/chat/completions

Once we have the machine, deployment is deliberately not very magical.

No Kubernetes.

No Terraform.

No Ansible.

Just a sequence of numbered Bash scripts.

./scripts/00-preflight.sh

sudo ./scripts/01-nvidia-driver.sh
sudo reboot

sudo ./scripts/02-luks-volume.sh
sudo ./scripts/03-docker.sh
sudo ./scripts/04-nvidia-toolkit.sh
sudo ./scripts/05-stack-config.sh
sudo ./scripts/06-pull-models.sh
sudo ./scripts/07-stack-up.sh
sudo ./scripts/08-test.sh

Each script does one job.

Each completed step leaves a marker behind, and the next script refuses to run if the previous prerequisite has not been completed.

Destructive steps also require the operator to explicitly type:

YES

before they continue.

For a single machine we preferred readable Bash over another layer of abstraction.

When something goes wrong, we can open the script and see exactly what it was trying to do.

00 — Preflight

The first script is essentially an interview.

It asks which disk to use, where to mount it, which domain to configure, which IPs should be allowed, which engines to enable and which models to load.

The answers end up in:

config.env

This is the configuration center of the entire stack.

The file is mode 600 and, obviously, is never committed to Git.

01 — NVIDIA

This installs the NVIDIA driver on the host and prepares the system to use the GPU.

It is the only step that requires a reboot.

Many providers already ship VM images with drivers installed, but we preferred to keep the step explicit and reproducible.

02 — LUKS

The secondary disk is partitioned, encrypted with LUKS2, formatted, mounted and registered in crypttab and fstab.

The stack uses both a passphrase and a keyfile.

That also allows automatic unlock during boot.

We will come back to the threat model of that choice later.

03 — Docker

Docker is installed and its storage is moved onto the encrypted disk.

Not only containers and volumes.

Images too.

That last detail cost us a few hours during testing, but we will come back to it.

The result we want is simple:

the system disk must remain disposable.

04 — NVIDIA Container Toolkit

This is the bridge between Docker and the GPU.

The script finishes by actually starting a container and running:

nvidia-smi

If the container cannot see the GPU, the process stops there.

Better to discover that now than five scripts later.

05 — Stack Generation

At this point config.env is transformed into the actual Docker Compose project.

The script generates:

  • docker-compose.yml;
  • LiteLLM configuration;
  • the Caddyfile;
  • secrets;
  • PostgreSQL configuration;
  • database backup support.

If the LiteLLM master key or PostgreSQL password do not exist yet, they are generated automatically.

06 — Model Download

Models are downloaded before the final stack is started.

Their caches live on the persistent volume.

This step can take a while because we can easily be talking about tens of gigabytes.

It can also be interrupted and resumed.

07 — Start the Stack

Finally:

docker compose up -d

Caddy, LiteLLM, PostgreSQL, Ollama and one or two vLLM instances start, depending on the configuration.

08 — Smoke Test

We do not stop at checking whether the containers are merely "running".

The script verifies liveness, the model list, a real chat completion, every configured model, and the behavior of Caddy's IP allowlist.

At the end we have something like:

curl https://llm.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "Content-Type: application/json" \
  -d '{
        "model": "qwen",
        "messages": [
          {"role": "user", "content": "Hello!"}
        ]
      }'

And this is where one of the main advantages of the project becomes obvious.

This is an OpenAI-compatible API.

For many clients, that means changing only:

base_url
api_key

and continuing to use the SDKs and libraries we already have.


5. Inside the Stack: Components and Responsibilities

At this point the scripts have finished their work and we finally have a functioning stack on the VPS.

So it is worth stopping for a moment and looking not at how we installed it, but at what we actually built.

There are only a few components, and their responsibilities are deliberately clear:

flowchart TD
    NET([Internet]) -->|HTTPS| CA["Caddy<br/>TLS + IP allowlist"]
    CA --> L["LiteLLM<br/>one OpenAI-compatible API · keys · budgets"]
    L --> O["Ollama<br/>flexibility · GGUF · swap models"]
    L --> V["vLLM<br/>throughput · resident model"]
    L --> PG[("PostgreSQL<br/>keys · config · usage")]
    O -.-> VOL[["🔒 LUKS volume — what survives the VM"]]
    V -.-> VOL
    PG -.-> VOL

The idea is that every part should do one job: Caddy protects and exposes the service, LiteLLM presents a single API to clients, Ollama and vLLM actually run the models, PostgreSQL stores application state, and the LUKS volume makes sure that everything important survives the VM.

Caddy: The Boundary with the Internet

Caddy is the only component in the stack that directly exposes public ports.

It terminates HTTPS, obtains and renews the TLS certificate automatically, and enforces the allowlist of authorized client IP addresses.

A client coming from an unauthorized network never reaches LiteLLM at all: it simply receives 403 Forbidden.

This first layer does not replace application-level authentication, but it significantly reduces the exposed surface.

LiteLLM: One Endpoint, Many Models

Behind Caddy sits LiteLLM, the actual application entry point.

The client sees an API compatible with OpenAI's and simply specifies the desired model in the request; LiteLLM decides which backend engine should receive it.

From the outside we can therefore send:

{
  "model": "teacher"
}

or:

{
  "model": "critic"
}

without the client needing to know whether teacher is running on vLLM and critic on Ollama.

LiteLLM also manages API keys and persists users, budgets and usage statistics in PostgreSQL, allowing us to distribute different credentials to applications or people without handing everyone the master key.

PostgreSQL: The State of the Service

PostgreSQL stores what makes LiteLLM more than just a reverse proxy: keys, configuration and usage information.

The database lives entirely on the persistent volume and is backed up every night with a compressed pg_dump and automatic retention.

Ollama: Flexibility

Ollama is the convenient choice when we want to experiment, switch models frequently or use GGUF quantizations.

It is simple, makes it easy to keep several models available, and lets us change them without redesigning the rest of the architecture.

vLLM: Throughput

vLLM approaches the problem differently.

The model tends to stay fixed and resident in GPU memory, while continuous batching and KV-cache management are designed to sustain multiple concurrent requests efficiently.

We use it when the priority is not changing models every five minutes, but getting good serving performance from the model we selected.

The stack can also start two separate vLLM instances on the same GPU, if memory allows it.

The LUKS Volume: What Must Survive

Finally there is the least glamorous but probably most important component: the disk.

The Compose project, database, models, HuggingFace cache, Docker images, certificates and configuration all live on the encrypted volume.

That is what makes it possible to treat the VPS as temporary compute.

If the machine no longer exists tomorrow, the architecture has not disappeared.

It is simply waiting for another GPU.


6. When One GPU Has to Host More Than One Model

The stack also supports a pattern we are using more and more often:

one model generates, another checks.

For example:

flowchart LR
    P([prompt]) --> T["teacher<br/>generates"]
    T --> CR["critic<br/>evaluates"]
    CR --> R([accept / reject])

The same pattern applies to generation plus verification, LLM-as-judge, dataset construction or distillation.

With vLLM, each process serves one model.

Two models therefore mean two containers.

And this is where we encountered one of the more interesting parts of the project: VRAM allocation.

A running model does not consume memory only for weights.

vLLM also preallocates space for the KV cache.

That brings us to:

--gpu-memory-utilization

The problem is that the behavior of this parameter has changed between engine versions.

With some versions, the second process effectively reasons about a cumulative ceiling for the whole GPU.

With newer behavior, each process reasons about its own requested share and compares that against free memory.

The result can be two very different errors:

No available memory for the cache blocks

or:

Free memory on startup is less than desired

We learned this in the traditional way:

by crashing both containers.

One rule, however, remains valid:

startup order matters.

That is why the second vLLM instance is started only after the first one becomes healthy and has already allocated its memory.


7. Day-2 Operations: Running and Maintaining the Stack

Getting something to start once is relatively easy.

The more interesting question is:

what is it like to operate tomorrow?

Most day-to-day changes go through config.env.

To add an Ollama model:

nano config.env

sudo ./scripts/06-pull-models.sh
sudo ./scripts/05-stack-config.sh

docker compose --env-file .env restart litellm

To change the allowed IPs:

nano config.env

sudo ./scripts/05-stack-config.sh

docker compose --env-file .env up -d caddy

To inspect service status:

docker compose --env-file .env ps

To follow logs:

docker compose --env-file .env logs -f

Every service has a healthcheck.

So docker compose ps does not merely tell us that a process exists.

It tells us whether the service is actually ready.

That matters especially with vLLM, where loading tens of gigabytes of weights into GPU memory can take several minutes.

Backups

LiteLLM uses PostgreSQL to store configuration and API keys.

Every night, an automatic compressed pg_dump is created.

The most recent fourteen copies are retained.

These backups live on the encrypted volume together with the rest of the stack.

They are useful for logical errors or database-level problems.

They do not, of course, replace a real off-site backup if our threat model includes complete loss of the provider's storage.


8. Validated by Breaking It

The repository did not start out like this.

At first, it was perfectly linted.

ShellCheck passed.

The YAML was valid.

docker compose config was happy.

Then we installed it on a real GPU VPS.

And a real VPS is an excellent tool for destroying illusions.

The Quoting Bug That Broke config.env

The first version of the configuration generator wrote some values without correct quoting.

As soon as a value contained spaces — for example a list of IPs or model names — Bash interpreted the file incorrectly.

The result was simple: every following script died immediately.

The code was perfectly linted.

And it was perfectly unusable with some of the defaults suggested by the program itself.

Docker Filled the Wrong Disk

We had carefully moved Docker's data root onto the encrypted volume.

Checked.

Printed.

Everything looked perfect.

Then we downloaded the vLLM image.

The system disk climbed to almost 100%.

Recent Docker versions also use the containerd image store, and the image layers were not ending up where we thought they were.

The script displayed the configured data root correctly.

But displaying something is not the same as verifying it.

Since then, the test no longer just reads the configuration.

It actually pulls an image and checks which filesystem receives the bytes.

The Murderous chown

The stack renderer used to run:

chown -R root:root

over the entire stack directory.

Unfortunately, that directory also contained:

pgdata/

PostgreSQL inside the container runs as its own user.

Regenerating the configuration therefore meant changing ownership of live database files while PostgreSQL was using them.

The result:

Permission denied

That bug now has a dedicated regression test.

CI creates a decoy file with the PostgreSQL ownership and fails if the script touches it.

And that is probably the most useful lesson from the entire project:

The gap between "linted" and "validated" closes only when the software is actually executed.

Every real-world bug should leave two things behind:

a fix and a test.


9. Security and Threat Model

The volume is encrypted with LUKS.

To allow automatic unlocking during boot, we keep a keyfile on the VM.

This is a deliberate trade-off.

It protects well if someone obtains:

  • the decommissioned data disk;
  • a snapshot of the data volume alone;
  • the volume mounted separately from the VM.

It does not protect against:

  • an attacker with root access to the running machine;
  • a complete VM snapshot that also contains the keyfile.

If that is not sufficient for your scenario, the solution is different:

  • interactive passphrase entry at boot;
  • remote unlocking;
  • systems such as Clevis/Tang.

The repository does not pretend that the chosen approach protects against threats it cannot actually address.

For our use case, the main goal is to prevent a disk containing models, database files and application data from being mounted elsewhere in cleartext.

And above all:

the keyfile and passphrase must be backed up outside the VPS.

If we lose both, the encrypted volume is not especially secure.

It is simply very, very empty from our point of view.


10. Disaster Recovery: Destroy the VM Without Fear

Now we reach the part I care about most.

We are done for the evening.

The GPU is still costing money.

So we press:

Destroy.

The VM disappears.

With it go the operating system, drivers and system disk.

The encrypted volume remains.

The next morning we create a new GPU VPS, preferably in the same region, attach the disk and update the DNS A record.

Then we clone the repository again:

git clone https://github.com/doradame/GPU_VPS_LiteLLM
cd GPU_VPS_LiteLLM

From the password manager we restore:

config.env
LUKS keyfile

Then:

sudo ./scripts/01-nvidia-driver.sh
sudo reboot

sudo ./scripts/02-luks-volume.sh
sudo ./scripts/03-docker.sh
sudo ./scripts/04-nvidia-toolkit.sh

The LUKS script detects that the encrypted volume already exists.

It does not format anything.

It simply opens and mounts it.

At that point:

cd /srv/llm/stack
docker compose --env-file .env up -d

The models are still there.

PostgreSQL is still there.

The API keys are still there.

The certificates are still there.

The Docker images are still there.

We do not have to download tens of gigabytes again.

We do not have to rebuild the database.

We have simply changed the computer underneath the stack.

And that is exactly the behavior we wanted.

What If I Change GPU Tomorrow?

There is no particular problem.

The volume is not tied to a specific GPU.

If the new GPU has roughly the same amount of VRAM, we can normally restart with the same configuration.

If it has more VRAM, we can increase the context window, concurrency and KV-cache budget.

If it has less, we need to resize the configuration or choose a smaller model.

One thing obviously does not change:

if the model weights physically do not fit in VRAM, no magical configuration parameter will make them fit.

If instead the new GPU is significantly newer and the pinned vLLM version does not support the architecture properly, we update the image version and regenerate the stack.


11. What This Stack Is Not

It is not Kubernetes.

It is not serverless.

It is not a hyperscale platform.

It is not designed for hundreds of tenants.

It is a stack intended for one person, a lab, a small team, internal applications or controlled AI pipelines.

It also does not replace commercial APIs when what we need is the best frontier model currently available.

It answers a different question:

Do I want to run these open-weight models under my own control, on hardware I can rent only when I need it?

If the answer is yes, this architecture starts to make a lot of sense.


12. Why It Was Worth It

We could have rented a managed endpoint and had a model available in a few minutes.

But that was not the experiment.

We wanted a repeatable recipe that could turn a fairly generic GPU VPS into:

a private LLM API
+
OpenAI-compatible
+
with multiple engines
+
with persistent storage
+
encrypted
+
and rebuildable

Now that recipe exists.

And more importantly, it has been tested in the way we prefer:

by breaking it.

The system disk filled almost completely.

The chown that killed PostgreSQL.

The changing behavior between different vLLM versions.

Those are all far more interesting than a screenshot where docker compose ps shows five green lines.

Today, those scars have become automated tests.

The repository is here:

https://github.com/doradame/GPU_VPS_LiteLLM

MIT licensed.

Inside you will find the scripts, complete documentation, the disaster recovery procedure and the CI pipeline.

But the result I actually cared about was much simpler.

Last night I finished using the GPU.

I checked that the volume was fine.

I pressed:

Destroy.

And went to sleep.

Made in MojaLab.

No models were harmed during testing. A couple were restarted in a loop, but they deserved it.