Colibrì: Running Large AI Models Without a GPU
Discover how Colibrì uses CPU, RAM and NVMe storage to run extremely large AI models without requiring a dedicated GPU server.
We normally associate AI with powerful GPU servers, expensive hardware and huge amounts of memory. But a project called Colibrì, created by an Italian developer, is exploring a very different idea: running extremely large AI models without depending on a GPU server.
The approach uses a combination of CPU, RAM and fast NVMe storage. Instead of keeping an enormous model completely in memory, the system brings in only the parts of the model that are actually needed. That changes the way we think about AI infrastructure.
GPU server vs. the Colibrì approach
GPU server:
Model → GPU VRAM → Inference
Colibrì-style approach:
Model → NVMe → RAM → CPU → InferenceIt is not necessarily faster than a GPU. In fact, performance can be much slower. The interesting part is something else: a model that doesn't fit into normal memory may still be usable on relatively modest hardware.
This could open interesting possibilities for local AI, private AI, AI agents, healthcare systems and organizations that want to reduce their dependence on expensive GPU infrastructure.
A simple example
Imagine a small company wants to run an AI assistant on its own server. Instead of purchasing an expensive GPU server, it could use a machine with a powerful CPU, 32–64 GB RAM and a fast NVMe SSD. The AI model is stored on the SSD.
When a user asks a question, the system does not need to keep the entire model in RAM. It can load the required model components, process the request and return the answer.
User: "Summarize this 50-page report"
↓
AI Agent
↓
Load required model data
↓
CPU + RAM + NVMe
↓
Generated SummaryIt may be slower than a GPU-based system, but the company could potentially avoid investing in a dedicated GPU server. This is particularly interesting for local AI, private AI, AI agents, healthcare systems and cost-sensitive deployments.
Real example: using Colibrì with GLM-5.2
Colibrì's documentation provides a practical example of running a very large Mixture-of-Experts model, GLM-5.2, locally. The documented setup can run CPU-only and streams model experts from disk instead of requiring the complete model to be resident in RAM.
Example hardware requirements from the current Quick Start:
- RAM: approximately 16 GB minimum; 24 GB or more recommended
- Free disk: approximately 380 GB for the int4 model
- Fast NVMe SSD recommended, because storage speed strongly affects generation speed
- GPU: not required; CPU-only operation is supported
How the example works
Company server
↓
CPU + RAM + Fast NVMe SSD
↓
Colibrì inference engine
↓
GLM-5.2 744B MoE model
↓
AI responseAfter the model is prepared, the Quick Start guide shows a local chat command such as:
COLI_MODEL=/nvme/glm52_i4 ./coli chatThe same project can also expose an OpenAI-compatible API, allowing an existing application to communicate with the local model.
Colibrì is not necessarily faster than a GPU server. Its key innovation is using storage, RAM and available GPU/CPU resources as a memory hierarchy, so very large models can be explored on hardware with limited fast memory.
Useful GitHub links
- Colibrì main repository: https://github.com/JustVugg/colibri
- Colibrì Quick Start guide: https://github.com/JustVugg/colibri/blob/main/docs/quickstart.md
The bigger lesson
The future of AI may not be about using more hardware. It may be about using hardware more intelligently.