Featured image of post Colibrì: Running 700B-Parameter LLMs on CPU Alone, Using SSD as ‘VRAM’

Colibrì: Running 700B-Parameter LLMs on CPU Alone, Using SSD as ‘VRAM’

A pure-C framework enabling prompt-based MoE inference across SSD/RAM/VRAM tiers.

Core Announcement

Core Announcement
Core Announcement|News screenshot

GitHub open-source project Colibrì has recently surged in popularity, reaching 32k Stars. It enables running ultra-large MoE (Mixture of Experts) models on consumer-grade laptops without any GPU. Developed independently by JustVugg, Colibrì is a zero-dependency pure-C implementation requiring no deep learning frameworks.

  • Release Status: Open-sourced; pre-built binaries for Linux/macOS/Windows available
  • Model Support: 9 model families including GLM-5.2/5.3, DeepSeek V4 Flash, Qwen, Inkling, and Kimi K3
  • Quantization: Uniform int4 precision
  • Minimal Hardware Requirement: GLM-5.2 runs on 16GB RAM (24GB recommended); 128GB RAM setups achieve ~1.8 token/s
  • Weight Availability: Framework open-source; models must be acquired separately (conversion tools provided)

Technical Innovation: Multi-tier Loading as Weight JIT

Technical Innovation: Multi-tier Loading as Weight JIT
Technical Innovation: Multi-tier Loading as Weight JIT|News screenshot

Colibrì’s breakthrough lies in decoupling model inference from memory constraints, specifically designed for MoE architectures like GLM-5.2.

MoE models have massive total parameters but activate only a small expert subset per token. Colibrì splits models into two layers:

  • Resident Layer: Attention, Embedding, shared Dense components—~9.9GB int4, always in RAM
  • On-Demand Layer: 19,456 routed experts—~370GB int4, stored entirely on NVMe SSD

During inference, the router identifies required experts. If not in RAM cache, they are loaded from SSD on-demand. After processing, experts are retained or evicted based on usage patterns. Developers describe this as a “JIT for model weights"—loading only hot components when needed.

To optimize I/O, Colibrì implements three-tier scheduling (VRAM/RAM/SSD):

  • Frequent experts reside in faster tiers; infrequent ones stay on SSD
  • LRU caching + frequency tracking for dynamic priority adjustment
  • Predictive preloading: Based on 71.6% predictability of router correlation between adjacent layers, background loading occurs during computation
  • Multi-SSD Parallelism: Duplicate model copies across two SSDs for throughput scaling

Key Metrics & Model Range

Key Metrics & Model Range
Key Metrics & Model Range|News screenshot

Colibrì continuously expands supported models, revealing counterintuitive ratios:

Model (int4)Total ParamsActive/TokenStorage UsageRAM MinGPU Required
GLM-5.2744B~40B~372GB16GBNo
Kimi K32.8T104B~1.6TB32GBNo

Note: Specific activation parameters and hardware requirements for DeepSeek V4 Flash, Inkling, and other models were not detailed in the original source.

Surprising Contrast: Initial inference on a pure-CPU machine (12-core + 25GB RAM) achieves only 0.05–0.1 token/s with cold cache, yet software-only optimizations push speeds to 5.8–6.8 token/s with full VRAM residency (6× RTX 5090), proving that the scheduling strategy contributes orders-of-magnitude performance gains.

Practical Recommendations

Practical Recommendations
Practical Recommendations|News screenshot

  • Ready to Try: Laptop users with 24GB+ RAM seeking to experiment with GLM/Qwen models without GPU; preference for zero-engine dependency and open-source tools
  • Wait Longer: Users requiring >5 token/s throughput, <500GB free disk space, or only 8–16GB RAM—current setup may be suboptimal
  • Pro tip: Web Dashboard and “Brain” visualization show real-time expert locations, memory tiers, and usage heatmaps

Final Thoughts

Colibrì demonstrates that next-gen LLM deployment hinges less on raw hardware and more on intelligent scheduling. The MoE + multi-tier storage path offers a concrete blueprint for accessible local inference.