Colibri Runs GLM-5.2 With 25 GB of RAM, Very Slowly

Colibri can run a 744-billion-parameter model with 25 GB of RAM by streaming weights from a 370 GB file. Cold generation takes 10 to 20 seconds per token.
artificial-intelligence
software-engineering
Author

Kabui, Charles

Published

2026-07-20

Keywords

colibri, glm-5-2, local-ai, mixture-of-experts, nvme-inference