FreeToken is an open-source inference engine, released by UC Berkeley and UT Austin researchers in August 2026, that splits a MoE model's experts across GPU, CPU and system memory according to each machine's measured bandwidth. It serves models from 35B to 753B parameters on consumer hardware.
Mixtral 8x22B is Mistral AI's Mixture of Experts model released in April 2024: 141B total parameters but only 39B active per token, an unrestricted Apache 2.0 licence, and multilingual performance ahead of Llama 3 70B in Spanish, French, Italian, and German. Production serving needs datacenter-class GPUs.
Gemini 1.5 Pro launched in February 2024 with a verified one-million-token context window. It retrieves over 95% of data up to 530,000 tokens in recall tests, reshaping RAG system design, making full-document analysis viable, and enabling new architectural patterns through context caching.
7 min2254.3
We use first- and third-party cookies to analyze site traffic. You can accept them, reject them, or configure your choice.
Learn more about cookies
Cookie preferences
NecessaryEssential for the site to work. Always on.
AnalyticsHelp us understand how the site is used (Google Analytics).