UC Berkeley Open-Sources FreeToken: Frontier-Scale MoE on Consumer Hardware

Posted on Sat 05 September 2026 in AI Infrastructure • Tagged with MoE, Inference, Edge AI, Open Source, LLM

UC Berkeley has open-sourced FreeToken, an edge-native inference engine that runs frontier-scale Mixture-of-Experts (MoE) models on consumer hardware. It treats GPU, CPU, host memory, and PCIe interconnects as a single unified inference platform.

The approach works because of how MoE models are structured. A model like DeepSeek-V4-Flash carries 284B total …


Continue reading

LLM Inference Optimization: What Actually Makes Your Model Fast

Posted on Sat 18 April 2026 in GenAI • Tagged with LLM, Inference, Optimization, Quantization, KV Cache, Speculative Decoding, Flash Attention

When you send a prompt to an LLM, three layers shape how fast you get a response: the hardware (GPUs, TPUs, LPUs), the model size and architecture, and the inference engine strategies sitting on top. Most of the latency battle is fought at that third layer — and the core problem …


Continue reading