The Hybrid Wins: Small Language Models in Real Products
A 3.8B parameter model now matches GPT-4o on extraction tasks. Phi-4, Gemma 4, and Qwen 3.5 changed the calculus for what runs locally and what runs in the cloud. The shift isn't 'small models won' — it's hybrid as the default architecture. Here's the pattern that actually ships.
