Tweet Overview
View this X/Twitter post from @deedydas published on lúc 02:43 16 tháng 7, 2025. This post contains 1 images.
Google DeepMind just dropped this new LLM model architecture called Mixture-of-Recursions. It gets 2x inference speed, reduced training FLOPs and ~50% reduced KV cache memory. Really interesting read. Has potential to be a Transformers killer.







