Tweet Overview
View this X/Twitter post from @deedydas published on 2025년 7월 16일 오전 02:43. This post contains 1 images.
Google DeepMind just dropped this new LLM model architecture called Mixture-of-Recursions. It gets 2x inference speed, reduced training FLOPs and ~50% reduced KV cache memory. Really interesting read. Has potential to be a Transformers killer.







