推文概览
查看 @deedydas 在 2025年7月16日 02:43 发布的这条 X/Twitter 推文。 这条内容包含 1 张图片。
Google DeepMind just dropped this new LLM model architecture called Mixture-of-Recursions. It gets 2x inference speed, reduced training FLOPs and ~50% reduced KV cache memory. Really interesting read. Has potential to be a Transformers killer.







