Rhythm Garg
@rhythmrg
Every redundant tensor, copy, checkpoint stall, or extra rollout node compounds into meaningful effects on the efficiency of our post-training stack. Excited that our work to support Kimi K3 is now making frontier open models much faster and cheaper to train on AC2 across model
Applied Compute@appliedcompute · Aug 28Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%.
At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to
At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to
1 32