Marianne Arriola
@mariannearr
Thrilled to see this OSS diffusion LLM from NVIDIA explore our work on encoder-decoder dLMs at scale (x.com/mariannearr/st…). Excited to see what the community builds with it!
NVIDIA AI@NVIDIAAI · Jul 1We took a 30B model and split it in two to write tokens in parallel instead of one at a time.
Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other
Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other
· Edited
0 17