Volodymyr Kuleshov 🇺🇦
@volokuleshov
Where do diffusion language models shine? One major use case is voice agents.
With diffusion LLMs running at >1000 tok/sec, the model can reason in real time before producing outputs. This boosts quality, and latency is still low enough for a live phone call.
Check it out 👇
With diffusion LLMs running at >1000 tok/sec, the model can reason in real time before producing outputs. This boosts quality, and latency is still low enough for a live phone call.
Check it out 👇
Inception@_inception_ai · Jul 14Smart enough to reason, fast enough for a phone call. On real voice agent prompts, nothing beats Mercury 2 on both.
The world's first reasoning diffusion LLM: full reasoning pass in <300ms at 1000+ tok/s on standard NVIDIA GPUs.
💬"The reasoning quality we need without
The world's first reasoning diffusion LLM: full reasoning pass in <300ms at 1000+ tok/s on standard NVIDIA GPUs.
💬"The reasoning quality we need without
Open quoted post →
2 29