RadixArk
@radixark
We trained a DSpark speculator for Inkling NVFP4, built end-to-end with SpecForge on a live SGLang target engine.
- 1.89x decode throughput over non-spec at bs=64 (8xB200, TP8)
- 7–14% faster than Inkling's built-in MTP at the same batch size
- 3.66 mean accept length, up to
- 1.89x decode throughput over non-spec at bs=64 (8xB200, TP8)
- 7–14% faster than Inkling's built-in MTP at the same batch size
- 3.66 mean accept length, up to
4 73