Jerry Tworek
@MillionInt
Oh, hi!
Samip@industriaalist · Sep 22We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning!
Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3.
We've
Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3.
We've
Open quoted post →
6 240