Amil Dravid @ COLM2026
@_AmilDravid
Check out this work co-led by @SophieLWang. We find that base models can match much of RL’s reasoning gains with the right first tokens, which we call 𝘵𝘰𝘬𝘦𝘯 𝘤𝘶𝘦𝘴. The effective cue differs across models and can be as simple as “.\n\nOkay,” or “To determine,” with no
Sophie Wang@SophieLWang · Oct 6Can “chicken” make a base model reason? 🐔
Yes!
Why? With the right first tokens, base models can match RL in reasoning. We trace this to learned training data associations and use a simple data edit to make “chicken” elicit reasoning too.
blog: sophielwang.com/cues
Yes!
Why? With the right first tokens, base models can match RL in reasoning. We trace this to learned training data associations and use a simple data edit to make “chicken” elicit reasoning too.
blog: sophielwang.com/cues
4 73