Goodfire
@GoodfireAI
Autonomous agent swarms are hacking companies. How can interpretability help us understand where this behavior comes from and how to address it?
Our cofounder @banburismus_ spoke with @JPBrebner from SPC about why interpretability matters right now, and how we can move toward
Our cofounder @banburismus_ spoke with @JPBrebner from SPC about why interpretability matters right now, and how we can move toward
Jonathan Brebner@JPBrebner · Aug 13What does the former co-lead of interpretability at Deepmind and now co-founder of @GoodfireAI think about how to keep AI from becoming (his words) "a generally evil guy"?
I got to speak with @banburismus_ at @spc the day after OpenAI announced the Hugging Face hack. Good
I got to speak with @banburismus_ at @spc the day after OpenAI announced the Hugging Face hack. Good
0 68