Adam Gleave
@ARGleave
Today's models are capable of hacking into production systems. Developers deploy guardrails to stop misuse, but if you find a jailbreak that bypasses this guardrail, there's no way to tell the developer. This is a serious oversight -- me & @richb_c talk about how to fix it.
2 14
AI Frontiers