Anthropic
@AnthropicAI
We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports.
Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real
Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real
anthropic.comInvestigating unintended model actions in our evaluations and internal use
379 3.2K