Florian Brand
@xeophon
this continues to be a problem for many (new) rl/env frameworks.
to my knowledge, only verifiers and inspect-swe have fixed this
to my knowledge, only verifiers and inspect-swe have fixed this
Prime Intellect@PrimeIntellect · Aug 25As models become more capable, reward hacks become an increasingly serious problem.
During a controlled experiment, we found a novel reward hack in which agents are able to gain web access in offline sandboxes.
During a controlled experiment, we found a novel reward hack in which agents are able to gain web access in offline sandboxes.
Open quoted post →
2 82