Calvin Huang
@calvinhhuangg
Long-running agents can log thousands of steps per run.
Analyzing every step to find what went wrong is slow and expensive.
@kalmiadev has been experimenting with @typesafeai Jev to help choose which steps matter for our harness to analyze.
Same bugs found on real customer
Analyzing every step to find what went wrong is slow and expensive.
@kalmiadev has been experimenting with @typesafeai Jev to help choose which steps matter for our harness to analyze.
Same bugs found on real customer
6 20