Building a code reviewer that engineers actually trust is much harder than it looks. Our engineer Dana wrote up the real story behind ours, including the parts that didn't work.
Early versions ran on frontier models and hallucinated defects that weren't there, at $2.07 per review. The fix wasn't a bigger model, but a smaller, more deliberate system: a bounded read-only agent, an ensemble of cheaper models that de-correlate better than four runs of one strong model, and a simple constraint requiring every claim to quote the exact changed line.
That last change alone took false positives from 38% down to 98% verified accuracy.
The result: a reviewer that catches real defects at eleven cents per review in production.
Full breakdown from Dana here: coval.ai/blog/how-we-bu…
Early versions ran on frontier models and hallucinated defects that weren't there, at $2.07 per review. The fix wasn't a bigger model, but a smaller, more deliberate system: a bounded read-only agent, an ensemble of cheaper models that de-correlate better than four runs of one strong model, and a simple constraint requiring every claim to quote the exact changed line.
That last change alone took false positives from 38% down to 98% verified accuracy.
The result: a reviewer that catches real defects at eleven cents per review in production.
Full breakdown from Dana here: coval.ai/blog/how-we-bu…
0 0 0 5 481 1