Built Agent Eval System Powering 70 Video AI Model Iterations
Designed, built, and operated a human-in-the-loop evaluation product as sole engineer — 14,390 judgments across 105 model builds. Grading throughput rose from ~10 to ~26 builds a month, saving $200K in vendor cost.