Build
The product
Evaluation datasets, indexing, internal review software, defect taxonomy, dashboard, automation, and AI agent skills.
AI evaluation infrastructure · Integrated case study
As sole engineer and operating lead, I turned subjective editorial judgment into a faster model-improvement loop—14,390 human decisions, 132 pull requests, and $200K in avoided vendor cost.
01 / At a glance
02 / The problem
Automated checks could confirm that a model ran. They could not reliably judge editorial timing, visual continuity, awkward repetition, narrative coherence, or whether one edit simply felt better than another.
The product needed human sensitivity at engineering speed.
03 / The operating change
| Old system | New system |
|---|---|
| External vendor separated from model development | In-house reviewers working close to AI and engineering |
| Manual assignment and spreadsheet coordination | Automated assignment and one operating dashboard |
| Annotations collected as output | Judgments analyzed as model-quality evidence |
| Defects relayed through coordination | Findings connected to Linear-linked issues |
| Approximately 10 graded builds per month | Approximately 26 graded builds per month |
04 / What I owned
Build
Evaluation datasets, indexing, internal review software, defect taxonomy, dashboard, automation, and AI agent skills.
Operate
Hiring, onboarding, training, assignment, quality control, tracked hours, and the finance payout workflow.
Improve
Analysis with the AI team, recurring failure patterns, Linear-linked tickets, and 17 weekly quality reports.
05 / The system I built
The product, automation, human operation, and engineering feedback loop worked as one system.
Candidate PR → hypothesis → isolated deployment. An engineer opened a candidate branch, stated what should improve and why, then deployed it to an isolated Cloud Run revision.
Candidate and baseline arms. The workflow pinned a stable render pool, generated 50 candidate outputs, reused or rendered the reference arm, and published matched video sets.
Three approval gates before assignment. The orchestrator filed the Linear request and build changelog, prepared videos and metadata, previewed capacity, assigned two randomized repetitions, and drafted the reviewer briefing.
Side-by-side editorial judgment. The approved assignment and briefing moved into the reviewer queue. Reviewers recorded preference strength, flaw severity across video, audio, text, and prompt adherence, free-text comments, and unusable outputs. Disagreements entered a third-opinion tie-break pool.
Win rates plus qualitative diagnosis. The Comparison Suite resolved disagreements, calculated candidate-versus-reference performance, built flaw scoreboards, summarized reviewer comments, applied a consistent tie convention, and published the weekly evaluation report.
From visible flaw to technical cause. The analysis handed high-priority findings to Linear and hg-debug. A flaw could be promoted into a labeled bug ticket, linked back to the model session, and traced through LLM steps, tool calls, asset selection, indexing, prompts, and rendering.
Evidence returned to the backlog. Root causes and prioritized defects moved back to the AI team, informing the next fix, candidate branch, stated hypothesis, and evaluation cycle.
Technical output: 132 pull requests across four connected products in 10 weeks, with 98% merged, as the sole engineer.
06 / The key judgment call
I retired two vendor-facing systems containing 57,690 legacy annotations, removed approximately 3,200 lines of obsolete tooling, and consolidated the work into one internal tool and a four-person team.
The decision saved the company $200K in vendor contract cost while creating a faster, tighter model-feedback loop.
07 / Results
08 / What this proves