All Notes

Resume
2026 · profile
Feyi Agbaje Resume
About Feyi Agbaje
2026 · profile
Systems Design Engineer & IESE MBA. 9+ years across enterprise software, AI research, and operations.
Versa
2026 · flagship
Versa is a live daily word game from Dear Barry Games. I built the content quality system for generation, model evaluation, human review, curation, and product analytics.
Inside the Frontier
2025-26 · flagship
I built a research Atlas that compares how frontier AI labs train, evaluate, and govern releases—and keeps every claim attached to evidence.
Roam
2026 · lab
I am building a Voice AI pipeline that turns open location data into narrated walking-tour assets, with quality checks designed for data, scripts, and synthesized audio.
Design Studio
2026 · flagship
I built a design-system studio to stop AI coding agents from inventing a new visual language every time they touch a product. It turns visual decisions into reusable tokens, semantic roles, and agent-readable constraints.
Maya Codex
2022 · enterprise
I led a joint research program asking where language models could genuinely help 3D artists learn Maya, where they would fail, and how those failures should change product strategy.
Bifrost Platform Foundations
2018–21 · enterprise
I helped turn an emerging procedural graph system into a product people could find, navigate, reuse, and adopt without breaking established Maya workflows.
Research Operations
2021–23 · enterprise
I turned a research bottleneck into reusable infrastructure: faster recruitment, a shared knowledge system, guarded self-service, analytics, and operational automation for a complex enterprise product organization.
Cached Playback
2018–19 · enterprise
I redesigned a technical caching feature around the way animators actually work, improving discoverability, learnability, control, and recovery while reducing a costly review loop.
Email Feyi 2026

Versa

Versa is a live daily word game from Dear Barry Games. I built the content quality system for generation, model evaluation, human review, curation, and product analytics.
flagshipEval-first AI product
Play Versa Try Versa Review
Current Versa Review interface showing the queue, candidate details, model insights, and human annotation panel.
Versa Review V3. The current workflow gives the human decision more visual priority than the model score.
Role
Co-founder · product + AI systems
Strategy · pipeline architecture · eval design · review UX
Focus
Eval-first AI product
Synthetic data · human review · production quality
Skills
AI evaluation · Human-in-the-loop
Product analytics · system design · failure analysis
Tools
Gemini · GPT · Claude · PostHog
LLM-as-judge · structured annotation · deterministic rules
00 / Live product

We launched a playable daily word game and began measuring how people use it.

I co-founded Dear Barry Games with my mum and launched Versa, a daily word game built around antonyms and the multiple meanings of one answer.

Promotion so far has been limited to two LinkedIn posts and sharing with family and friends. The first 30 days therefore give me an early product baseline, not evidence of scale.

56
people started a puzzle
85.7%
completed after starting
15
returned to play on another day
Data note
PostHog person-level activity from 26 July to 25 August 2026, in UTC. Test accounts are excluded.
01 / Quality system

Generating puzzles was easy. Knowing which ones were good enough to publish was harder.

I built a pipeline that combines deterministic quality gates, same or near-family and cross-family model evaluation, human review, curation, and post-launch measurement.

Hard gates catch structural failures before publication. Reviewers can still override strong model scores when an antonym is wrong, the meanings are too similar, or a clue gives away the answer.

1,408
original candidates evaluated
437
candidates reviewed by people
54.5%
failed at least one explicit quality gate
70.5%
of human-flagged bad-antonym cases had model antonym scores of 95+
Why human review stayed
Model confidence was not enough. Human reviewers found semantic failures in highly scored candidates, and 55% of multiply-reviewed candidates produced conflicting human verdicts.
02 / Versa Review

The review system had to work for the people doing the review.

Reviewers inspect clues and model evidence, record a verdict, change difficulty, add quality tags, leave notes, and resolve disagreements.

My parents reviewed real content, often on their phones. I rebuilt Versa Review around the devices they used: a multi-panel desktop view, tablet panels, and a focused queue → candidate → annotation flow on phones.

Product outcome
The same review workflow now adapts to desktop, tablet, and phone instead of forcing every reviewer into a desktop interface.
Selected screens

One workflow, adapted to each screen

3 views
Versa Review desktop interface showing the queue, candidate evidence, and human annotation panel.
Desktop Queue, evidence, and annotation stay visible together.
Versa Review tablet interface with the annotation panel open beside the candidate.
Tablet Panels preserve context without crowding the review task.
Versa Review phone interface focused on the reviewer annotation workflow.
Phone The interface focuses on one decision at a time.
03 / What I learned

The first batch exposed problems in both the prompts and the generation architecture.

Reviewing the prompts more carefully before a large run would have saved time and money. Each stage also needs its own evaluation so a failure can be traced to generation, judging, curation, or delivery.

The duplicate analysis found only 445 distinct answers across 1,408 candidates. Parallel agents produced 963 excess candidates beyond the first occurrence of each answer, even though every ordered clue sequence was unique. The system created variation, but far too much of it converged on the same target words.

445
distinct answers across 1,408 candidates
68.4%
excess duplicate candidates
100%
of ordered clue sequences were unique
Architecture change
I would use one generation process with an exclusion list for previous answers, while keeping the same or near-family and cross-family evaluators.
04 / Next iteration

The next iteration is a smaller, measured experiment.

I will revise the prompt using observed failures, then test a small batch against the current shortlist rate before paying for another large run.

I am still evaluating the best generation approach. One option is to compare a few models across easy, medium, and hard puzzles, auto-grade the results, and choose the best quality-to-cost tradeoff.

Longer-term goal
Create a continuous improvement process with clear definitions of success and failure, instrumented from generation through live play.
01
Validate the prompt
Run a small test with explicit success criteria before scaling generation.
02
Simplify generation
Use one generator, exclude the 445 existing answers, and keep independent evaluators.
03
Build regression sets
Use launch-quality puzzles as positive examples and observed failures as cases the system must catch.
04
Close the loop
Add controlled puzzle editing to the review product and connect pre-release signals to completion, clue use, mistakes, sharing, and retention.

Working on a hard product problem?

I’m exploring GTM Strategy, AI Product, Product Strategy, and Forward Deployed roles.