All Notes

Resume
2026 · profile
Feyi Agbaje Resume
About Feyi Agbaje
2026 · profile
Systems Design Engineer & IESE MBA. 9+ years across enterprise software, AI research, and operations.
Versa
2026 · flagship
Versa is a live daily word game from Dear Barry Games. I built the content quality system for generation, model evaluation, human review, curation, and product analytics.
Inside the Frontier
2025-26 · flagship
I built a research Atlas that compares how frontier AI labs train, evaluate, and govern releases—and keeps every claim attached to evidence.
Roam
2026 · lab
I am building a Voice AI pipeline that turns open location data into narrated walking-tour assets, with quality checks designed for data, scripts, and synthesized audio.
Design Studio
2026 · flagship
I built a design-system studio to stop AI coding agents from inventing a new visual language every time they touch a product. It turns visual decisions into reusable tokens, semantic roles, and agent-readable constraints.
Maya Codex
2022 · enterprise
I led a joint research program asking where language models could genuinely help 3D artists learn Maya, where they would fail, and how those failures should change product strategy.
Bifrost Platform Foundations
2018–21 · enterprise
I helped turn an emerging procedural graph system into a product people could find, navigate, reuse, and adopt without breaking established Maya workflows.
Research Operations
2021–23 · enterprise
I turned a research bottleneck into reusable infrastructure: faster recruitment, a shared knowledge system, guarded self-service, analytics, and operational automation for a complex enterprise product organization.
Cached Playback
2018–19 · enterprise
I redesigned a technical caching feature around the way animators actually work, improving discoverability, learnability, control, and recovery while reducing a costly review loop.
Email Feyi 2025-26

Inside the Frontier

I built a research Atlas that compares how frontier AI labs train, evaluate, and govern releases—and keeps every claim attached to evidence.
flagshipAI evaluation + governance
Open live Atlas
Inside the Frontier research product overview
The Atlas opens with the findings, then lets readers inspect the evidence behind them.
Role
AI Research Intern · Product builder
Research synthesis · system design · implementation
Focus
AI evaluation + governance
Cross-lab comparability · grounded research
Skills
AI strategy · RAG / retrieval
Evaluation · safety & governance · data synthesis
Tools
Python · BM25 + BGE · FastAPI
Structured extraction · evidence audit · OpenAI API
00 / Question

Make frontier AI evidence comparable without pretending it is uniform.

I reviewed public technical reports and system cards from leading AI labs, then structured the evidence so readers could compare releases without losing the surrounding context.

40
reviewed source documents
2,111
indexed evidence chunks
1,015
evaluation occurrences
Research question
How do leading AI labs measure capability and risk, and how does that evidence connect to release decisions?
01 / Finding

The same benchmark was often not the same test.

Prompts, tools, attempt budgets, graders and thresholds varied across labs. A shared benchmark name was only the start of the comparison—not proof that two scores belonged in a ranking.

Strategy implication
Benchmark rankings can look more precise than the underlying public evidence allows.
Evaluation comparability matrix across frontier AI labs
Protocol audit: A shared benchmark name is a starting point. The surrounding protocol determines whether scores are truly comparable.
02 / Atlas

Move from finding to source without leaving the product.

The Atlas combines concise findings, grounded question answering and a searchable evaluation catalog. Readers can open the source behind a claim instead of taking the synthesis on trust.

Selected screens

Inside the Atlas

3 views
Inside the Frontier grounded question and answer interaction
Ask the Atlas A research question returns an answer grounded in the reviewed source library.
Inside the Frontier source drawer
Evidence drawer Readers can inspect the source excerpt behind a claim.
Inside the Frontier data explorer
Data explorer Structured records make evaluation evidence searchable by lab and release.
03 / Method

Design the evidence boundary before the interface.

I kept curation, extraction, comparison and retrieval as separate stages. That made missing evidence visible and stopped fluent synthesis from outrunning the public record.

Research principle
Where the public record stops, the product should say so.
01
Curate
Set the source boundary.
02
Extract
Structure the evidence.
03
Compare
Audit the protocol.
04
Retrieve
Return claims with sources.

Working on a hard product problem?

I’m exploring GTM Strategy, AI Product, Product Strategy, and Forward Deployed roles.