All Notes

Resume
2026 · profile
Feyi Agbaje Resume
About Feyi Agbaje
2026 · profile
Systems Design Engineer & IESE MBA. 9+ years across enterprise software, AI research, and operations.
Versa
2026 · flagship
Versa is a live daily word game from Dear Barry Games. I built the content quality system for generation, model evaluation, human review, curation, and product analytics.
Inside the Frontier
2025-26 · flagship
I turned a primary-source AI corpus into a grounded Atlas for comparing how frontier labs train, evaluate, and govern model releases without separating the claim from the evidence.
Daybreak
2026 · flagship
I took a GLP-1 companion from a weak generic tracker to a focused product thesis, then carried the strategy through product requirements, privacy boundaries, interaction design, and a deployed build.
Design Studio
2026 · flagship
I built a design-system studio to stop AI coding agents from inventing a new visual language every time they touch a product. It turns visual decisions into reusable tokens, semantic roles, and agent-readable constraints.
Maya Codex
2022 · enterprise
I led a joint research program asking where language models could genuinely help 3D artists learn Maya, where they would fail, and how those failures should change product strategy.
Bifrost Platform Foundations
2018–21 · enterprise
I helped turn an emerging procedural graph system into a product people could find, navigate, reuse, and adopt without breaking established Maya workflows.
Research Operations
2021–23 · enterprise
I turned a research bottleneck into reusable infrastructure: faster recruitment, a shared knowledge system, guarded self-service, analytics, and operational automation for a complex enterprise product organization.
Cached Playback
2018–19 · enterprise
I redesigned a technical caching feature around the way animators actually work, improving discoverability, learnability, control, and recovery while reducing a costly review loop.
Email Feyi 2025-26

Inside the Frontier

I turned a primary-source AI corpus into a grounded Atlas for comparing how frontier labs train, evaluate, and govern model releases without separating the claim from the evidence.
flagshipAI evaluation + governance
Open live Atlas
Inside the Frontier research product overview
Research productThe product opens with decision-relevant findings while retrieval and explorer layers preserve access to the underlying evidence.
Role
AI Research Intern · Product builder
Research synthesis · system design · implementation
Focus
AI evaluation + governance
Cross-lab comparability · grounded research
Skills
AI strategy · RAG / retrieval
Evaluation · safety & governance · data synthesis
Tools
Python · BM25 + BGE · FastAPI
Structured extraction · evidence audit · OpenAI API
00 / Question

Make frontier AI evidence comparable without pretending it is uniform.

Frontier lab reports contain useful evidence, but benchmark names can hide changes in prompts, tools, attempt budgets, graders, thresholds, and release context.

I built a grounded research system to compare how leading labs train, evaluate, and govern releases while keeping claims traceable to reviewed primary sources.

40
reviewed source documents
2,111
indexed evidence chunks
1,015
evaluation occurrences
Research question
How do leading AI labs measure capability and risk, and how does that evidence connect to release decisions?
01 / Research system

Build the evidence system before trusting the synthesis.

The workflow separates curation, chunking, structured extraction, comparison, and synthesis. Each evaluation occurrence is stored by lab and release, with protocol fingerprints that preserve how the evaluation was actually run.

The reader product retrieves from the reviewed source library and exposes source evidence instead of presenting model-generated synthesis as authority.

01
Curate
Define the corpus and evidence boundary.
02
Extract
Structure releases, evaluations, training, and governance evidence.
03
Compare
Audit protocols before comparing scores.
04
Retrieve
Ground questions in reviewed chunks with traceability.
02 / Comparability

The same benchmark was often not the same test.

A shared benchmark name often concealed materially different testing conditions. That made comparability itself a product requirement: the Atlas should warn when two numbers should not be treated as a direct ranking.

Strategy implication
Benchmark rankings can look more precise than the underlying public evidence allows.
Evaluation comparability matrix across frontier AI labs
Protocol audit: A shared benchmark name is a starting point. The surrounding protocol determines whether scores are truly comparable.
03 / Atlas

Turn technical research into something a non-researcher can inspect.

The deployed Atlas combines high-level synthesis, grounded question answering, and structured exploration. A reader can move from conclusion to evidence without reading every source document first.

Atlas product views macOS Browser
Inside the Frontier source drawer
Evidence drawer:
Readers can inspect the source excerpt behind a claim.
1 / 3
04 / Result

A research product that connects AI capability, evaluation, and governance.

The project demonstrates primary-source research, RAG and retrieval design, structured extraction, AI evaluation literacy, safety and governance analysis, evidence auditing, and technical communication.

AI strategy
Connect technical evidence to release and governance decisions.
Retrieval
Ground synthesis in a defined source corpus.
Evaluation
Compare protocol, not just benchmark names and scores.
Communication
Make primary-source research usable by product and strategy audiences.

Working on a hard product problem?

I’m exploring GTM Strategy, AI Product, Product Strategy, and Forward Deployed roles.