Applied AI Engineering
Hosted by R3i Ventures — AI Lab
Prototypes are easy. You will build an evaluation harness first, then the feature, then close the gap between demo behaviour and real behaviour — with cost, latency and failure modes measured rather than hoped for.
Skills assessed
The six sprints
Each sprint has one named outcome. Your mentor signs it off, sends it back with written challenges, or fails it — the same three options the permanent team gets.
- W114 hrs · 3 tasks
Eval before build
Outcome — A graded eval set of 50 real cases before a line of feature code.
Collect real inputs
ResearchFifty genuine cases, including the ugly ones.
5 hours
Define grading
BuildWhat does good look like? Write the rubric.
5 hours
Baseline run
ReviewScore the naive prompt. Record it. Never delete it.
4 hours
- W214 hrs · 3 tasks
Prototype
Outcome — A working feature that beats the baseline on your own eval.
Build the loop
BuildTool definitions, retries, structured output.
6 hours
Iterate against eval
BuildChange one thing at a time; log every score.
6 hours
Cost and latency read
ReviewPer-call cost and p95 latency, measured.
2 hours
- W314 hrs · 3 tasks
Failure analysis
Outcome — A taxonomy of how it fails, ranked by user harm.
Read every failure
ResearchAll of them. By hand. No shortcuts.
6 hours
Build the taxonomy
BuildGroup failures by cause, not by symptom.
4 hours
Rank by harm
ReviewWhich failure would you least want screenshotted?
4 hours
- W414 hrs · 3 tasks
Guardrails
Outcome — The top failure class reduced by half, measured on held-out cases.
Design the mitigation
BuildPrompt, tool, validation or refusal — pick deliberately.
5 hours
Implement and measure
BuildHeld-out set only. No peeking at the training cases.
6 hours
Regression check
ReviewConfirm you did not break the cases that worked.
3 hours
- W514 hrs · 3 tasks
Production readiness
Outcome — Deployed behind a flag with observability and a kill switch.
Ship behind a flag
BuildStaged rollout, kill switch tested for real.
5 hours
Production telemetry
BuildLog inputs, outputs, scores, cost — privacy-safe.
5 hours
Shadow traffic
ReviewRun silently against live traffic; compare.
4 hours
- W614 hrs · 4 tasks
Technical review
Outcome — A written eval report and a defence to the engineering panel.
Write the eval report
BuildBaseline, deltas, remaining failure modes, cost.
5 hours
Code review
ReviewFull review with a staff engineer; act on it.
4 hours
Panel defence
PresentPresent the numbers and what you would do next.
2 hours
Case study writeup
BuildPublish to your FutureFit skills passport.
3 hours
Ready to take the track?
Applications close two weeks before the Autumn 2026 cohort starts.