R3i FutureFitVirtual InternshipsBrowse tracks
← All tracks
EngineeringAdvancedAutumn 2026

Applied AI Engineering

Hosted by R3i Ventures — AI Lab

Prototypes are easy. You will build an evaluation harness first, then the feature, then close the gap between demo behaviour and real behaviour — with cost, latency and failure modes measured rather than hoped for.

Skills assessed

Evaluation designPrompt engineeringLatency & cost controlFailure analysisGuardrails

The six sprints

Each sprint has one named outcome. Your mentor signs it off, sends it back with written challenges, or fails it — the same three options the permanent team gets.

  1. W1

    Eval before build

    Outcome — A graded eval set of 50 real cases before a line of feature code.

    14 hrs · 3 tasks
    • Collect real inputs

      Research

      Fifty genuine cases, including the ugly ones.

      5 hours

    • Define grading

      Build

      What does good look like? Write the rubric.

      5 hours

    • Baseline run

      Review

      Score the naive prompt. Record it. Never delete it.

      4 hours

  2. W2

    Prototype

    Outcome — A working feature that beats the baseline on your own eval.

    14 hrs · 3 tasks
    • Build the loop

      Build

      Tool definitions, retries, structured output.

      6 hours

    • Iterate against eval

      Build

      Change one thing at a time; log every score.

      6 hours

    • Cost and latency read

      Review

      Per-call cost and p95 latency, measured.

      2 hours

  3. W3

    Failure analysis

    Outcome — A taxonomy of how it fails, ranked by user harm.

    14 hrs · 3 tasks
    • Read every failure

      Research

      All of them. By hand. No shortcuts.

      6 hours

    • Build the taxonomy

      Build

      Group failures by cause, not by symptom.

      4 hours

    • Rank by harm

      Review

      Which failure would you least want screenshotted?

      4 hours

  4. W4

    Guardrails

    Outcome — The top failure class reduced by half, measured on held-out cases.

    14 hrs · 3 tasks
    • Design the mitigation

      Build

      Prompt, tool, validation or refusal — pick deliberately.

      5 hours

    • Implement and measure

      Build

      Held-out set only. No peeking at the training cases.

      6 hours

    • Regression check

      Review

      Confirm you did not break the cases that worked.

      3 hours

  5. W5

    Production readiness

    Outcome — Deployed behind a flag with observability and a kill switch.

    14 hrs · 3 tasks
    • Ship behind a flag

      Build

      Staged rollout, kill switch tested for real.

      5 hours

    • Production telemetry

      Build

      Log inputs, outputs, scores, cost — privacy-safe.

      5 hours

    • Shadow traffic

      Review

      Run silently against live traffic; compare.

      4 hours

  6. W6

    Technical review

    Outcome — A written eval report and a defence to the engineering panel.

    14 hrs · 4 tasks
    • Write the eval report

      Build

      Baseline, deltas, remaining failure modes, cost.

      5 hours

    • Code review

      Review

      Full review with a staff engineer; act on it.

      4 hours

    • Panel defence

      Present

      Present the numbers and what you would do next.

      2 hours

    • Case study writeup

      Build

      Publish to your FutureFit skills passport.

      3 hours

Ready to take the track?

Applications close two weeks before the Autumn 2026 cohort starts.

Apply now