New Cohort Starts:

Donate
← Back to the skill map
Reliability & Operations/AI & reliability

LLM Evals

8 micro-topics in Reliability & Operations, each with the evidence that proves you have it and the written reason behind every link. Before you start, it rests on 6 other domains. Downstream, it holds up 4 domains.

8 topics  ·  depth 0–9 of 16  ·  8 internal links  ·  6 in  ·  5 out

Before you start

Load-bearing links first. Each one says what in this domain rests on what outside it, and why.

Delivery & CI/CD

1 link

load-bearing

Model Foundations

1 link

load-bearing

Recognize silent LLM failure rests on Reason about nondeterministic output

Variation is why fluent output can be wrong.

Model Integration

1 link

load-bearing

Use a model as a judge rests on Get structured output

A judge returns a structured score.

Context Engineering

1 link

supporting

Sequence evals by cost rests on Treat the context window as a budget

Both are resource allocation decisions.

Testing Foundations

1 link

supporting

Build an evaluation dataset rests on Reason about the testing pyramid

An eval set is a test suite for a nondeterministic system.

Validation & Pydantic

1 link

supporting

Write deterministic checks rests on Generate JSON Schema from models

Structural checks are the same validation applied to output.

What you will be able to do

In prerequisite order. Each idea names the artifact that closes it, the market demand that put it on the map, and what it stands on.

Depth 0  ·  meta  ·  scaffolded  ·  M22 §22.4

Run human evaluation

Samples output for review where it matters

Evidence
A human review round changes a shipping decision
Market anchor
LLM evaluation
Rests on
supporting

Validate the judge

Human labels are the reference the judge is measured against.

Depth 4  ·  conceptual  ·  unassisted  ·  M22 §22.4

Recognize silent LLM failure

Knows fluent output can be wrong

Evidence
Identifies a confident wrong answer in review
Market anchor
LLM evaluation
Rests on
load-bearing

Reason about nondeterministic output · Model Foundations

Variation is why fluent output can be wrong.

Depth 5  ·  representational  ·  unassisted  ·  M22 §22.4

Build an evaluation dataset

Collects representative cases with expected outcomes

Evidence
Assembles a 20-case set covering known failure modes
Market anchor
LLM evaluation
Rests on
load-bearing

Recognize silent LLM failure

You build a dataset because you cannot trust a single result.

supporting

Reason about the testing pyramid · Testing Foundations

An eval set is a test suite for a nondeterministic system.

Depth 6  ·  procedural  ·  unassisted  ·  M22 §22.4

Write deterministic checks

Asserts on structure, format and constraints

Evidence
Catches malformed output without a model
Market anchor
LLM evaluation
Rests on
load-bearing

Build an evaluation dataset

Checks run over the dataset.

supporting

Generate JSON Schema from models · Validation & Pydantic

Structural checks are the same validation applied to output.

Depth 8  ·  conceptual  ·  scaffolded  ·  M22 §22.4

Use a model as a judge

Scores output with a rubric-driven model

Evidence
Runs a judged eval over a dataset
Market anchor
LLM evaluation
Rests on
load-bearing

Write deterministic checks

You reach for a judge only when cheap checks cannot decide.

load-bearing

Get structured output · Model Integration

A judge returns a structured score.

Depth 9  ·  meta  ·  scaffolded  ·  M22 §22.4

Validate the judge

Checks the judge against human labels

Evidence
Reports judge agreement rates before trusting it
Market anchor
LLM evaluation
Rests on
load-bearing

Use a model as a judge

A judge must be checked before it is trusted.

Depth 9  ·  meta  ·  scaffolded  ·  M22 §22.4

Gate changes on evals

Blocks a regression before it ships

Evidence
A prompt change that lowers the score fails CI
Market anchor
MLOps
Rests on
load-bearing

Use a model as a judge

Gating requires a score.

load-bearing

Automate workflows with CI · Delivery & CI/CD

The gate runs in CI.

Depth 9  ·  meta  ·  scaffolded  ·  M22 §22.4

Sequence evals by cost

Runs cheap checks before expensive ones

Evidence
Most failures are caught before a judge is invoked
Market anchor
LLM evaluation
Rests on
load-bearing

Write deterministic checks

The ladder orders checks by cost.

load-bearing

Use a model as a judge

The expensive rung is the judge.

supporting

Treat the context window as a budget · Context Engineering

Both are resource allocation decisions.

What rests on this domain

Everything downstream that names an idea here as a prerequisite, grouped by where it lives.

Retrieval (RAG)

2 links

load-bearing

Measure retrieval quality rests on Build an evaluation dataset

Metrics are computed over an evaluation set.

load-bearing

Measure generation faithfulness rests on Use a model as a judge

Faithfulness is scored by a judge.

Harness Engineering

1 link

load-bearing

Add an independent checker rests on Write deterministic checks

A checker needs a definition of correct.

AI Safety & Governance

1 link

load-bearing

Measure and mitigate bias rests on Build an evaluation dataset

Fairness metrics run over an evaluation set.

Agents

1 link

supporting

Reflect on results rests on Write deterministic checks

Self-evaluation borrows from how you score output.

Retool. Retrain. Relaunch.

375 ideas. 17 weeks. No tuition, ever.

Vets Who Code is a veteran-run 501(c)(3). The accelerator is free, remote, and we don’t take a share of your first paycheck.

Free · Remote · 17 weeks · EIN 86-2122804