Inside the teaching cell
848 exams. Graded blind. One subject still failing.
Three consoles do the training: a Trainer that scores the model against held-out labels it never sees, an Orchestrator where a human subject-matter expert picks the lessons, and an obstacle course that tests whether the AI can act rather than just answer. These are the real screens from a sealed run — including the cause family that hasn't graduated.
The Curriculum Trainer grades the model against answers it can't see.
Source-of-truth faults — NetBox environment changes plus injected faults — are diagnosed blind and then autograded against a held-out label. Progress is tracked per cause family, because "84% accurate" across a whole domain hides which subjects the model is actually weak at.
NOW GRADING
- student model
- us.deepseek.r1-v1:0
- adapter
- base (prompt-only persona)
- provider · routing
- bedrock · aws
- domain pack · vendor
- aaa_core · all
Source-of-truth faults diagnosed BLIND against a held-out label, then autograded. Tracked by cause family.
848 graded84% accuracy15 mastered (cause family)
the AI masters up to
PhD
overall 84% over 844 exams · next class: graduated (PhD)
teacher agreement: κ 1.00 over 841 sealed
statistically confirmed: Elementary (90% conf) · macro-F1 0.84
The teacher agreement figure is the one to look at. Cohen's kappa of 1.00 across 841 sealed exams means the automatic grader and the human reviewer reached the same verdict every time. If that number drifted, the grading itself would be suspect and every accuracy figure above it would be worth less.
A human picks the lessons. Nothing is auto-signed.
The Orchestrator is the SME's teaching console. The expert selects lessons, sets rounds and delay, and launches. The launched runs still grade blind, and they land in a human sign-off queue rather than being credited automatically — which is what stops the system from grading its own homework.
Lessons are banded by difficulty. Elementary lessons inject a single held-out fault. Middle School lessons inject two competing faults and score three points for correctly picking the primary — a much harder test, because the wrong answer is also present in the evidence.
The SME is part of the teaching cell (teacher · peer · aid). Launched runs grade BLIND and land in the Human Sign-Off queue — never auto-signed.
L1 · Elementary 18 lessons
L2 · Middle School 4 lessons
L3 · High School 6 lessons
Then it has to do the job, not describe it.
Naming a cause family correctly is stage one. The obstacle course tests whether the model can drive read-only tools to gather its own evidence, propose a guarded remediation that a human then approves, and stay inside the no-leakage line when questioned adversarially. The taxonomy itself is sealed at twenty families, and the copilot can propose a new one — but only an SME can admit it.
COURSES
AAA Forensic Investigation LIVE
Drive read-only tools (ReAct) to gather evidence and conclude the cause family.
Stage 1 — read-only investigate
AAA Remediation (guarded action) LIVE
After concluding, propose a guarded remediation; graded vs the playbook; queued for human approval (never auto-run).
Stage 2 — propose fix, human-approved
Diagnoser Panel QA (tool fluency) LIVE
Grade whether the Diagnoser knows every panel — what each proves, which panel for a fault, evidence to cause — plus an adversarial leg that must refuse the wrong panel and refuse a leaked answer-key hint. Gate ≥85% per band.
Stage 1 — tool fluency exam
Code-Understanding (SWE pack) PLANNED
Locate a design fact or injected code change blind; graded against the diff and test.
Stage 1 — inject-and-grade-on-code
CAUSE-FAMILY PROPOSALS — HUMAN APPROVAL (0)
No cause-family proposals in flight. The copilot drafts a NEW family when evidence fits none of the active taxonomy — the SME approves or rejects here.
ACTIVE CAUSE FAMILIES — SME REVIEW (20)
The sealed canonical taxonomy (core 20 · v1.0). Approve confirms the family; Reject flags it for review. Non-destructive — the canonical set is not mutated.
Why we're showing you the failing subject.
authorization_denied · 27% (25/92)
Every other family on the Trainer board is above 95%. This one is at twenty-seven percent after ninety-two exams, and it is sitting in a screenshot on our marketing site because removing it would break the thing the product is for.
A cause family does not ship to live diagnosis until it clears all six graduation gates — breadth, mastery, volume, accuracy, stability and sealed integrity. authorization_denied has not cleared them, so TestPulse will not return it as a confident verdict. It stays in training. The board is what tells you that, and a board that only ever showed green would tell you nothing.
If a vendor shows you an AI accuracy chart with no failures on it, the honest question is what got cropped.
Watch the whole cell run, end to end.
The 4:26 walkthrough covers all three consoles, the SME sign-off, and the sealed adapter running offline on an Acer GN100 with the network unplugged.