BMI2 Download the beta

Inside the teaching cell

848 exams. Graded blind. One subject still failing.

Three consoles do the training: a Trainer that scores the model against held-out labels it never sees, an Orchestrator where a human subject-matter expert picks the lessons, and an obstacle course that tests whether the AI can act rather than just answer. These are the real screens from a sealed run — including the cause family that hasn't graduated.

console 1 of 3 · run CFG-COA_20260722

The Curriculum Trainer grades the model against answers it can't see.

Source-of-truth faults — NetBox environment changes plus injected faults — are diagnosed blind and then autograded against a held-out label. Progress is tracked per cause family, because "84% accurate" across a whole domain hides which subjects the model is actually weak at.

Curriculum Trainer RADIUSTACACS+ Generate report

NOW GRADING

student model
us.deepseek.r1-v1:0
adapter
base (prompt-only persona)
provider · routing
bedrock · aws
domain pack · vendor
aaa_core · all
Pulser (test cases) · 31 SoT (cause families) · 15 Agentic (walk-thru) · 3

Source-of-truth faults diagnosed BLIND against a held-out label, then autograded. Tracked by cause family.

Certified mastery: PhD · 441 lifetime exams · 84% · 12/12 families

sealed: git 2026-06-30 · SME-sealed (contiguous PhD, all 12 subjects PhD)

848 graded84% accuracy15 mastered (cause family)

auth_semantic_failure99% (166/167)
command_authorization_failure99% (89/90)
policy_path_delay100% (74/74)
resolver_path_degradation100% (73/73)
endpoint_network_readiness_delay100% (69/69)
timing_budget_violation100% (63/63)
dhcp_path_degradation100% (33/33)
authorization_denied27% (25/92)
directory_path_degradation100% (18/18)
accounting_path_failure95% (18/19)
transport_degradation95% (18/19)

the AI masters up to

PhD

overall 84% over 844 exams · next class: graduated (PhD)
teacher agreement: κ 1.00 over 841 sealed
statistically confirmed: Elementary (90% conf) · macro-F1 0.84

The teacher agreement figure is the one to look at. Cohen's kappa of 1.00 across 841 sealed exams means the automatic grader and the human reviewer reached the same verdict every time. If that number drifted, the grading itself would be suspect and every accuracy figure above it would be worth less.

console 2 of 3

A human picks the lessons. Nothing is auto-signed.

The Orchestrator is the SME's teaching console. The expert selects lessons, sets rounds and delay, and launches. The launched runs still grade blind, and they land in a human sign-off queue rather than being credited automatically — which is what stops the system from grading its own homework.

Lessons are banded by difficulty. Elementary lessons inject a single held-out fault. Middle School lessons inject two competing faults and score three points for correctly picking the primary — a much harder test, because the wrong answer is also present in the evidence.

Curriculum Orchestrator SME teaching console — pick lessons, launch, test the AI

The SME is part of the teaching cell (teacher · peer · aid). Launched runs grade BLIND and land in the Human Sign-Off queue — never auto-signed.

0 selected rounds 1 delay 12 s ▶ Launch lessons dry-runclear

L1 · Elementary 18 lessons

auth_semanticsingletc AUTH-SEM · score 1 · held-out fault: auth_semantic_failure
auth_deniedsingletc AUTHZ-DENY · score 1 · held-out fault: authorization_denied
dynamic_authsingletc CFG-COA · score 1 · held-out fault: dynamic_authorization_path_degradation
policysingletc CFG-POLICY · score 1 · held-out fault: policy_path_delay
accountingsingletc CFG-ACCT · score 1 · held-out fault: accounting_path_failure
transportsingletc TAC-05D · score 1 · held-out fault: transport_degradation
ldap_serversingletc CFG-LDAP · score 1 · held-out fault: directory_path_degradation
nassingletc VL-E2E-01 · score 1 · held-out fault: nas_or_relay_instability

L2 · Middle School 4 lessons

compete_acct_policycompetingtc CFG-ACCT · score 3 · gold accounting_path_failure vs competing policy_path_delay
compete_dns_dhcpcompetingtc CFG-DNS · score 3 · gold resolver_path_degradation vs competing dhcp_path_degradation
compete_tac_authz_transportcompetingtc TAC-02 · score 3 · gold command_authorization_failure vs competing transport_degradation

L3 · High School 6 lessons

console 3 of 3

Then it has to do the job, not describe it.

Naming a cause family correctly is stage one. The obstacle course tests whether the model can drive read-only tools to gather its own evidence, propose a guarded remediation that a human then approves, and stay inside the no-leakage line when questioned adversarially. The taxonomy itself is sealed at twenty families, and the copilot can propose a new one — but only an SME can admit it.

Agentic Curriculum — Obstacle Course refresh

COURSES

AAA Forensic Investigation LIVE

Drive read-only tools (ReAct) to gather evidence and conclude the cause family.

Stage 1 — read-only investigate

AAA Remediation (guarded action) LIVE

After concluding, propose a guarded remediation; graded vs the playbook; queued for human approval (never auto-run).

Stage 2 — propose fix, human-approved

Diagnoser Panel QA (tool fluency) LIVE

Grade whether the Diagnoser knows every panel — what each proves, which panel for a fault, evidence to cause — plus an adversarial leg that must refuse the wrong panel and refuse a leaked answer-key hint. Gate ≥85% per band.

Stage 1 — tool fluency exam

Code-Understanding (SWE pack) PLANNED

Locate a design fact or injected code change blind; graded against the diff and test.

Stage 1 — inject-and-grade-on-code

CAUSE-FAMILY PROPOSALS — HUMAN APPROVAL (0)

No cause-family proposals in flight. The copilot drafts a NEW family when evidence fits none of the active taxonomy — the SME approves or rejects here.

ACTIVE CAUSE FAMILIES — SME REVIEW (20)

The sealed canonical taxonomy (core 20 · v1.0). Approve confirms the family; Reject flags it for review. Non-destructive — the canonical set is not mutated.

accounting_path_failure✓ approvedApproveReject
auth_semantic_failure✓ approvedApproveReject
authorization_denied✓ approvedApproveReject
certificate_status_uncertainty✓ approvedApproveReject
directory_path_degradation✓ approvedApproveReject
…15 more

Why we're showing you the failing subject.

authorization_denied · 27% (25/92)

Every other family on the Trainer board is above 95%. This one is at twenty-seven percent after ninety-two exams, and it is sitting in a screenshot on our marketing site because removing it would break the thing the product is for.

A cause family does not ship to live diagnosis until it clears all six graduation gates — breadth, mastery, volume, accuracy, stability and sealed integrity. authorization_denied has not cleared them, so TestPulse will not return it as a confident verdict. It stays in training. The board is what tells you that, and a board that only ever showed green would tell you nothing.

If a vendor shows you an AI accuracy chart with no failures on it, the honest question is what got cropped.

Watch the whole cell run, end to end.

The 4:26 walkthrough covers all three consoles, the SME sign-off, and the sealed adapter running offline on an Acer GN100 with the network unplugged.