Every module produces or corrects a real artifact — a prompt, a rubric, a golden answer, an audit. Every artifact is scored against published rubrics with written feedback. The method is the lesson.
Golden answer first · Reverse-engineer criteria · Atomize (one idea per criterion) · Difficulty-check against a live model · Evaluate blind.
Fabrication · Omission · Misapplication · Presentation. Every QC review maps errors to a family before a word of feedback is written.
"Could a stranger judge this criterion true/false without seeing the prompt, the sources, or any other criterion?" Drilled until reflexive.
How models are trained (pretraining → SFT → RLHF → evals) in plain language; who buys expert data and why; where the money concentrates; honest market risks.
Prompts, input files, golden answers, rubrics, spec sheets; attempter/reviewer/auditor roles; how a package flows through QA.
Realistic persona and context, timelessness, jurisdiction clarity, explicit deliverables, objective gradability. Exercise: repair three broken prompts.
The rubric as the training signal; self-contained criteria; one idea per criterion; verb-first, objective wording. Exercise: five criteria, auto-checked.
Application mechanics, the expert profile, time and rate economics, contractor basics, burnout prevention.
What we never do — and why integrity is the graduate’s competitive asset. Signed as a condition of certification.
Legal and Finance tracks share the spine; each applies it through domain labs.
Tasks that stump strong models legitimately: multi-step dependency chains, document-grounding, plausible red herrings. Legal lab: layperson-voice, jurisdiction-identifiable scenarios. Finance lab: cascading calculations across documents.
Realistic synthetic source material; document-authenticity standards; file hygiene. Graded: a 3-document input set.
The GRADE Loop end-to-end: composition, logical-progression ordering, bounded ranges, the name-swap test, offline evaluability.
Affirmative failure-mode negatives; criticality discipline; weighting so reasoning carries ~75–80% of points. Legal lab: IRAC-decomposed rubrics. Finance lab: formula-vs-solution separation.
Perfect-answer-first; citation integrity verified against primary sources. Graded: a golden answer + rubric pair, peer-graded blind.
SFT vs. RLHF evaluation; strict annotation formats; preference justifications; timed calibration practice.
Reading feedback without ego; the Four Failure Families on your own work; self-review checklists; professional escalation.
Portfolio assembly; platform strategy; rate negotiation with real market data; the attempter → reviewer → auditor ladder. Capstone review.
Verify first, annotate second; independent verification of every claim against primary sources; ambiguity analysis; documenting "considered but not flagged."
Full-package audits: prompt ↔ inputs ↔ golden ↔ rubric consistency; criterion-by-criterion rubric audit; findings vs. preferences; severity discipline.
Consolidated, specific, warm, actionable author-facing feedback; the economics of review throughput.
Three packages audited independently, calibrated against reference audits — inter-rater agreement becomes your credential evidence.
The same spine, re-cut for your context: a diagnostic of your instructions and QA outcomes, custom labs built on your task types (NDA-firewalled), pre/post approval-rate measurement, and an optional train-the-trainer license.
Enterprise programs →A proctored, scenario-based exam and blind portfolio review leading to the "SME Standard Certified" credential — with a public registry and a 2-year maintenance cycle. The strategic goal: platforms fast-track credential holders.
Partner on recognition →