Inputs
Prompts, responses, traces, policies, professional standards, known failures, current rubrics, expert requirements, and customer platform rules.
Convert professional standards, model failures, and qualified human decisions into demonstrations, preferences, rubrics, and held-out evaluations with a traceable review path.
Each program names the required qualification, work unit, review method, output, and acceptance question before recruitment or production begins.
Prompts, responses, traces, policies, professional standards, known failures, current rubrics, expert requirements, and customer platform rules.
Expert sourcing and verification, demonstrations, reference answers, preferences, rubrics, evaluations, red teaming, and agent workflow grading.
Task sets, judgments, rationales, weighted criteria, graded references, held-out evaluations, failure taxonomies, and qualification manifests.
Production-like qualification, calibration, separated author and reviewer roles, blinded comparison where useful, agreement analysis, and adjudication.
Tasks completed in the approved environment or exported in the buyer's schema with rubric versions, role provenance, and a QA report.
The same standard should connect what teaches the model, what grades it, and what exposes its failures.
Author prompts, reference responses, demonstrations, counterexamples, tool-use traces, and structured rationales when the program requires them.
Create preference pairs, pointwise judgments, error labels, and reasons against stable criteria rather than reviewer taste alone.
Build weighted rubrics, graded examples, evaluator instructions, acceptance rules, and adjudication guidance from buyer standards.
Create held-out tasks, regression sets, policy checks, domain red-team cases, and grading for agent actions or completed work products.
Credentials can open the gate. Production-like work and calibrated review determine who stays in the program.
Name the model behavior, artifact, expert requirement, policy boundary, and acceptance method.
Output: task and expert specificationReview representative prompts, responses, traces, current guidance, and systematic disagreement.
Output: failure and ambiguity mapVerify the agreed signals, run a production-like task, compare decisions, and revise the rubric.
Output: approved reviewer pool and sampleSeparate authoring, review, and adjudication where the risk requires it, then track agreement and corrections.
Output: reviewed expert artifactsCheck schema, task coverage, duplicates, leakage risks, rubric versions, and delivery in the approved environment.
Output: artifact set, QA report, and provenanceExpert disagreement can expose ambiguity in the task, gaps in the rubric, or a real split in professional judgment. It should not be averaged away without inspection.
Representative sample
This example shows the delivery shape. The actual discipline, criteria, evidence, identity policy, and acceptance rule come from the engagement.
Start with one failure cluster and the qualification needed to judge it correctly.
The program defines acceptable evidence, then uses a production-like qualification task and calibration. A credential alone does not prove task performance.
The engagement defines the provenance level. Qualification and review-role evidence can be delivered without exposing personal identity where privacy or safety requires de-identification.
Yes, when the customer environment supports the required access, contributor, review, and audit workflow.
Rubrics are versioned. The team records why criteria change, re-calibrates reviewers after material changes, and keeps production artifacts tied to the version used.
Receive a sample task set, draft rubric, reviewer plan, and pilot price.