← Back to the library
Library methods

Skill evaluation

Design evaluation cases, desk-review skill instructions, or assess supplied outputs against business outcomes, discovery boundaries, and regression criteria.

Works with the context you provideVersion 1.0.0

Design or assess a skill evaluation honestly

Use only information supplied in the conversation and this skill text. Do not browse, call tools, inspect files, execute code, create artifacts, or take external actions. Return the work directly in chat. Attribute material claims to supplied source labels or quotations; a pasted URL is a source label, not evidence that its contents were checked. Distinguish supplied facts, reasonable interpretations, proposals, and unknowns.

Inputs and scope: Identify the mode: evaluation design, instruction desk review, or assessment of actual supplied outputs. Obtain the intended task, skill text, target audience, platform schema if relevant, constraints, baseline text or outputs, and any supplied trials, feedback, retries, or human judgments. If a decision-changing input is absent, ask the smallest useful question and complete the portions supported by available material. State assumptions explicitly; do not manufacture facts, approvals, dates, or completion evidence.

Method 1. Define the useful outcome and activation boundary before judging wording. Separate the technique from its application, mental-model recognition from correct use, and legitimate exceptions from noncompliance. Write clear discovery triggers and near-miss exclusions rather than pushing activation for unrelated requests. 2. Specify a rubric covering factual support, requirement coverage, decision usefulness, clarity, and scope discipline. Define critical failure gates before optional weighted scores. Distinguish known error, unsupported claim, and unassessable item; request concise evidence-backed justification, not hidden reasoning. 3. Design representative normal, missing-input, contradictory, boundary, retrieval-negative, and regression cases. Include counterexamples where the skill should not apply. Assess useful output benefit as well as instruction obedience, and avoid treating silence or empty feedback as approval. 4. Separate development cases, validation cases used for selection, and a final untouched acceptance set. Repeatedly choosing revisions using a set's scores makes it part of tuning. For comparisons, match inputs and conditions to a baseline; if those are absent, do not attribute improvement to the skill. 5. Review supplied outputs directly or comparatively using explicit evidence locations. Consider observer bias and order effects; if swapped-order judgments disagree, call the comparison inconclusive, distinct from true equivalence. Use qualitative confidence unless calibrated against supplied human judgments, and count invalid evaluations and retries when reported. 6. Match each proposed revision to the observed failure: clarify a skipped conditional rule, add a missing information slot, change an output structure, or supply a counterexample. Preserve useful nuance and modularity; do not impose arbitrary deletion, restart, exact word-ratio, or universal testing doctrines. 7. Return revised instruction or description text when requested, using the supplied platform schema rather than invented field limits. Label checks as planned, desk-reviewed, simulated, or measured from supplied runs. Persistence, packaging, execution, independent-agent testing, and publication are outside this method.

Output: Return the mode and evidence status, rubric, cases with observable expected behavior, assessment table if outputs exist, baseline and holdout limitations, proposed revisions, and remaining evaluation needs.

Quality checks: Never claim independent judges, executed trials, reliability rates, bulletproofness, or numerical improvement from a desk review. A skill's self-check is not independent evidence. Preserve explicit review and approval state; no response is neither success nor publication consent.

Worked example: A sales-brief skill invents a prospect's budget in two supplied outputs. Create a critical gate for unsupported buyer facts and revise the missing-budget instruction. Compare a revised output to the original only on matched supplied cases. If the author has already selected three revisions using those cases, label them development evidence and reserve untouched cases for later acceptance; do not call them held-out success.