Design it step by step
Define the reader and objective
Invented task: select a first API exercise for someone who has never used an API. Rate actionability, beginner accessibility, and source traceability separately. Define observable evidence for each.
Write ordered levels
For actionability: 0=no steps, 1=concepts only, 2=steps with missing input or outcome, 3=input, steps, and expected outcome. Define the other dimensions similarly. Interpret Score against your levels, not an assumed 0–1 scale.
Normalize before weighting
All three dimensions here range from 0 to 3, so divide by 3. Illustrative weights are 0.5, 0.3, 0.2. Keep the raw scores and contributions; these weights are an exercise, not an optimum.
Apply hard conditions first
Unclear reuse permission or clearly outdated API instructions trigger review before ranking. A high score is not high confidence. Do not turn missing dimensions into zeros; retry, review, or exclude pending data.
Test sensitivity and human preference
A=(3,1,2) scores 0.733 and B=(2,3,3) scores 0.833. With weights (0.8,0.1,0.1), A=0.9 and B=0.733, reversing the order. Ask human reviewers for reasons. Changing weights cannot repair incorrect component judgments.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| A: actionability 3, accessibility 1, sources 2 | 0.5×1 + 0.3×⅓ + 0.2×⅔ = 0.733 | Second under these weights |
| B: actionability 2, accessibility 3, sources 3 | 0.5×⅔ + 0.3×1 + 0.2×1 = 0.833 | First under these weights |
| C: accessibility missing | Missing is not zero; total is incomplete | Wait for a complete evaluation |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
# Offline arithmetic exercise: these are invented ratings.
weights = (0.5, 0.3, 0.2)
ratings = {"A": (3, 1, 2), "B": (2, 3, 3), "C": (2, None, 3)}
assert abs(sum(weights) - 1) < 1e-9
for item_id, values in ratings.items():
if any(value is None for value in values):
print(item_id, "pending review")
continue
assert all(0 <= value <= 3 for value in values)
contributions = [w * value / 3 for w, value in zip(weights, values)]
print(item_id, round(sum(contributions), 3), contributions)
# A: 0.733; B: 0.833; C: pending reviewDesign a decisionCommon mistakes
- Mixing raw 0–3 and 0–10 scales lets the larger scale dominate.
- Double-counting nearly identical dimensions.
- Calling a weighted score a probability of success. It is a preference function, not a calibrated probability.
TRY / THINK / COMPARE
Think first, then compare
A tutorial scores 0.91, but its key API source cannot be accessed. Recommend it immediately?
Show explanation
Review the source first. If traceability is required, a high total cannot compensate. Keep scores for diagnosis but hold the recommendation.
Before handing it over
- Levels have observable definitions.
- Ranges and weights are reproducible; weights sum to one.
- Hard constraints, missing values, and confidence are separate.
- Compare at least two reasonable weight settings.
Common questions
Is the total a confidence value?
No. It combines rubric ratings under your preferences; confidence describes certainty in a judgment.
How should I choose weights?
Define the reader task and validate against labeled data and failures. There is no universal setting.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.