J•JEV Field Guide
Home/Jev composite scoring: rubrics, normalization, and explainable ranking

After the basics · Worked guides

Jev composite scoring: rubrics, normalization, and explainable ranking

Rank tutorial candidates with explicit Score levels, normalization, weights, hard constraints, and missing-value handling.

Source checked 2026-10-0214 minutes
What you will build

Build a ranking rubric that explains why an item ranks first and can be recalculated under different weights.

Before you start · How to choose Choice, Score, or Noul

Design it step by step

01

Define the reader and objective

Invented task: select a first API exercise for someone who has never used an API. Rate actionability, beginner accessibility, and source traceability separately. Define observable evidence for each.

02

Write ordered levels

For actionability: 0=no steps, 1=concepts only, 2=steps with missing input or outcome, 3=input, steps, and expected outcome. Define the other dimensions similarly. Interpret Score against your levels, not an assumed 0–1 scale.

03

Normalize before weighting

All three dimensions here range from 0 to 3, so divide by 3. Illustrative weights are 0.5, 0.3, 0.2. Keep the raw scores and contributions; these weights are an exercise, not an optimum.

04

Apply hard conditions first

Unclear reuse permission or clearly outdated API instructions trigger review before ranking. A high score is not high confidence. Do not turn missing dimensions into zeros; retry, review, or exclude pending data.

05

Test sensitivity and human preference

A=(3,1,2) scores 0.733 and B=(2,3,3) scores 0.833. With weights (0.8,0.1,0.1), A=0.9 and B=0.733, reversing the order. Ask human reviewers for reasons. Changing weights cannot repair incorrect component judgments.

Worked example · Invented by this site

ObservationJudgmentApplication action
A: actionability 3, accessibility 1, sources 20.5×1 + 0.3×⅓ + 0.2×⅔ = 0.733Second under these weights
B: actionability 2, accessibility 3, sources 30.5×⅔ + 0.3×1 + 0.2×1 = 0.833First under these weights
C: accessibility missingMissing is not zero; total is incompleteWait for a complete evaluation

A copyable design draft

Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.

# Offline arithmetic exercise: these are invented ratings.
weights = (0.5, 0.3, 0.2)
ratings = {"A": (3, 1, 2), "B": (2, 3, 3), "C": (2, None, 3)}
assert abs(sum(weights) - 1) < 1e-9

for item_id, values in ratings.items():
    if any(value is None for value in values):
        print(item_id, "pending review")
        continue
    assert all(0 <= value <= 3 for value in values)
    contributions = [w * value / 3 for w, value in zip(weights, values)]
    print(item_id, round(sum(contributions), 3), contributions)
# A: 0.733; B: 0.833; C: pending review
Design a decision

Common mistakes

  1. Mixing raw 0–3 and 0–10 scales lets the larger scale dominate.
  2. Double-counting nearly identical dimensions.
  3. Calling a weighted score a probability of success. It is a preference function, not a calibrated probability.

TRY / THINK / COMPARE

Think first, then compare

A tutorial scores 0.91, but its key API source cannot be accessed. Recommend it immediately?

Show explanation

Review the source first. If traceability is required, a high total cannot compensate. Keep scores for diagnosis but hold the recommendation.

Before handing it over

  • Levels have observable definitions.
  • Ranges and weights are reproducible; weights sum to one.
  • Hard constraints, missing values, and confidence are separate.
  • Compare at least two reasonable weight settings.

Common questions

Is the total a confidence value?

No. It combines rubric ratings under your preferences; confidence describes certainty in a judgment.

How should I choose weights?

Define the reader task and validate against labeled data and failures. There is no universal setting.

Sources and further reading

Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.