Design it step by step
Define labels and actions first
Invented message: “The charge failed, and I also want to stop next month’s renewal.” Decide whether your system needs one owning queue or multiple intents. Write precedence and ambiguity rules before labeling.
Keep three quantities distinct
Option probability, answer confidence, and repeated-label agreement are different quantities. Noul has no confidence field. Record its statement direction: a high risk probability does not authorize an action.
Measure fixed-input variation separately
Hold state, questions, model version, and policy fixed. Test wording or irrelevant-field changes in another experiment. Save distributions, failures, timing, and cache status.
Design abstention before extra calls
Review an ambiguous ticket or clarify renewal intent. Voting can preserve shared bias. Add bounded calls only if validation demonstrates a better error–cost tradeoff.
Evaluate accuracy and coverage together
Invented repeats billing, billing, cancel establish a majority, not truth. On untouched cases, measure errors among automated actions, total coverage, and review burden, sliced by language and ambiguity.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| All repeats say billing; human rule says cancel | Perfect agreement can still be wrong | Retain the error and inspect the task |
| billing, billing, cancel | A majority creates no new fact | Apply the review policy |
| High Noul for a risk statement | Risk probability is not permission | Use the risk-handling branch |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
# Offline example: invented labels, not API results.
from collections import Counter
labels = ["billing", "billing", "cancel"]
majority, count = Counter(labels).most_common(1)[0]
print(majority, count / len(labels))
# Agreement alone does not establish accuracy.
review_required = len(set(labels)) > 1
print("review:", review_required)Design a decisionCommon mistakes
- Equating agreement with accuracy.
- Calling top-option probability confidence.
- Repeating until the desired answer appears.
TRY / THINK / COMPARE
Think first, then compare
Five identical answers on one input: is the system reliable?
Show explanation
Not established. Add ground truth, varied cases, and an independent test set; report the limited observation honestly.
Before handing it over
- Label and ambiguity rules are explicit.
- Distributions, versions, and cache status are recorded.
- Additional calls have a budget cap.
- The test set did not choose the threshold.
Common questions
Do repeats always improve accuracy?
No. Test whether they repair mistakes, not merely stabilize labels.
Is abstention a failure?
It is a designed outcome; assess coverage and review workload too.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.