Design it step by step
Separate claims from verified facts
Invented ticket: “The left earbud is silent. I bought it ten days ago and want a replacement.” Keep the user report, verified order date, and diagnostic status separate. A purchase claim does not establish eligibility.
Define three atomic questions
Use Choice for the main issue, Noul for an explicit replacement request, and Score for diagnostic detail using described levels. Issue, preference, and information quality are separate judgments.
Batch only questions with available context
The API supports multiple questions in a request. These three share the ticket state. A result from a channel test is a new fact; wait for that result before judging whether the test passed.
Make relevance a code condition
A sufficiently confident device-issue classification opens diagnostics. Replacement intent is a flag for staff, not permission to replace. Ignore diagnostic scores for shipping issues. Missing fields or API failures go to review.
Compare the complete workflows
Hold tickets and question versions fixed. Compare batching with sequential requests using p50/p95 latency, total cost, errors, and fallback rate. Log model versions, results, and actions; do not claim a published demo as your measurement.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| Device issue; replacement requested; no test | Preference is not eligibility | Open diagnostics; keep the request flag |
| Shipping delay; diagnostic score also returned | The diagnostic score is irrelevant | Use the shipping branch only |
| Unclear issue; order state missing | More questions cannot create facts | Fetch order data or request review |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
{
"model": "jev-latest",
"state": {
"user_report": "Left earbud is silent; I want a replacement.",
"order_verified": true,
"diagnostic_status": "not_tested"
},
"questions": {
"issue": {
"type": "choice",
"instructions": "What is the main issue in the user report?",
"criteria": {
"device": "Device behavior or malfunction",
"shipping": "Delivery, tracking, or missing parcel",
"billing": "Payment or charges",
"other": "None of the above"
}
},
"replacement_requested": {
"type": "noul",
"instructions": "Does the user explicitly request a replacement?"
}
}
}Design a decisionCommon mistakes
- Feeding a predicted category into another question as a fact. Share the original state and compose results in code.
- Using every returned field. Select the branch first, then consume its relevant results.
- Treating one request as one business action. Actions still need permissions, conditions, and failure handling.
TRY / THINK / COMPARE
Think first, then compare
The user says “It may be out of battery or broken.” Can this request determine whether charging fixed it?
Show explanation
No. Classify the issue and information quality now, then obtain the charging-test result. A prediction is not an observation.
Before handing it over
- Each question targets one concept and traceable facts.
- Every answer has a relevance condition.
- Failure, missing fields, and uncertainty have a defined path.
- Cost and latency comparisons use the same data and versions.
Common questions
Is batching always faster?
No. It can reduce round trips, but measure your workload and service limits.
Do questions automatically validate one another?
No. Define consistency constraints and conflict handling explicitly.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.