Design it step by step
List the facts the answer needs
Invented query: “Can an X2 earbud bought in Japan receive warranty service in China?” Needed facts include model, sales channel, region rules, and effective date. Retrieval finds candidates; a judgment layer filters evidence; a generator writes the answer. Jev cannot create an absent policy.
Filter by metadata without losing regional evidence
Retain document and passage IDs, URL, version, and text spans. A Chinese query may require a Japanese policy. If regional terms are missing, expand retrieval or report missing evidence instead of inferring eligibility from a general FAQ.
Separate relevance, support, and conflict
Official RAG recipes ask multiple questions about a query–passage pair, then filter in code. Record applicability separately: does the passage concern X2, establish cross-border coverage, or conflict with another applicable policy? A repair address does not prove eligibility.
Generate only covered claims
Invented P1 covers X2 bought from a Japanese direct store at designated Japanese service centers within one year. P2 offers paid repair in mainland China. Neither proves free cross-border warranty. State the known scope and the evidence gap.
Verify individual claims after generation
Check citation IDs and quoted spans in code, then assess whether each passage supports its claim. Split sentences containing region, price, and time claims. An existing citation does not establish support; keep conflicts visible with their dates and scope.
Evaluate stages separately
Fix queries and label the necessary evidence. Measure its presence in candidates, its ranking, unsupported answer claims, and appropriate fallback. Slice by model, region, and language. A filtering threshold cannot repair absent Japanese-source retrieval.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| P1: one-year coverage at Japanese centers | Supports a Japanese service scope and term | State that scope with P1 |
| P2: paid repair in mainland China | Supports paid repair, not free warranty | Separate repair fees from eligibility |
| Draft: free worldwide warranty | Neither source supports worldwide or free | Remove the claim; request confirmation |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
{
"query": "Can a Japan-bought X2 receive warranty service in China?",
"passage": {
"id": "P1",
"source": "invented-policy-for-this-exercise",
"scope": {"model": "X2", "purchase_channel": "Japan direct store"},
"text": "One-year coverage at designated service centers in Japan."
},
"claim_to_check": "X2 has free worldwide warranty service.",
"required_facts": ["purchase_region", "service_region", "fee", "term"]
}Design a decisionCommon mistakes
- Treating topical relevance as full support.
- Mixing versions, countries, channels, or effective dates.
- Accepting an answer because citations are well formed; format and semantic support are different checks.
TRY / THINK / COMPARE
Think first, then compare
A passage says a repair center accepts X2. Does it prove free repair for a Japan-bought device?
Show explanation
No. Model acceptance does not establish region eligibility or price. Retrieve applicable terms or clearly report the gap.
Before handing it over
- Required facts and evidence gaps are explicit.
- Sources retain versions, regions, and text spans.
- Claims map to evidence; conflicts remain visible.
- Retrieval, ranking, and generation failures are logged separately.
Common questions
Does this replace a vector database?
No. Retrieve candidates first, then use judgments for filtering or ranking.
Does strong support guarantee truth?
No. A source may be outdated, wrong, or inapplicable. Check provenance and scope too.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.