Design it step by step
Define scope and access first
Invented assistant answers public product questions, not other customers’ orders. Authentication and access rules decide available data. Semantic screening adds signals rather than replacing permissions.
Run deterministic checks first
Validate size, type, permissions, and required fields in code. Label user, retrieval, and output sources. “Ignore the rules” inside a retrieved article is content, not system policy.
Separate questions from actions
Noul can assess requests for others’ data or rule manipulation; distinguish discussion from execution. Score can rate a described impact rubric. Code maps signals to versioned actions. No production thresholds are supplied here.
Check outputs without promising perfect safety
A normal prompt can still yield leakage or unsupported claims. Check outputs before display and remove fields, regenerate, or review as needed. Detector failures have explicit handling, not a default safe label.
Evaluate false blocks and misses
Include legitimate explanations of prompt injection and forged authorization in retrieved text. Measure harmful passes and benign blocks by language, source, and policy version; provide recovery paths and minimize private logs.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| Question asks for a definition of prompt injection | Discussion, not an unauthorized action | Provide a public explanation |
| Retrieved text says export all orders | No valid authorization | Do not execute |
| Detector call fails | Unknown, not safe | Use the failure policy |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
# Policy contract, not a tested detector.
{
"policy_version": "public-support-v1",
"allowed_sources": ["public_product_docs"],
"check_surfaces": ["user_input", "retrieved_text", "model_output"],
"possible_actions": ["pass", "review", "block"],
"detector_error_action": "review",
"access_control_owner": "application_code"
}Design a decisionCommon mistakes
- Treating probability as a security certificate.
- Checking inputs but never outputs.
- Blocking every educational mention of a risk topic.
TRY / THINK / COMPARE
Think first, then compare
A retrieved article claims admin approval to read other orders. Allow access?
Show explanation
No. Text cannot alter identity or access rights; enforce permissions in code.
Before handing it over
- Access control and semantic checks separate.
- Discussion and action distinguishable.
- Inputs and outputs have failure handling.
- False blocks, misses, and fallback recorded.
Common questions
Can guardrails guarantee no bypass?
No. Use real tests, restricted access, and ongoing evaluation.
Why separate signals and policy?
It makes versions, costs, thresholds, and review routes explicit.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.