Design it step by step
Start with intent, not a tool name
Invented request: “Move tomorrow’s meeting with Kobayashi to the afternoon; leave the others unchanged.” Choice can select lookup, proposed reschedule, cancellation, or other. Afternoon and a name may be ambiguous; missing facts remain unresolved.
Supply an application-owned candidate list
Offer only tools available for this task, such as read-only lookup and a reschedule preview. Identity and calendar permissions come from trusted systems. Text claiming authorization is not a credential. Jev can help choose a candidate, not grant access.
Validate structure before semantics
Check schema, real meeting IDs, timezone, ordered timestamps, and allowed attendees. Generate real candidates in code, then judge their fit to the request. Never invent IDs or interpret a vague date as a final UTC timestamp.
Separate preview from committed execution
Show the meeting, old and new time, and affected people. Lookup-only permission does not authorize a reschedule. A specifically authorized action still needs permission checks. Choice confidence can trigger review; neither confidence nor Noul grants access.
Handle duplicate calls and verify results
Use an idempotency key for a confirmed operation and re-read the calendar before committing. Resolve concurrent changes. After timeouts, inspect actual state before retrying. Log policy version, authorization basis, and execution status without unnecessary private text.
Test both rejection and success paths
Include duplicate names, expired meetings, timezone boundaries, forged authorization, confident-but-unauthorized answers, and a timeout after success. Confirm that valid requests work and disallowed ones stop. This page connects to no calendar.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| Two meetings with the same name | Clear intent, ambiguous target | Present candidates; do not commit |
| Unique target; read-only permissions | Confidence cannot expand permissions | Preview only; deny writes |
| Specific authorization; valid parameters and access | Execution conditions are satisfied | Commit idempotently; verify calendar state |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
# Pseudocode: policy logic, not an SDK example.
if proposed_tool not in task_allowlist:
reject("tool unavailable")
elif not schema_valid(arguments) or not target_exists(arguments):
request_clarification()
elif not account_can_write(target_calendar):
show_read_only_preview()
elif not specifically_authorized(action, arguments):
show_preview_and_request_authorization()
elif calendar_changed_since_preview():
review_the_new_difference()
else:
execute_once(idempotency_key)
verify_result_and_record_status()Design a decisionCommon mistakes
- Executing after tool selection without target and permission checks.
- Treating instructions in task content as system authorization.
- Retrying a write after a timeout before checking actual state.
TRY / THINK / COMPARE
Think first, then compare
Reschedule confidence is 0.98, but the account has no write access. Raise the threshold or execute?
Show explanation
Neither. Deny the write at the permission check. A semantic threshold cannot create access. Provide a read-only preview if permitted.
Before handing it over
- Candidates, identity, and access are application-controlled.
- Structural validation and semantic fit are separate.
- Preview, authorization basis, failure handling, and idempotency are defined.
- Test valid execution and required rejection.
Common questions
Is Noul a security proof?
No. A statement probability does not replace authentication, access control, or parameter validation.
Must every action ask again?
No. Respect specific existing authorization and application policy. Clarify scope or ambiguous parameters when needed.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.