JJEV Field Guide
Home/Public cases

CASE NOTES / 08

Cases with source code or official evaluations

Speed and outcomes belong to the original authors or TypeSafe. We explain the decision point and limits of reproduction.

Browser use

browser-use/jev-ultrafast

Jev selects an operation and page element from structured browser state; a text model handles typing. The project publishes its own performance notes.

View sourceRelated guide

Game state

fhshaik/typesafe-mario

Structured emulator state is mapped to legal controller choices. This is a game-specific experiment, not a vision benchmark.

View sourceRelated guide

Code review

devagrawal09/jev-review

A local review workflow asks bounded questions about code changes. Review signals still need developer judgment.

View sourceRelated guide

Context

tamaratran/fast-jev-compaction

Scores tool outputs for retention while keeping surviving text verbatim. Evaluate information loss on your own tasks.

View sourceRelated guide

Security incidents

TypeSafe workflow eval: Security incidents

Official evaluation decomposes an incident decision into questions and code rules; labels are model consensus, not independent ground truth.

View sourceRelated guide

Customer service

TypeSafe workflow eval: Customer service

Official workflow illustrates parallel judgments on a conversation and account context. It does not establish results for every support system.

View sourceRelated guide

Invoice processing

TypeSafe workflow eval: Invoice processing

Official workflow combines document judgments with explicit rules; payment authority should remain outside a tutorial model call.

View sourceRelated guide

Agent traces

TypeSafe workflow eval: Agent trace

Official evaluation assesses when a completed agent run needs human review. Different trace formats require their own tests.

View sourceRelated guide

We have not run a Jev API benchmark.