Jev: Typed Decision Gates for Agent Safety
A fast reflex
Jev is a low-cost decision model, around 420 ms per call, that returns typed probabilities rather than prose. It classifies command action and blast radius before an agent tool call; the reasoning model handles the work that passes the gate.
Three command outcomes
The provisional bands are safe ≥0.85: RUN silently; 0.5–0.85: CONFIRM with a human; and safe <0.5: BLOCK. The related prompt gate either runs a high-quality prompt or proposes a context-grounded rewrite and waits for approval. Decisions and prompt versions are recorded in JSONL.
The miss that matters
A destructive command scored 0.54 in shadow data, landing in CONFIRM rather than BLOCK. This is a reviewable miss, not evidence the gate is calibrated. Logs, human override, and a review of roughly 20 prompts are needed before treating the thresholds as reliable.
How the pieces connect
Explore the original interactive diagram. Its controls include guided views, search, trace, and theme switching. On a narrow screen, use the diagram controls or open it full size.
More notes will be added as this work develops.