// HACKER NEWS — CYBERSECURITY
Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key.
If you classify text with Jev, every answer is a network call to one vendor and comes back in about 300 ms, at any load. For a batch job that is fine. For an agent loop that decides, acts, and decides again, or a game tick, or anything that classifies then acts, 300 ms per step is the whole budget.
Jevstiller sits in front of that call, learns a small local model from Jev’s own answers, and lets it answer what it is sure about in about 15 ms on a CPU. The interesting part is not the small model. It is the contract:
Set one number, say 98%. Jevstiller returns the label Jev would have returned on at least that share of requests.
This post is about what it takes to make that sentence true, why the obvious way of picking a confidence threshold does not make it true, and what it costs.
Per task, over a window of traffic, call c the share of requests the local model answers (coverage), and e the share of those where its label differs from Jev’s. Requests it does not answer go to Jev and agree with Jev by definition. So the system’s agreement with Jev is
Your target A* gives a budget β = 1 − A*: the share of all requests that may end up with an answer Jev would not have given. At 98%, that is 2 in 100. The router’s job is to answer as much as it can while keeping c · e under β.
Two things are deliberately absent from that sentence. It says nothing about the model’s accuracy against the truth: if Jev is wrong, the local model is wrong the same way, and the status report says so next to every number. And it is a statement about agreement over all requests, not about the local model’s accuracy on the requests it chose to answer. The first framing is what you can verify; the second is what most tools report.
The usual recipe for a cascade like this: hold out some data, sweep a confidence threshold, keep the loosest one whose measured disagreement is within budget, ship it. It is what you would write in an afternoon, and it is what the point-estimate rule in our benchmark does.
Then it breaks the budget about half the time. On five public tasks, twenty random train/calibration/test splits each, the point-estimate rule exceeded the 2% budget on 6 to 12 of 20 splits per task, by up to a full percentage point: