Recommending a skill with Jev: one call, one threshold, 100 questions
One Jev call picks a skill from 141 candidates, or says none applies: 90.8% first-answer correct against a strong generative model's 93.8%, at 1/9 the cost and 3.6× the speed. Two rules came out of the measurement: give the Choice a none option, and judge models are a good fit for routing.