Benchmarks & research

jev-ood-calibration

@scienthoon2PythonMITupdated 2026-09-19

Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.

scienthoon/jev-ood-calibration

Where it calls Jev

import { experimental_evaluate as evaluate } from 'ai';

scripts/jev_eval.mjs:21

The link points at the commit we read, so the line number still holds.

This one cannot run here

Its question set is assembled at runtime, or never written out literally in the code, so there is nothing to lift. The source link above will show you.

Other projects in this category