评测与研究

jev-ood-calibration

@scienthoon2PythonMIT更新于 2026-09-19

对TypeSafe的Jev在未见过任务上的独立校准测试,包括900条规则生成的支持工单和3个公开基准,提供原始响应、带噪声基线的ECE、温度重拟合及类型误校准符号,成本约0.06美元。

英文原文

Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.

scienthoon/jev-ood-calibration

它在哪儿调用了 Jev

import { experimental_evaluate as evaluate } from 'ai';

scripts/jev_eval.mjs:21

链接指向我们抓取当天的那个 commit,行号是准的。

这个项目没法在这儿跑

它的 question 组合是运行时拼出来的,或者代码里没有直接写出来,所以没法原样搬过来。源码链接在上面,可以自己去看。

同类的其他项目