Benchmarks & research

jev-korean-benchmark

@mahlernim5Pythonupdated 2026-09-17

Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence

mahlernim/jev-korean-benchmark

Where it calls Jev

from typesafe_sdk import Choice, RetryPolicy, TypeSafeClient

jevbench/medqa_run.py:10

The link points at the commit we read, so the line number still holds.

This one cannot run here

Its question set is assembled at runtime, or never written out literally in the code, so there is nothing to lift. The source link above will show you.

Other projects in this category