Open· Last updated

Public TypeSafe evals

The launch post describes workflow evals: a fixed compute graph (“workflow”), with large external models as reference probabilities. TypeSafe says Jev sits on a Pareto frontier in those plots and that homepage multiples (they mention 193.6× faster and 444.6× cheaper) come from this work and are likely high-end real-world gains.

They also list caveats in the same article: workflows were authored by people on their capabilities team; the reference mix (they name large OpenAI and Anthropic models) may bias the comparison; LLM baselines use TypeSafe’s System One wrapper.

They invite readers to a workflow evals site for examples, disagreements, full queries, and each workflow.

How jev.pro will treat evals

If you run your own harness, publish your labels and prompts. Do not claim they are TypeSafe official numbers.

Sources

Public TypeSafe or adjacent documentation only. No private claims.