Jev.aitools.fyi

Command Palette

Search for a command to run...

Product profile
Python24 starsUpdated 2026-10-07

Decision Index

Reproduce the public Decision Index 0.3 suite locally or as one Hugging Face Job: 37 chance-corrected benchmarks, HTTP /v1/systemone engines, resumable runs.

Overview

decision-index (github.com/apolinario/decision-index, MIT, PyPI decision-index 0.3) is apolinario's reproduction kit for the Decision Index leaderboard for typed decision engines. Inputs are state plus Choice or Noul questions with explicit criteria; outputs are one probability per option. The live board edition described in the README is 0.3: Full score is 20% public benchmarks from this kit, 50% private same-skill tests, and 30% private new-domain tasks run only by maintainers. This repository rebuilds and runs the public index: 37 benchmarks in five areas, chance-corrected and coverage-adjusted, with about 7 GB of pinned downloads and no redistributed row files. README states not affiliated with TypeSafe AI. The project's own sources describe the board but do not publish a separate live board URL, so this listing links the GitHub repo only.

Problem: Teams cite Decision Index ranks in README tables but cannot reproduce the public suite or score their own /v1/systemone server with the same math.

Built for: Benchmark authors and open decision model maintainers who need resumable runs, chance-corrected area scores, and parity checks against edition 0.3 public index rules.

First indexed on Jev Directory: 2026-09-22

Creator and team

Name
apolinario
Handle
@apolinario

How Jev is used

Role in the product flow
Batch evaluation harness calling engines via transformers or HTTP POST /v1/systemone with frozen prompts
Primitives
ChoiceNoul
State in
Frozen suite rows with state blobs and criteria per benchmark; engines must not truncate or drop options.
Decision out
Per-row probabilities scored into chance-corrected benchmark skills and a weighted public index.
  1. pip install -e .[transformers,rebuild] and python -m decision_index suite rebuild
  2. suite import staged rows for edition 0.3 with hash verification
  3. python -m decision_index pipeline --engine http --option base_url=... for your server
  4. python -m decision_index score --edition 0.3 on results.jsonl

This kit scores hosted Jev, Kev, Mapika decider, and any engine that implements System One wire format. iammrduncan-typesafe-ai-benchmark and other single-project eval repos measure one author's harness; Decision Index is the cross-model public suite Kev and decider README rows cite. mapika-decider lists Decision Index ranks in its Standing section. posthog-jeeves publishes separate JevBench tables. Use the http engine when your model already exposes POST /v1/systemone locally.

Sourced performance claims

  • README parity: tests/test_index021.py checks math against published entrants including Jev at 57.89 on edition 0.2.1 public index; tests/test_index03.py checks 0.3 public-index math for 11 entrants including Jev; score --edition 0.3 over a full lab run reproduces Cloudflare clef public index 61.71.Source: github.com/apolinario/decision-index README Parity section
  • README rules: models whose median, mean, or 80th-percentile latency exceeds 1,000 ms per request on maintainer RTX PRO 6000 measurement are not added to the board as Jev-like.Source: github.com/apolinario/decision-index README Quickstart
  • GitHub repo apolinario/decision-index had 24 stars and 57 forks on 2026-10-07, last pushed 2026-10-07 per GitHub API.Source: GitHub API, 2026-10-07

Features and stack

Features

  • Edition flags 0.3, 0.2.1, 0.2, and 0.1 with pinned manifests in hub/
  • Checkpoint resume in results.jsonl with error retries
  • Reference transformers and http engines plus Hugging Face Job helper
  • Chance correction, gold star weights, and GSM8K rebuild in 0.3
  • Submission flow via submissions/README.md pull request

Stack

  • Python 3.10+
  • PyTorch optional for transformers engine
  • Hugging Face Hub downloads for suite rebuild
  • HTTP System One servers for open models

Pricing: MIT kit; you pay for download bandwidth, GPUs, and Hugging Face Jobs if used.

FAQ

Where is the live leaderboard URL?
This listing checked the README and repository homepage fields on 2026-10-07. They describe Decision Index 0.3 and Full score rules but do not name a standalone board URL, so only the GitHub repo is linked here.
Does this compute the Full score?
No. The kit runs and scores the public 20% portion. Private halves are run by maintainers after you submit complete public runs per submissions/README.md.
What Jev score should I expect on 0.2.1 vs 0.3?
README parity tests cite Jev at 57.89 on the 0.2.1 public index. For 0.3, the README says lab rescoring reproduces Cloudflare clef at 61.71 public index, not a single Jev headline number in the intro.

Related learn guides

Original Jev guidance that pairs with this product pattern.

  • System One model

    The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.

  • Jev vs LLM classification

    When to gate with System One probabilities instead of asking a chat model to label things.

Related products

Hand-picked neighbors with rich profiles or overlapping tags.

IntegrationsPython702
Demo
Mapika open System One family on Qwen3.5: one forward pass, Choice/Score/Noul, POST /v1/systemone via decider-ai on PyPI, weights on Hugging Face.
integrationsopen-sourcedecision-model
IntegrationsPython6,740
Demo
Jared Palmer's open Jev-like Qwen decision models: System One-compatible serve, HF weights 0.8B to 27B, eval tables vs hosted Jev.
integrationsopen-sourceself-hosted
IntegrationsPython414
PostHog's open 9B Jev-like model that reasons before it scores options, with CISPO training, a block-4 drafter, and Jev-compatible POST /v1/systemone.
integrationsopen-sourcedecision-model
BenchmarksPython3,844
DemoFeatured
Semantic ifs on a home 3090: open models doing structured branches without pretending to be TypeSafe.
benchmarksopen-sourceself-hosted
By Rishit