Decision Index
Reproduce the public Decision Index 0.3 suite locally or as one Hugging Face Job: 37 chance-corrected benchmarks, HTTP /v1/systemone engines, resumable runs.
Overview
decision-index (github.com/apolinario/decision-index, MIT, PyPI decision-index 0.3) is apolinario's reproduction kit for the Decision Index leaderboard for typed decision engines. Inputs are state plus Choice or Noul questions with explicit criteria; outputs are one probability per option. The live board edition described in the README is 0.3: Full score is 20% public benchmarks from this kit, 50% private same-skill tests, and 30% private new-domain tasks run only by maintainers. This repository rebuilds and runs the public index: 37 benchmarks in five areas, chance-corrected and coverage-adjusted, with about 7 GB of pinned downloads and no redistributed row files. README states not affiliated with TypeSafe AI. The project's own sources describe the board but do not publish a separate live board URL, so this listing links the GitHub repo only.
Problem: Teams cite Decision Index ranks in README tables but cannot reproduce the public suite or score their own /v1/systemone server with the same math.
Built for: Benchmark authors and open decision model maintainers who need resumable runs, chance-corrected area scores, and parity checks against edition 0.3 public index rules.
First indexed on Jev Directory: 2026-09-22
Creator and team
- Name
- apolinario
- Handle
- @apolinario
How Jev is used
- Role in the product flow
- Batch evaluation harness calling engines via transformers or HTTP POST /v1/systemone with frozen prompts
- Primitives
- ChoiceNoul
- State in
- Frozen suite rows with state blobs and criteria per benchmark; engines must not truncate or drop options.
- Decision out
- Per-row probabilities scored into chance-corrected benchmark skills and a weighted public index.
- pip install -e .[transformers,rebuild] and python -m decision_index suite rebuild
- suite import staged rows for edition 0.3 with hash verification
- python -m decision_index pipeline --engine http --option base_url=... for your server
- python -m decision_index score --edition 0.3 on results.jsonl
This kit scores hosted Jev, Kev, Mapika decider, and any engine that implements System One wire format. iammrduncan-typesafe-ai-benchmark and other single-project eval repos measure one author's harness; Decision Index is the cross-model public suite Kev and decider README rows cite. mapika-decider lists Decision Index ranks in its Standing section. posthog-jeeves publishes separate JevBench tables. Use the http engine when your model already exposes POST /v1/systemone locally.
Sourced performance claims
- README parity: tests/test_index021.py checks math against published entrants including Jev at 57.89 on edition 0.2.1 public index; tests/test_index03.py checks 0.3 public-index math for 11 entrants including Jev; score --edition 0.3 over a full lab run reproduces Cloudflare clef public index 61.71.Source: github.com/apolinario/decision-index README Parity section
- README rules: models whose median, mean, or 80th-percentile latency exceeds 1,000 ms per request on maintainer RTX PRO 6000 measurement are not added to the board as Jev-like.Source: github.com/apolinario/decision-index README Quickstart
- GitHub repo apolinario/decision-index had 24 stars and 57 forks on 2026-10-07, last pushed 2026-10-07 per GitHub API.Source: GitHub API, 2026-10-07
Features and stack
Features
- Edition flags 0.3, 0.2.1, 0.2, and 0.1 with pinned manifests in hub/
- Checkpoint resume in results.jsonl with error retries
- Reference transformers and http engines plus Hugging Face Job helper
- Chance correction, gold star weights, and GSM8K rebuild in 0.3
- Submission flow via submissions/README.md pull request
Stack
- Python 3.10+
- PyTorch optional for transformers engine
- Hugging Face Hub downloads for suite rebuild
- HTTP System One servers for open models
Pricing: MIT kit; you pay for download bandwidth, GPUs, and Hugging Face Jobs if used.
Links
FAQ
- Where is the live leaderboard URL?
- This listing checked the README and repository homepage fields on 2026-10-07. They describe Decision Index 0.3 and Full score rules but do not name a standalone board URL, so only the GitHub repo is linked here.
- Does this compute the Full score?
- No. The kit runs and scores the public 20% portion. Private halves are run by maintainers after you submit complete public runs per submissions/README.md.
- What Jev score should I expect on 0.2.1 vs 0.3?
- README parity tests cite Jev at 57.89 on the 0.2.1 public index. For 0.3, the README says lab rescoring reproduces Cloudflare clef at 61.71 public index, not a single Jev headline number in the intro.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev vs LLM classification
When to gate with System One probabilities instead of asking a chat model to label things.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.