Jev.aitools.fyi

Command Palette

Search for a command to run...

Product profile
Python606 starsUpdated 2026-09-25

TypeLLM

SGLang JSON Schema generation on existing AR LLMs: enums, numbers, depends_on graphs, JevBench 228/231 with thinking. Not a System One server.

Overview

TypeLLM (github.com/TypeLLM/TypeLLM, Apache-2.0, typellm.ai) adds type-safe generation on top of existing autoregressive weights via SGLang constrained decoding and JSON Schema. pip install typellm exposes TypeLLMClient against an HTTP SGLang endpoint: string, integer, number, boolean, and enum fields with optional thinking budgets, image input for VLMs, depends_on dependency graphs with shared-prefix KV reuse, and permutation averaging on enum questions. It is inspired by TypeSafe Jev interface ideas but is not a one-step decision head: you keep native generation while outputs stay in schema. Public JevBench evals report 195/231 tasks without thinking and 228/231 with thinking on 231 public tasks (see evals/jevbench in the repo).

Problem: Autoregressive LLMs excel at prose but leak invalid JSON, wrong enum labels, and option-order bias when you force them to behave like classifiers.

Built for: Engineers already serving open models on SGLang who want schema-guaranteed fields without swapping to a dedicated System One checkpoint.

First indexed on Jev Directory: 2026-09-25

Creator and team

Name
TypeLLM
Handle
@TypeLLM

How Jev is used

Role in the product flow
Constrained autoregressive decode over shared context instead of POST /v1/systemone on a decision model
Primitives
ChoiceScoreNoul
State in
Free-form context string (plus optional images for vision models) and a questions map with JSON Schema types, instructions, enums, and depends_on edges per typellm.ai docs.
Decision out
Schema-valid values and enum distributions; JevBench-style tasks map enums to choices and numeric rubrics to scores without answer-token prose.
  1. Serve a compatible model with SGLang prefix caching enabled
  2. pip install -U typellm and point TypeLLMClient at the SGLang base URL
  3. Declare questions with types, optional thinking, and dependency graphs
  4. Run generate(); reuse KV for shared prefixes and optional permutation averaging on enums

TypeLLM competes in the same mental lane as typed decisions but through constrained AR on models you already host, not a frozen decision checkpoint. nokia-applied-research-anyjev and featherless-simple-jev chase logit-native or head-fitted probabilities on vLLM; githubnext-localjev fakes System One with chat JSON. TypeLLM is for teams that refuse a second model class yet still want JevBench-shaped guarantees on enums and rubrics. Read evals/jevbench/METHOD.md before you equate bench accuracy with production calibration on your tickets.

Sourced performance claims

  • JevBench README reports 195/231 public tasks correct without thinking and 228/231 with thinking on Qwen3.8-27B configuration documented in evals/jevbench.Source: github.com/TypeLLM/TypeLLM evals/jevbench/README.md
  • README lists negligible output-token cost for categorical fields, depends_on graphs, image input, and permutation averaging blog at typellm.ai/blog/fair-die.Source: github.com/TypeLLM/TypeLLM README
  • Public GitHub repo TypeLLM/TypeLLM had about 606 stars when this listing was drafted.Source: GitHub API September 2026

Features and stack

Features

  • TypeLLMClient over SGLang with string, int, number, boolean, and enum outputs
  • depends_on graphs and shared-prefix reuse documented September 2026
  • Optional thinking mode and vision image input for Qwen VL checkpoints
  • Permutation averaging for fairer enum distributions
  • Published JevBench per-task answers and METHOD.md in the repo

Stack

  • Python
  • SGLang
  • JSON Schema
  • typellm PyPI package

Pricing: Open source Apache-2.0 client; you pay for your own SGLang GPU hosting.

FAQ

Is TypeLLM a System One server?
No. It constrains autoregressive generation via SGLang. You do not get POST /v1/systemone unless you wrap it yourself.
How does it relate to JevBench?
The repo ships a full JevBench run with per-task logs. Scores measure schema-typed answers on public tasks, not hosted api.typesafe.ai latency.
Do I need to retrain my LLM?
README emphasizes no architecture or weight changes; compatibility depends on serving the base model through SGLang with the documented features.

Related learn guides

Original Jev guidance that pairs with this product pattern.

  • System One model

    The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.

  • Jev vs LLM classification

    When to gate with System One probabilities instead of asking a chat model to label things.

Related products

Hand-picked neighbors with rich profiles or overlapping tags.

Agent toolingPython601
Nokia research: Decider API turns hub LLMs into calibrated Jev-style choices with L0 flip fixes and optional L2 heads on vLLM.
agent-toolingopen-sourcedecision-model
IntegrationsPython484
Demo
Simple Jev (github.
integrationsclassifierhuggingface
Agent toolingTypeScript771
GitHub Next Bun server: local POST /v1/systemone over oMLX DiffusionGemma chat JSON, TypeSafe SDK compatible, README honest on calibration.
agent-toolingself-hostedsystem-one
BenchmarksPython3,844
DemoFeatured
Semantic ifs on a home 3090: open models doing structured branches without pretending to be TypeSafe.
benchmarksopen-sourceself-hosted
By Rishit