AnyJev
Nokia research: Decider API turns hub LLMs into calibrated Jev-style choices with L0 flip fixes and optional L2 heads on vLLM.
Overview
AnyJev (github.com/nokia-applied-research/AnyJev, Apache-2.0, PyPI anyjev 0.0.2) from Nokia Applied Research with Tencent Hunyuan coauthors turns hub models into typed deciders: `Question.choice`, yes/no, and score rubrics return distributions you can threshold. Levels raw, L0, L1, and L2 add position-bias fixes and calibration; L2 fits a closed-form head on hidden states (often after `python -m anyjev.truncate`). README banner on BANKING77 cites order-flip rate 0.230 to 0.073 at L0 with zero labels and auto-decidable share at 5% error rising from 7.7% raw to 52.0% with labels. `python -m anyjev.pipeline` truncates, serves, fits, and measures accuracy, ECE, and latency on your machine.
Problem: Raw next-token logits from general LLMs flip when you shuffle options and lie about confidence, so automation thresholds become roulette.
Built for: Teams serving open weights on vLLM or Hugging Face who want Jev-style typed decisions with real probabilities and optional 100 to 300 label head fits, without a full Qwen fine-tune.
First indexed on Jev Directory: 2026-09-25
Creator and team
- Name
- Nokia Applied Research
- Organization
- Nokia
- Handle
- @nokia-applied-research
How Jev is used
- Role in the product flow
- Decider over vLLM or HF backends: cyclic option shifts at L0, temperature at L1, linear head on hidden states at L2
- Primitives
- ChoiceNoulScore
- State in
- Plain-text state strings plus `Question` definitions with named criteria lists or rubric levels per README Python API.
- Decision out
- Per-question probability maps (for example route.distribution on billing vs technical) with documented calibration metrics on held-out sets.
- pip install anyjev[hf] and optionally truncate blocks with python -m anyjev.truncate
- Serve with vLLM embed pooler for L2 or generate task for raw/L0/L1
- Decider.fit_head on 100 to 300 labels for L2, or run zero-label L0
- decide() on live traffic; optional unlabelled maintenance per README routing docs
AnyJev is the research-grade open stack when you refuse Jared Palmer-sized fine-tunes but still want probabilities that survive option reordering. jaredpalmer-kev and togethercomputer-tev1 bet on trained Qwen checkpoints; featherless-simple-jev assembles logits from arbitrary HF models without Nokia's L0 rotation math. wfzyx-von is a purpose-built encoder; AnyJev wraps models you already host. theoleecj-semif chases full System One servers; AnyJev's Decider API targets pipeline operators measuring ECE on BANKING77 before they wire agents. Cross-link githubnext-localjev for Mac bridges that fake System One with chat JSON instead of logit reads.
Sourced performance claims
- README BANKING77 table lists raw order-flip 0.230 vs L0 0.073 with zero labels, ECE 0.240 raw vs 0.095 L1, and auto-decidable at 5% error 7.7% raw vs 52.0% with labels.Source: github.com/nokia-applied-research/AnyJev README banner and results table
- README states L2 head fit on 100 to 300 labels is a closed-form solve with no gradient updates to base weights.Source: github.com/nokia-applied-research/AnyJev README With labels L2
- Public GitHub repo nokia-applied-research/AnyJev had 601 stars and 84 forks when this listing was drafted.Source: GitHub API September 2026
Features and stack
Features
- Decider + Question.choice Python API with vLLM and HF backends
- L0/L1/L2 levels documented with flip-rate and calibration tables
- truncate and pipeline CLIs for serve-measure loops on your hardware
- Prebuilt heads in anyjev-heads (~100 KB each) for select Qwen3 sizes
- Apache-2.0 with PyPI anyjev and CI workflow in repo
Stack
- Python
- vLLM embed and generate servers
- Hugging Face transformers
- PyPI anyjev
Pricing: Open source library; you pay for GPUs, vLLM hosting, and label collection.
Links
FAQ
- How is AnyJev different from featherless-simple-jev?
- Simple Jev builds classifier HTTP from logits on Featherless or your HF server. AnyJev adds L0 rotation debiasing, optional L2 heads, and pipeline measurement focused on calibration, not just API shape.
- Do I need labels?
- L0 works with zero labels for large flip-rate gains per README. L1 and L2 expect on the order of 100 to 500 labels for temperature and head fits.
- Is this TypeSafe hosted Jev?
- No. AnyJev is Nokia's open research code on models you serve. Compare numbers locally with python -m anyjev.pipeline before trusting banner benchmarks.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev vs LLM classification
When to gate with System One probabilities instead of asking a chat model to label things.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.