tev1
Together's open Jev-inspired recipe: fine-tune Qwen3.5-4B for letter Choice on state plus options, host as Tev1-4B-experimental.
Overview
tev1 (github.com/togethercomputer/tev1, MIT) is Together's open recipe for a Jev-inspired decision model, not a TypeSafe API client. Hassan El Mghari (@nutlope) published together/Tev1-4B-experimental on Together serverless with weights at huggingface.co/togethercomputer/Tev1-4B-experimental. You send state, a question, and 2 to 24 lettered options; the model returns one answer letter via examples/decide.py (temperature 0, max_tokens 8, thinking off, regex parse). Training recipe new v1 starts from Qwen/Qwen3.5-4B with LoRA SFT on 37,840 train and 4,568 val examples. README states this is an independent implementation that does not use Jev answers as training labels.
Problem: Teams want fast letter-pick classifiers without paying TypeSafe per call, but also without guessing JSON from a chat model.
Built for: Builders who need Jev-shaped Choice decisions on their own Together endpoint, or who want to fine-tune Qwen3.5-4B with an open recipe for about twenty dollars.
First indexed on Jev Directory: 2026-09-24
Creator and team
- Name
- Hassan El Mghari
- Organization
- Together
- Handle
- @nutlope
How Jev is used
- Role in the product flow
- Letter Choice over shared state and explicit option lists on a fine-tuned 4B endpoint
- Primitives
- Choice
- State in
- JSON with state string, question string, and options array (label, key, description) per examples/ and decide.py.
- Decision out
- Single option letter (and semantic key in the helper script) parsed from a short completion; logprobs are preferences, not calibrated confidence per README.
- Format state, question, and 2 to 24 lettered options as JSON
- Call Together chat completions with thinking disabled and tight token cap
- Apply the repo system prompt: state is data, return one letter only
- Parse the letter with regex in decide.py and map to option keys
- Optional: run scripts/evaluate.py on a labeled holdout you control
Tev1 is the DIY cousin of hosted System One and the Featherless logits stack: same mental model of state plus explicit options, but the weights and bill live on Together. nutlope-1kpapers still calls TypeSafe for atlas topics; tev1 lets you own the classifier head after a cheap LoRA run. featherless-simple-jev assembles JSON from logits without a chat completion; tev1 learns letter answers from supervised examples. kylejeong-jev-as-judge and classifier.dev remain the paths when you want the closed jev model or HTTP primitives without training. Be honest in architecture reviews: Tev1 does not call api.typesafe.ai and dev bench scores are not production SLAs.
Sourced performance claims
- new v1 recipe uses 37,840 training and 4,568 validation examples from the v1 plus v2.1 union on Qwen/Qwen3.5-4B with LoRA SFT.Source: github.com/togethercomputer/tev1 README and runs/new-v1/README.md
- Development benchmarks on the published endpoint scored 880/1,000 main decisions (88%) and 300/300 policy-transfer; reused during training, not a held-out eval.Source: github.com/togethercomputer/tev1 runs/new-v1/README.md
- Together blog cites about seventeen dollars and about twenty five minutes to fine-tune the sample dataset; serverless together/Tev1-4B-experimental lists about four cents per 1M input tokens with output free.Source: together.ai blog how-to-train-your-own-jev and Together model pricing September 2026
- Public GitHub repo togethercomputer/tev1 had forty two stars and six forks when this listing was drafted.Source: GitHub star count September 2026
Features and stack
Features
- Open MIT repo with dataset builders, train_together.py, and decide.py
- Hosted weights Tev1-4B-experimental on Together serverless
- Hugging Face model card and full weight download
- Documented new v1 recipe with saved dev benchmark reports
- Blog walkthrough from clone to deployed endpoint for about seventeen dollars
Stack
- Python 3.12+
- uv
- Qwen/Qwen3.5-4B
- Together fine-tuning and inference
- LoRA SFT
Pricing: Repo training example targets about seventeen dollars per blog; serverless inference bills per Together model card (input priced, output free on the experimental endpoint).
Links
FAQ
- Is Tev1 the same as TypeSafe Jev?
- No. README calls it a Jev-inspired independent implementation. It does not use Jev training labels and does not call the TypeSafe API unless you wire that yourself.
- Can I trust the 88 percent main benchmark?
- runs/new-v1/README.md labels those rows as reused development benchmarks, not untouched holdout tests. Run evaluate.py on your own split before quoting accuracy in prod.
- How do I try it without training?
- Call together/Tev1-4B-experimental on Together serverless or download weights from Hugging Face, then mirror decide.py settings (temperature 0, max_tokens 8, thinking off).
- How is this different from Simple Jev?
- Simple Jev serves Choice, Score, and Noul from open-model logits on Featherless. Tev1 is a single fine-tuned Qwen that completes one letter per question on Together.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev vs LLM classification
When to gate with System One probabilities instead of asking a chat model to label things.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.