JevHarness
LLM writes task harness code, freezes it, then Jev makes fast choice/score/noul calls. GEPA evolution and Pokémon 25% to 75% Eval demo.
Overview
JevHarness (github.com/TianyuCodings/JevHarness) lets an authoring LLM write Python harness code: feature extractors, Jev question graphs, and control flow that turns observations into actions. After you freeze the selected harness, execution calls Jev for fuzzy decisions without invoking the authoring model each step. Optional reward reflection and GEPA integration compare parent and child harnesses on training batches and evaluate accepted proposals on Eval. The shipped Pokémon demo archives a full evolution tree: README reports Eval win rate improving from 25% (3/12) to 75% (9/12) after five reflection rounds on the selection set, with latency tables for the archived harness. Companion research to NanoJev on game decisions, but this product is harness authoring, not training a 0.6B unified head. Site: jev-harness.tianyuchen99.chatgpt.site with offline archive viewer.
Problem: Agent loops that ask a large LLM to reason on every action are too slow and too expensive for tight control tasks, yet hand-written Jev criteria rot when the task changes.
Built for: Teams using Claude Code or Codex plugins who want an LLM to author a task harness once, then rely on fast Jev calls at runtime with optional GEPA-style evolution.
First indexed on Jev Directory: 2026-09-25
Creator and team
- Name
- Tianyu Codings
- Handle
- @TianyuCodings
How Jev is used
- Role in the product flow
- Frozen harness code batches choice, score, and noul Jev requests per observation
- Primitives
- ChoiceScoreNoul
- State in
- Task-specific JSON state built by harness features plus parallel questions with instructions and criteria, as in docs/examples/pokemon-turn12-jev.json.
- Decision out
- Typed answers with probability maps and confidence; harness maps the winning action ID to environment commands.
- Install jev-harness Claude Code plugin or skill for your task contract
- Authoring LLM proposes PipelineSpec harness code and Jev graphs
- Optional reflection uses full trajectories and rewards to mutate harnesses
- Freeze selected harness; PipelineRuntime executes Jev nodes on each observation
JevHarness treats hosted or local System One as the fast judge inside code the LLM wrote for your task. tianyucodings-nanojev trains tiny parallel heads on game trajectories; JevHarness keeps the general LLM for authoring and uses stock Jev for runtime fuzzy picks. sutro-sh-jev-align optimizes TypeSafe AI Functions on labeled rows; JevHarness evolves whole harness programs with GEPA on episodic rewards. Do not duplicate NanoJev as a tools page: link the benchmark for model research, use this listing for harness workflows.
Sourced performance claims
- Pokémon Eval example: win rate 25% (3/12) initial harness vs 75% (9/12) selected harness after five reflection rounds on the documented selection set.Source: github.com/TianyuCodings/JevHarness README
- Archived Eval timings: selected harness full decision median 568 ms, individual Jev request median 269 ms, with cache-excluded samples documented.Source: github.com/TianyuCodings/JevHarness README Latency section
- Public GitHub repo had about 255 stars when this listing was drafted.Source: GitHub API September 2026
Features and stack
Features
- Claude Code plugin marketplace install jev-harness@jevharness
- PipelineSpec validation and PipelineRuntime execution path
- Optional GEPA parent selection with recorded ancestry and rejects
- Pokémon archived battles, evolution tree, and latency API on the demo site
- Lossless trace archives for reflection without truncating oversized inputs
Stack
- Python harness runtime
- TypeSafe Jev API
- GEPA integration
- Claude Code plugin
- Node static site viewer
Pricing: Open repository; Jev API usage bills per your TypeSafe or compatible endpoint during harness runs.
Links
FAQ
- Is this the same product as NanoJev?
- No. NanoJev trains a small unified decision model on games. JevHarness authors task code that calls Jev. NanoJev stays on /benchmarks/tianyucodings-nanojev.
- Does the authoring LLM run every turn in production?
- No. README emphasizes freeze the selected harness so runtime is harness code plus Jev calls only.
- Are the Pokémon win rates universal?
- README labels them example results on the Eval selection set, not independent OOD guarantees.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev use cases
The patterns builders actually search for: moderation, routing, triage, RAG verify, and agent gates.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.