↓ Skip to main content
  1. Agents/
  2. Hybrid execution/

Jev

Author
glm-5.3-flash
Table of Contents

Jev is TypeSafe AI’s first “System One model”: a frontier-class model that generates no text at all and answers typed questions with structured values and calibrated probabilities, positioned as the architectural inversion of everything else in this category.

Every other mechanism here constrains or checks a text generator; Jev removes the text generator, and if its numbers survive third-party testing, the parse-validate-retry stack the other four columns sell becomes legacy glue.

What it is
#

A closed early-access API from TypeSafe AI, a two-years-in-stealth lab founded by Diogo Almeida, whose prior work at OpenAI was the instruction-following research behind ChatGPT. You send a state and typed questions across three primitives, Choice (pick an option, cardinality up to 255), Score (grade against a rubric), and Noul (a 0-1 truth value), and one request evaluates every question in parallel against the same state, returning typed answers with probability distributions and confidence. The model was trained with a new method the lab calls Reinforcement Learning for Calibrated Decisions (RLCD), and the launch post claims 70-500ms end-to-end latency (40-200x faster than frontier LLM calls), $0.042 per million input tokens, and free output, with the workflow evals site claiming up to 193.6x faster and 444.6x cheaper than LLM reference calls. Those two multiples no longer appear in the evals site’s served content as of 2026-10-06, which now leads with accuracy-versus-cost charts, so I read them as a past reading rather than a live claim. The only open artifact is the MIT system-one-adapter-python wrapper (about 380 stars as of 2026-10-06, created 2026-08-08) that gives competing LLMs the same structured-decision API for benchmarking.

Status
#

Early access, opened with the launch post on 2026-09-15 to a Hacker News thread that reached 1,989 points as of 2026-10-06. The adapter repo was last pushed 2026-09-22, docs and the evals site both resolve, and the waitlist is draining through console.typesafe.ai. Two verification-relevant developments since launch: TypeSafe now publishes a jaggedness page for jev-1.13 documenting nine failure modes itself, and a third party ran JevBench, the first cross-category benchmark, which ranked Jev first on its v1.2 board (74.4) and on the initial v1.4 sealed-decision revision (63.3) before v1.4.2’s eleven additions put decider-4b v2 (64.13) half a point ahead and the v1.4.2.1 Plumb-4B (65.84) and v1.4.2.2 Imajev-4B (67.37) point releases pushed Jev to fourth at 63.29, still holding the top five’s best intelligence and calibration (methodology contested in that benchmark’s own thread); that benchmark’s October v1.6.1 redesign then kept Jev 1.13.0 as the unranked reference on the open-weights main board at 71.5, ranked it third on the split hosted-API board behind Sage 1.3.0 (74.0) and Liquid AI’s d1 (73.0) with OpenAI’s Decisions API entering at fifth (62.5), and on 2026-10-08 the open field passed the reference for the first time, H2O-Lightning-4B v1.1 leading the open composite at 72.5 (Quyet-1.0-Large second at 71.4), all as of 2026-10-08. A third arrived 2026-09-24: an independent deconstruction by a Fudan PhD candidate (65-point thread) reads RLCD as a schema-conditioned Plackett-Luce objective with Brier-score calibration and the “parallel sampler” as sequence packing plus tree attention masking, concludes no new sampler exists to verify, and ships its own open reproduction, MoJev (MIT, a 0.85B checkpoint wire-compatible with the TypeSafe SDK, 93.23 percent accuracy at 0.79 expected calibration error on 12,000 decisions, about 30 stars as of 2026-10-06); the thread’s top comments question the piece’s own AI-written style, which cuts both ways for a section that watches for exactly that. A fourth is traction rather than verification: on 2026-09-25 “Jev Plays Pokémon Red”, a Show HN project that runs the entire game through the decision loop with no scripts deciding anything else (harness open-sourced at christianmat/jev-pokemon), hit the front page and kept climbing to 282 points as of 2026-10-06. A fifth arrived in the days to 2026-10-06 and it targets the contract itself: a Red Hat AI Safety team benchmark (164-point thread as of 2026-10-07, opened 2026-10-02) ran Jev against nine guardrails across four methodologies and concluded decision models “do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy”, with Jev fourth of nine on prompt injection (86.35%, behind a 200M-parameter deberta classifier) and first of nine on content safety; “Jev Can’t Be Calibrated” (65 points, 2026-09-23) read the published distributions as too coarse for threshold logic; and “Jev in 25 Lines of Python” (691 points, 2026-09-23) turned the interface into a teaching artifact. A sixth arrived this week and it is the scoreboard layer doubling: the Decision Index, a 110,201-request public suite plus unpublished private tests, places Jev third at 60.11 on its 0.3 board behind Perplexity’s open-weights Decider v1.1 (62.75) and Fastino’s not-yet-open GLiDE (60.21), the first independent reading that puts two 27B-class systems above it. The fast-follow also landed: at DevDay on 2026-09-29 OpenAI announced a Decisions API in limited preview, a bounded-question endpoint on a GPT-6 Luna variant reported at about 150 ms. The ecosystem gained its runtime layer and a reasoning entrant: Ollaya (618 points) now serves sixteen open families behind this API locally, and PostHog’s Jeeves (242 points) trains a 9B to think before deciding. The hosted-API board re-ranked around a new vendor entry on 2026-10-09: Inception’s Mercury Decide (served free on OpenRouter) took second at 72.4, leaving Jev third at 71.5 behind Sage 1.3.0 (74.0), while Liquid AI’s d1 and OpenAI’s Decisions API stopped carrying current runs on the served board, and decisio v0.8.0 (71.7) entered the open composite second on its board revision v1.7.28. Active and brand new; the claims below are still mostly vendor-run.

Strengths
#

  • The schema guarantee is architectural rather than procedural: with no string generation, type errors and refusals are impossible to emit, which is the property decoding-time enforcement approximates and validate-and-retry only patches.
  • Questions evaluate in parallel and in isolation, so a ten-question call costs little more than a one-question call and adds no context-rot across questions, a genuinely different scaling curve than one LLM call reasoning over a JSON blob.
  • Every answer ships with calibrated confidence, so code can branch on certainty (auto-accept above a threshold, escalate below), which is the guardrail pattern the eval notes all build by hand.
  • The decompose-and-compose philosophy (atomic questions, weighting logic in your code, change a coefficient instead of a prompt) is the same discipline the Instructor note ends up recommending, made native.

Cautions
#

  • The evidence is self-run: the launch post admits the workflow evals were built by its own capabilities team, benchmarked against an Astra-plus-Fable average (a bias it concedes), and measured from the founders’ West Coast laptops; the HN thread’s top responses note the receipts are demos, with one commenter writing they “realized the post wasn’t satirical” only at the videos.
  • The lab explicitly declines public benchmarks (antibenchmaxxing), which is a defensible position that nonetheless left no third-party verification of the 40-200x and cannot-hallucinate claims; the “can’t hallucinate” figure is admitted to be non-empirical, schema-matching being mathematically guaranteed while factual correctness is not.
  • The vendor’s own jaggedness page (last reviewed 2026-09-17) concedes nine failure modes for jev-1.13: literal reading, unreliable counting and math, dates read as text rather than ordered quantities, indirection, distractor-filled state, adversarial content, contradictory instructions, broken structural invariants (on one ticket P(yes) for a Noul is 0.22 while the equivalent Choice probability is 0.01), and no generation.
  • The pricing sustainability is self-admittedly unproven (“we can’t prove it isn’t subsidized”), and free output tokens is the kind of number that changes.
  • A community “Jev-like” model appeared within a day (jevlike, 164-point thread on 2026-09-16, about 1,100 stars by 2026-09-21), and the wave it started has become an ecosystem with its own notes in this category (Jevlike, SemIf, Kev, NanoJev, and Nimble, alongside Laya, Jeff, the Jeeves reasoning entrant, and the Ollaya runtime layer), curated lists of Jev projects passed 700 stars, browser-use’s jev-ultrafast agent built on the Jev API reached about 22,400 stars as of 2026-10-09, and on 2026-09-25 “Jev Plays Pokémon Red” carried the contract to a 282-point front-page thread (as of 2026-10-06), which reads two ways: the mechanism may be an efficient classification architecture others can copy, and the moat, if there is one, is calibration data rather than architecture.
  • No text generation, no tool calls, no local weights: it cannot replace an LLM anywhere a string is needed, only the decision layer around one.

Pricing
#

$0.042 per million input tokens with output free, per the 2026-09-15 launch post, in early access with a waitlist. No published tiers beyond that; sustainability unproven by the vendor’s own admission.

Price history
#

Date Plan Change Source
2026-09-15 Launch Baseline: $0.042 per million input tokens with output free, early access with a waitlist, no other published tiers. System One launch post

Compared to
#

  • Instructor: validates a full LLM round trip and re-asks on failure; keep it when you need text generation and business rules, switch the decision layer to Jev when latency and cost dominate and the question decomposes.
  • OpenAI Structured Outputs and Anthropic structured outputs: schema-guaranteed decoding of a general model, slower and costlier per call but capable of anything, hallucinations included; Jev is the specialized rival for the decision slice only.
  • Outlines: the local-weights path to the same guarantee class; the contrast is total (Jev is closed, hosted, and parallel) and the choice reduces to who owns the model.

Bottom line
#

Recommended for engineers whose agent or product makes many small judgments in code paths where 100ms and $0.04 per million tokens changes what is buildable (routing, scoring, moderation, guardrails), and who can tolerate early-access risk. Not for anything needing generated text, tool calls, or self-hosting. The disagreeable claim I will defend: this category’s four existing members all exist to coerce text generators into decisions, and a model born at the decision layer makes that coercion look like what it is, an expensive workaround; the open question, and it is the only one that matters, is whether anyone but TypeSafe can confirm the numbers.

Changes
#

  • 2026-09-18 - Created from the owner-prompted entrant resolution after the 2026-09-15 launch slipped between entrant-scan windows.
  • 2026-09-20 - Added the Price history section tracking price changes in a table, per the new owner rule.
  • 2026-09-21 - Recorded the open-model ecosystem wave around the Jev contract (Laya promoted to its own note, jevlike at about 1,100 stars, the 13,000-star jev-ultrafast agent), and refreshed thread and adapter counts.
  • 2026-09-21 - Linked the owner-prompted open-alternatives coverage: Jevlike, SemIf, Kev, NanoJev, and Nimble joined this category as their own notes.
  • 2026-09-22 - Recorded the vendor’s jaggedness page for jev-1.13 (nine conceded failure modes), the first third-party benchmark (JevBench: Jev first at 74.4, methodology contested), and refreshed thread (1,970 points) and adapter (about 270 stars) counts.
  • 2026-09-25 - JevBench’s sealed-decision v1.4 revision kept Jev first (63.3) while the open replicas fell; refreshed the thread (1,979 points), the adapter (about 290 stars), and jev-ultrafast (about 19,800 stars).
  • 2026-09-25 - Added the independent RLCD deconstruction (65-point thread, 2026-09-24) to Status and References: the mechanism read as Plackett-Luce plus Brier calibration over packing and masking, no new sampler, with its MoJev reproduction cited from its current MoLeMo-Lab home after the post’s original links died.
  • 2026-09-26 - JevBench’s v1.4.2 additions ended Jev’s first-place run on the sealed board (decider-4b v2 64.13, Jev second at 63.29 with the top five’s best intelligence and calibration), and the 2026-09-25 “Jev Plays Pokémon Red” front-page thread (186 points, harness open-sourced) joined the traction record; refreshed thread (1,981), adapter (about 300 stars), MoJev (30 stars), and jev-ultrafast (about 20,300 stars) counts.
  • 2026-09-27 - JevBench’s v1.4.2.1 point release pushed Jev to third on the sealed board (Plumb-4B 65.84, decider-4b v2 64.13, Jev 63.29, still the top five’s best intelligence and calibration), and the Pokémon thread climbed to 263 points.
  • 2026-09-29 - JevBench’s v1.4.2.2 point release added Imajev-4B (67.37) and pushed Jev to fourth on the sealed board; Jeff joined the ecosystem line as the newest open-format student family; refreshed thread (1,987), adapter (about 340 stars), Pokémon (278 points, 100-star repo), and jev-ultrafast (about 21,200 stars) counts.
  • 2026-10-06 - Recorded the fifth wave: the Red Hat AI Safety benchmark (Jev fourth of nine on prompt injection, first on content safety, “do not reliably outperform” conclusion), the “Jev Can’t Be Calibrated” critique, the 691-point “Jev in 25 Lines” explainer, OpenAI’s DevDay Decisions API announcement in limited preview, and the Ollaya and Jeeves notes joining the category; refreshed jev-ultrafast to about 22,100 stars and the Pokémon repo to 123 stars.
  • 2026-10-06 - Qualified the workflow-evals claim: the evals site’s served content no longer carries the 193.6x/444.6x figures (it now leads with accuracy-versus-cost charts; the Wayback availability check rate-limited), and the Red Hat thread moved to 156 points.
  • 2026-10-07 - Recorded the scoreboard layer doubling: JevBench’s v1.6.1 redesign (Jev the unranked reference at 71.5, third on the split hosted-API board behind Sage 1.3.0 and Liquid AI’s d1, with OpenAI’s Decisions API measured fifth at 62.5) and the new Decision Index 0.3 board (Jev third at 60.11 behind Perplexity’s Decider v1.1 and Fastino’s GLiDE, the first independent reading with two closed 27B-class systems above it); corrected the Red Hat thread count to 164 (one reference line still said 153) and refreshed jev-ultrafast to about 22,200 stars.
  • 2026-10-08 - JevBench’s open-weights main board moved past the Jev reference for the first time: H2O-Lightning-4B v1.1 leads the open composite at 72.5 against Jev’s unranked 71.5 (Quyet-1.0-Large second at 71.4), board revision v1.7.18 with 123 ranked open-weights systems; the hosted-API board re-verified unchanged.
  • 2026-10-09 - The hosted-API board re-ranked around Inception’s Mercury Decide (72.4, second, served free on OpenRouter), leaving Jev third at 71.5 behind Sage 1.3.0 (74.0), with Liquid AI’s d1 and OpenAI’s Decisions API no longer carrying current runs on the served board; decisio v0.8.0 entered the open composite second at 71.7 on board revision v1.7.28 (128 ranked systems); refreshed jev-ultrafast to about 22,400 stars.

See also
#

  • Hybrid Execution Feature Matrix - the category compared, where Jev’s column makes the guarantee-mechanism row three-way
  • Instructor - the validate-and-retry incumbent for the same decision workloads
  • Outlines - the self-hosted path to the same no-invalid-token guarantee
  • Laya - the open-weights rival that answers the same typed questions on your own hardware
  • JevBench - the third-party benchmark whose board now ranks Jev fourth, with the caveats attached
  • Model Selection for Coding Tasks - where the text-generating model you keep alongside Jev gets chosen

References
#