Get the app

Meet Jev: The Non-Generative 'System One' Model Replacing LLMs in Agent Loops

TypeSafe AI ditches autoregressive text generation for parallel, typed decision-making—delivering 200x faster execution and 400x lower cost for programmatic AI workflows.

For the past four years, the entire artificial intelligence industry has operated under a single architectural dogma: every problem should be solved by autoregressively predicting the next token in a string. We have forced massive text-generation transformers to serve as classification switches, fuzzy routers, output validators, and conditional logic gates.

Diogo Almeida—a former OpenAI researcher and co-author of the seminal InstructGPT paper that laid the foundation for ChatGPT—just released the counterargument. After two years in stealth, his startup TypeSafe AI launched Jev, inaugurating what they describe as a fundamentally new architectural category: System One Foundation Models.

Jev makes a radical trade-off that runs directly counter to current frontier trends: it completely abandons free-form string generation. It cannot chat, it cannot draft essays, and it cannot write explanations. Instead, it ingests unstructured or structured system state and outputs type-safe, schema-enforced probabilistic decisions at speeds up to 200 times faster and 400 times cheaper than traditional LLMs.

The Problem with LLMs as Software Glue

Modern agentic workflows and automated pipelines spend the vast majority of their compute cycles on decisions rather than creative generation. In an average multi-agent loop, a model is repeatedly queried to answer simple programmatic questions:

  • Does this customer inquiry require human escalation?
  • Which downstream tool should be invoked next?
  • Did the generated SQL query satisfy safety guardrails?
  • Is this code patch ready for deployment?

When developers use frontier chat models (like GPT-4o, Claude 3.5 Sonnet, or Gemini) to perform these tasks, they inherit the fundamental inefficiencies of autoregressive text generation. Every decision requires sequentially sampling dozens of tokens, introducing 3 to 10 seconds of round-trip latency. Furthermore, developers must construct complex JSON schemas, parse markdown output blocks, handle occasional schema drift, and manage non-deterministic refusals.

As Almeida points out, using a 100-billion-parameter conversational agent to evaluate a boolean conditional is the computational equivalent of driving an eighteen-wheeler to the corner store to buy milk.

   Traditional Agent Loop (Sequential Autoregressive LLM):
   [State] ──> [Token 1 ──> Token 2 ──> Token N (JSON String)] ──> [Parser / Regex] ──> [Branch]
   Latency: 2,000ms – 10,000ms | Hallucination Risk: Non-Zero | Cost: High

   System One Loop (Jev Parallel Decision Sampler):
   [State] ──> [Direct Parallel Tensor Evaluation] ──> [Typed Output: {is_urgent: 0.999}]
   Latency: 70ms – 200ms | Hallucination Risk: 0% | Cost: Near Zero ($0.042 / MTok)

Under the Hood: RLCD and Non-Autoregressive Sampling

The name "System One" is drawn from Daniel Kahneman’s dual-process cognitive framework. While frontier reasoning models (such as OpenAI's o-series or DeepSeek-R1) leverage RLVR (Reinforcement Learning with Verifiable Rewards) to execute slow, deliberative "System 2" chain-of-thought generation, Jev is purpose-built for instantaneous "System 1" intuition.

To build Jev, TypeSafe AI replaced the standard LLM training and inference stack with three distinct innovations:

  • Reinforcement Learning for Calibrated Decisions (RLCD): Modern chat LLMs are optimized via RLHF (Reinforcement Learning from Human Feedback), which trains models to produce verbose, pleasing text. RLCD instead optimizes the model’s internal logits to reflect calibrated, epistemically honest probabilities. If Jev assigns an 85% probability to a classification, historical accuracy on that distribution strictly tracks 85%.
  • Hardware-Aware Parallel Sampler: Because Jev does not condition token $N$ on token $N-1$, it bypasses sequential KV-cache decoding entirely. When an application passes multiple questions against a single state context, Jev evaluates all queries in parallel in a single forward pass. Adding ten extra questions to a query barely shifts the end-to-end response time.
  • Typed Output Primitives: The model exposes three native output primitives declared prior to execution:
    • Noul: A strict boolean assertion returning an exact calibrated probability ($P \in [0, 1]$).
    • Choice: Categorical selection across up to 255 pre-declared options, returning confidence scores and full probability distributions.
    • Score: Continuous ordinal rating across custom ranges with uncertainty intervals.

Because the output surface is bounded entirely by the predefined schema, TypeSafe claims syntax errors and hallucinations are mathematically impossible.

The Raw Numbers: 193.6x Faster, 444.6x Cheaper

According to TypeSafe AI's benchmarks, the operational delta between running a System One decision model and querying frontier LLMs is stark:

  • Latency: End-to-end request response times drop from 3,000ms–15,000ms down to 70ms–500ms, representing a 40x to 193.6x speedup on equivalent classification and extraction tasks.
  • Unit Economics: Input tokens are priced at $0.042 per million tokens ($42 per billion tokens), while output tokens are free ("too cheap to meter"), representing a 444.6x cost reduction compared to mid-tier frontier models.
  • Parallel Question Scaling: Evaluating 20 distinct properties over a large document batch executes in sub-second latency, enabling map-reduce operations over terabytes of unstructured logs and data feeds.

The Ecosystem Integration: LangChain and Beyond

Framework maintainers have moved fast to adopt the architecture. LangChain rolled out official support via langchain-typesafe, introducing the TypeSafeClassifier module to decouple agent routing from generative execution.

from langchain_typesafe import Noul, TypeSafeClassifier

# Instantiate the System One decision engine
classifier = TypeSafeClassifier()

response = classifier.invoke({
    "state": "The primary database connection pool exhausted on worker node 4. Alerts triggering.",
    "questions": {
        "is_p0_incident": Noul(instructions="Does this indicate immediate production customer downtime?"),
        "requires_security_team": Noul(instructions="Is this related to an unauthorized intrusion attempt?")
    }
})

print(response.nouls["is_p0_incident"].noul) # Output: 0.984

In experimental setups, developers have integrated Jev as a high-speed router. Jev inspects user input in 80ms, determines whether the request is simple or complex, and routes to lightweight or heavyweight generative models accordingly—cutting overall pipeline latency in half.

Other early implementations include sub-100ms real-time game controllers (including a viral demo running automated navigation in Doom), live browser-use action dispatchers, and automated guardrail arbiters evaluating prompt jailbreaks before they ever reach larger language models.

The Bifurcation of AI Architectures

For years, the industry operated under the assumption that a single, unified general-purpose model would eventually handle every stage of the software stack. Jev demonstrates that the future of production AI is likely a dual-tier architecture:

  1. Generative & Reasoning Engines (System 2): Massive autoregressive models deployed sparingly for high-context human interaction, creative drafting, and long-horizon verifiable problem-solving.
  2. Deterministic Decision Engines (System 1): Ultra-fast, type-safe models like Jev operating as the compiled "smart if-statements" and nervous systems of modern software.

By treating classification and decision-making as first-class architectural primitives rather than a side effect of chat generation, TypeSafe AI has pointed out an obvious truth: machines don't need models to talk to them—they just need them to decide.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play