What a System One model decides, and what it refuses

System One models return typed probabilities for bounded questions and refuse free-form text, code, and reasoning explanations.

A System One model evaluates a shared state against bounded questions and returns typed probabilities. TypeSafe's Jev uses noul, choice, and score; it does not write replies, produce code, or explain its reasoning. The application owns the thresholds and the next action.

On 2026-10-06, TypeSafe's model page mapped both jev-latest and jev-preview to jev-1.13.0. A 2026-10-05 industry report also described Laya and Cloudflare's Clef as newer System One implementations. The behavior and limits below are specifically for TypeSafe's Jev 1.13.

Key facts about TypeSafe System One models

  • TypeSafe's launch page is dated 2026-09-15, and calls Jev its first System One model. The current model identifier is jev-1.13.0.
  • TypeSafe says Jev uses Reinforcement Learning for Calibrated Decisions (RLCD), which trains calibrated decisions rather than generated prose.
  • A state can be a string, JSON object, or array of text values. Jev allows 64,000 tokens per request, with 32,000 tokens for the state plus the longest question.
  • The API is POST /v1/systemone. Questions using the same state are evaluated independently and in parallel.
  • Jev costs $0.042 per million input tokens, or $42 per billion. TypeSafe says output tokens are free.
  • TypeSafe says the allowed answer structure is defined before inference, so schema type errors are mathematically impossible.

How TypeSafe System One questions work

State is the content being judged. It can be a support message, a passage, or application data. Questions define the judgments, and one request can mix all three question types against the same state.

Question type

Use it for

Returned answer

Limit

noul

A yes/no proposition

noul, a probability from 0 to 1

Binary

choice

One option from a fixed set

choice, probabilities, confidence

Up to 255 options

score

A position on an ordered rubric

score, legend, probabilities, confidence

2 to 10 levels

A noul value near 1 means yes, a value near 0 means no, and a value near 0.5 means the model gives both outcomes similar probability. Noul has no separate confidence field because its one probability already describes the two-outcome distribution.

choice returns the option with the highest probability plus the full distribution across the options. Its confidence is computed from how concentrated that distribution is. score uses the positions of the rubric levels, starting at 0, and returns their probability-weighted mean. A three-level result can therefore be fractional, such as 1.43. Code can rank the result or round it when it needs one level.

The useful boundary is in application code. For example, a program can send one urgency question and one routing question, then keep uncertain cases for a person:

# Illustrative application logic after a System One response
urgent = answers["is_urgent"].noul
team = answers["routing_team"].choice

if urgent >= 0.80:
    escalate_to_oncall(ticket_id)
elif urgent <= 0.20:
    route_to_backlog(ticket_id, team=team)
else:
    route_to_human_triage(ticket_id)

What TypeSafe System One models refuse

TypeSafe's System One documentation says these models do not write replies, produce source code, or generate explanations of their reasoning. Jev's jaggedness guide adds that chaining choice questions to force text generation is slow and does not work well. Use a generative model for a patch, email, summary, or other open-ended output.

Jev accepts text input only. Its model documentation lists strings, JSON objects, and arrays of text values, and says images, audio, and video are not supported directly. Convert those inputs into text or structured fields before sending them as state.

Jev 1.13's nine documented failure modes

TypeSafe's Jev 1.13 jaggedness page names these limits and gives a practical remedy for each one.

Failure mode

What can go wrong

TypeSafe's remedy

Literal reading

Jev answers the written question, not the implied intent

State the exact condition and boundary cases

Math and numbers

Counting, numeric comparisons, and exact interpolation are unreliable

Keep arithmetic and parsing in code

Date and time comparison

Dates are read as text, not ordered quantities

Extract components, then compare in code

Indirection

Double negatives and multi-hop references reduce accuracy

Reduce hops and identify the relevant state

Large irrelevant state

Unrelated detail distracts from the judgment

Filter the state before sending it

Adversarial content

Injected or misleading state can move the answer

Write precise criteria and test edge cases

Contradictory instructions and criteria

Conflicting definitions confuse the model

Align instructions and criteria

Choice option order

The first option can receive a positional bias

Reorder options and check consistency

Generation

Chained choices are slow and poor at text generation

Use a generative model or extract candidates in code

TypeSafe's 193.6x benchmark claim

TypeSafe says the 193.6x faster and 444.6x cheaper figures on its site come from workflow evaluations and represent the high end of real-world gains. The evaluation page averages results across four workflows: security incidents, agent trace observability, invoice processing, and customer service.

The methodology matters. TypeSafe generated reference labels by averaging GPT-6 Astra and Claude Fable 5.1, both at high thinking. Other models used their providers' default reasoning settings and TypeSafe's adapter for structured decisions. TypeSafe reports Jev end-to-end latency of 70 to 500 milliseconds from West Coast developer laptops, compared with 3 to 329 seconds for the frontier models in that comparison. These are TypeSafe's measurements, not an independent benchmark.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.