Use Clef-flash for the fastest text-and-image routing, Clef when you need the larger model or a longer context window, and Jev when input cost and text-only routing matter more than latency. In Cloudflare's 2026-10-01 edge benchmark, Clef-flash had 38.8 ms median latency, Clef 209.3 ms, and Jev 524.1 ms. Those figures are not a universal accuracy ranking: Jev still led Clef on the When2Call task.
Clef, Clef-flash, and Jev key facts
- On 2026-10-01, Cloudflare released
@cf/cloudflare/clef(27B) and@cf/cloudflare/clef-flash(9B) under the Apache 2.0 license on Workers AI. - In Cloudflare's 43 benchmark runs, Clef-flash achieved 38.8 ms median latency (122.4 ms p95), Clef achieved 209.3 ms median (238.6 ms p95), and Jev achieved 524.1 ms median (536.0 ms p95).
- TypeSafe's benchmark page reports 70 to 500 ms per call for Jev 1.13 on its hosted API.
- The published input rates are $0.042 per million tokens for
typesafe/jev, $0.090 for@cf/cloudflare/clef-flash, and $0.240 for@cf/cloudflare/clef. - Clef and Clef-flash accept up to four images, whereas TypeSafe Jev 1.13 accepts text and structured text only.
- TypeSafe documents Jev 1.13 with a 64k total request budget and a 32,000-token limit for
stateplus the longest question. Cloudflare liststypesafe/jevat 32,000 tokens. - Cloudflare lists Clef and Clef-flash at 65,536 tokens and marks both models as vision-capable.
Clef and Jev's routing interface
System One decision models evaluate one input state against typed questions and return probabilities rather than prose.
Both Clef and Jev follow the System One API schema, evaluating up to 64 typed questions per call across three primitives:
noul: A boolean question returning the probability that the answer is true.choice: A categorical question returning the selected label, per-option probabilities, and a confidence score.score: An ordinal rubric returning a probability-weighted score across defined levels.
Because Clef matches Jev's API, an agent router switches models on Cloudflare Workers by updating the model string:
const response = await env.AI.run("@cf/cloudflare/clef-flash", {
model: "clef-flash",
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: {
type: "noul",
instructions: "Is this support request urgent?",
},
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
},
},
},
});Clef and Jev latency benchmarks
Latency on the hot path determines whether an orchestrator can evaluate guardrails and branch tools inline before dispatching downstream workers.
Model | Parameter size | Median latency | p95 latency | Measurement source |
|---|---|---|---|---|
| 9B | 38.8 ms | 122.4 ms | Cloudflare (43 benchmark runs, 2026-10-01) |
| 27B | 209.3 ms | 238.6 ms | Cloudflare (43 benchmark runs, 2026-10-01) |
| Closed | 524.1 ms | 536.0 ms | Cloudflare (43 benchmark runs, 2026-10-01) |
Jev 1.13.0 (hosted API) | Closed | 70 to 500 ms (band) | Not reported | TypeSafe (vendor benchmark) |
These figures do not measure the same route. TypeSafe's 70 to 500 ms figure is a vendor-reported hosted API range. Cloudflare's 43-run comparison used Workers AI. In that comparison, Jev reached 524.1 ms median, while Clef and Clef-flash were 2.5x and 13x faster at median.
Jev, Clef, and Clef-flash pricing
The published input rates are:
- TypeSafe Jev (
typesafe/jev): $0.042 per million input tokens ($42 per billion tokens). - Cloudflare Clef-flash (
@cf/cloudflare/clef-flash): $0.090 per million input tokens. - Cloudflare Clef (
@cf/cloudflare/clef): $0.240 per million input tokens.
At those published input rates, Jev costs less than half of Clef-flash and less than a fifth of Clef. That makes Jev the lowest-cost option among these three for text routing where 500 ms response times fit the SLA.
Clef, Clef-flash, and Jev input limits
Two structural differences dictate pipeline fit: multimodal input and context allocation.
TypeSafe Jev 1.13 is text-only. Agent workflows handling UI captures or documents must run OCR or visual captioning before calling Jev. In contrast, Clef and Clef-flash include a vision encoder and accept up to four images. In Cloudflare's Browser Run example, Clef rendered and classified a web page screenshot in 2.2 seconds, compared to 4.7 seconds for gpt-oss-120b.
Catalog listings also show conflicting context numbers:
- Cloudflare's Workers AI catalog lists
typesafe/jevwith a 32,000 token context window. - TypeSafe's documentation specifies Jev 1.13 with a 64k token limit per request, with the 32,000-token state boundary described above.
The two figures describe different limits. TypeSafe permits 64k tokens across the entire request, but caps the input state plus the longest single question at 32,000 tokens. Clef and Clef-flash have a 65,536-token context window in Cloudflare's catalog.
Clef and Jev decision quality
Cloudflare's Jev Decision Index shows a mixed result. Clef or Clef-flash scored highest on 7 of 10 evaluated benchmarks, while Jev led on When2Call and agent trace observability:
- BFCL (case exact): Clef-flash scored 98.76, Clef scored 98.47, and Jev scored 95.75.
- BANKING77 (macro-F1): Clef scored 94.20, Clef-flash scored 90.93, and Jev scored 79.74.
- Home appliances (case exact): Clef-flash scored 97.73, Clef scored 82.95, and Jev scored 52.27.
On Cloudflare's copy of TypeSafe's workflow evaluations, the Clef variants beat Jev in three of four domains: invoice processing (64.7 for Clef versus 61.8 for Jev), customer service (76.3 for Clef and 77 for Clef-flash versus 76.0 for Jev), and security incidents (62.9 versus 61.7). Jev led agent trace observability, scoring 71.6 against Clef-flash's 69.8 and Clef's 68.5.
A Beri comparison read on 2026-10-06 describes the same split, with Clef ahead on intent and tool-call tests and Jev ahead on the "should I act" When2Call test. It is corroboration, not a new benchmark.
Choosing Clef, Clef-flash, or Jev for agent routing
- Use Clef-flash ($0.09/Mtok) for latency-sensitive paths that can tolerate the benchmark's 122.4 ms p95, or when evaluating screenshots alongside text.
- Use Clef ($0.24/Mtok) when the larger model or 65,536-token context window fits the workload. Test the decision types first because Jev leads some judgment-oriented benchmarks.
- Use Jev ($0.042/Mtok) for budget-conscious, text-only routing pipelines where its 70 to 500 ms hosted latency band fits the SLA.
Sources
- Cloudflare Workers AI Changelog: https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/ (read 2026-10-06)
- Cloudflare Blog: Introducing Clef: https://blog.cloudflare.com/clef-decision-models/ (read 2026-10-06)
- Cloudflare Workers AI Model Catalog - Clef: https://developers.cloudflare.com/workers-ai/models/clef/ (read 2026-10-06)
- Cloudflare Workers AI Model Catalog - Clef-flash: https://developers.cloudflare.com/workers-ai/models/clef-flash/ (read 2026-10-06)
- Cloudflare Workers AI Model Catalog - Jev: https://developers.cloudflare.com/ai/models/typesafe/jev/index.md (read 2026-10-06)
- TypeSafe AI Models Documentation: https://docs.typesafe.ai/models (read 2026-10-06)
- TypeSafe AI Jev Benchmark: https://www.jevtypesafeai.com/jev/benchmark (read 2026-10-06)
Last verified: 2026-10-06.