Valyu
A carved stone hand operates a decision press: several candidate tiles converge through a narrow gate, and a single checked answer tile emerges, illustrating a fast bounded decision.
Tutorial

Jev vs OpenAI's Decisions API: What to expect

>_ Prosper

Jev and OpenAI's Decisions API tackle the same problem: you define bounded answers, supply context, and get back a decision software can use. Jev is TypeSafe AI's purpose-built System One model, released in early access on September 15, 2026. TypeSafe lists $0.042 per million input tokens, free output, and 70–500ms end-to-end latency. OpenAI announced the Luna-based Decisions API at DevDay on September 29. It accepts text or images and is in limited preview. Launch coverage reports roughly 150ms per decision, but OpenAI has not published a matched benchmark against Jev or a price for the Decisions API.

Short version: Jev has public documentation and pricing; the Decisions API has an official preview announcement, but little implementation detail. You can evaluate Jev's documented interface today. You cannot yet make a fair cost, accuracy, or latency comparison between the two.

Checked October 1, 2026. Jev figures below are TypeSafe's published figures unless noted. Decisions API details are from OpenAI's DevDay announcement and attributed launch coverage; preview details may change.

TL;DR

Jev (TypeSafe AI)Decisions API (OpenAI)
AnnouncedSeptember 15, 2026September 29, 2026 (DevDay)
StatusEarly access, live on Vercel AI Gateway and DigitalOceanLimited preview for selected API customers
Underlying modelJev, a new model class trained with RLCD"A version of GPT-6 Luna"
Reported latencyTypeSafe: 70–500ms end to endLaunch coverage: ~150ms vs ~1.6s for Luna via the regular API; no matched test
Input price$0.042 / MTokNot published
Output priceFreeNot published
Input typesText only (string, JSON, array of strings)Text and images
Question typesChoice (up to 255 options), Score (2 to 10 levels), Noul (yes/no probability)User-defined questions with finite predefined answers; full schema unpublished
Multiple questionsDocumented in one parallel callOpenAI describes a "set of user-defined questions"; request/response shape unpublished
Probabilities and confidenceTypeSafe says probabilities are calibrated; Choice/Score have a distribution and derived confidence; Noul returns a yes probability onlySome coverage reports confidence scores; OpenAI has not documented their meaning or calibration
Context limits64k tokens total, 32k for state plus longest questionNot published
Standard rate limits100k tokens/s, 40 requests/s as currently listed; TypeSafe says limits can changeNot published
Public docsYes, including a failure-mode pageNo Decisions API entry in the public API docs index or pricing page as of Oct 1

Think of one customer ticket saying a payment integration failed and they need help today. Both products are designed to return a routing decision instead of prose. Jev's published contract lets you ask about urgency in the same call. OpenAI mentions multiple questions but has not published the request schema or confidence semantics.

One ticket, two bounded-decision interfaces. The difference here is how much of each contract is public, not proof that OpenAI is limited to one question.

What is Jev?

Jev is a decision model from TypeSafe AI that returns typed, probabilistic answers instead of generated text. You send state (a string, a JSON object, or an array of text) and a set of typed questions. Jev answers the questions in parallel. TypeSafe says the output shape is guaranteed by construction: a returned value fits the defined answer space, but it can still be the wrong valid answer.

TypeSafe came out of stealth on September 15, 2026 with a reported $40M seed round led by DCVC. Its CEO Diogo Almeida is a former OpenAI researcher; TypeSafe credits him as a co-inventor of RLHF and InstructGPT, methods behind ChatGPT and GPT-4. TypeSafe says Jev is trained with RLCD (Reinforcement Learning for Calibrated Decisions), an objective aimed at calibrated probabilities rather than preferred prose.

The whole API is three primitives:

  • Choice: pick one of up to 255 options. Returns the choice, probabilities across options, and a confidence statistic.
  • Score: place something on a scale of 2 to 10 described levels. Returns a position that can land between levels, probabilities, and confidence.
  • Noul: a yes-or-no probability. It does not return a separate confidence field.

I wrote a full walkthrough of these, with patterns and failure modes, in How to Use Jev.

What is the OpenAI Decisions API?

The Decisions API is OpenAI's preview for applying Luna to bounded decisions. OpenAI describes a set of user-defined questions with finite predefined answers, with text or images as context. It names three uses: classifying content, routing requests, and choosing an agent's next action. OpenAI has not yet published the full request or response schema.

The speed figure reported in launch coverage is about 150 milliseconds per decision versus 1.6 seconds for Luna through the regular API. That is a comparison inside OpenAI's stack, not a benchmark against Jev. In the DevDay demo, 10,000 customer requests go through the Responses API on the left and the Decisions API on the right.

Same CSV, same prompt: "Route these 10,000 customer requests." 1.6s per request vs 150ms per request.

The Decisions API is done while the Responses API sits under 1,000. Note the "shown 15x realtime" label on the demo.

"~10x faster decisions." Faster than OpenAI's own Responses API. Read that sentence twice.

Here is what OpenAI has not published for the Decisions API: a public price, rate limits, answer-set limits, the precise multi-question request shape, the meaning or calibration of any returned confidence value, context limits, regional availability, or independent accuracy and latency comparisons. As of October 1, it has no dedicated entry in the public developer docs index, API changelog, or pricing page. The regular GPT-6 Luna model is documented separately; its limits and prices cannot be assumed to apply to this preview API.

Jev vs Decisions API: the comparison that matters

1. Speed: the 10× is not about Jev

Look at what the 10x compares. It is the Decisions API against GPT-6 Luna through OpenAI's own Responses API. It says nothing about Jev.

The comparison has a denominator: OpenAI's own regular API. Neither vendor has published a matched head-to-head test.

TypeSafe reports 70–500ms end-to-end response time for Jev. Launch coverage quotes roughly 150ms for Decisions API. Those figures overlap, but they were not measured under the same workload, region, input size, or concurrency. We cannot rank the APIs on latency from these figures.

The demo illustrates why a bounded-decision interface can be attractive for routing instead of a regular text-generating call. If you ran 10,000 calls strictly one after another at the reported per-call times, 1.6 seconds each would add up to about 4.4 hours and 150ms each to 25 minutes. This is illustrative arithmetic, not a throughput benchmark: concurrency and rate limits change the wall-clock result.

2. Price: Jev has one. OpenAI does not.

Jev costs $0.042 per million input tokens, and output is free. Adding questions or Choice options still increases the input-token count; a meaningful cost estimate depends on the size of the state and how many requests you make.

OpenAI has not listed a price for the Decisions API. The frequently cited $0.10 input and $0.50 output per million tokens are GPT-6 Luna's standard short-context text prices, not a published Decisions API tariff. Jev's listed input price is known; an apples-to-apples cost comparison must wait for OpenAI's pricing and measured request sizes.

3. What comes back: probabilities, confidence, and missing details

This is the part most launch-day coverage skipped, and it is the part that changes how you build.

Jev's Choice and Score answers include probabilities across the options or levels and a separate confidence value derived from how concentrated that distribution is. Noul returns a single probability of "yes," with no separate confidence field. TypeSafe says its model's probabilities are calibrated, but we have not found an independent calibration study. The practical pattern is still useful: set action thresholds in your own code, test them on your own data, and scale them to the cost of being wrong. A read-only lookup and a money-moving action should not share an automatic-approval rule; high-stakes actions need explicit confirmation.

The classification and the action are separate decisions. The cost of being wrong determines where your application draws the line.

Some launch coverage says the Decisions API returns confidence scores. OpenAI's own announcement does not document a confidence field or claim calibration, so do not design a production threshold around it yet. Equally, Jev's confidence of 0.85 is not a claim that this individual answer has an 85% chance of being right: it is a statistic of the returned distribution. Calibration of probabilities is a claim about outcomes across groups of predictions, not a guarantee on one decision. Validate any threshold against labeled examples from your workflow.

4. Shape of the call: documented fan-out vs an unpublished schema

Jev documents many independent questions over one state in a single parallel request. More questions cost input tokens, but TypeSafe says they add little latency. Its cookbook reports that batching 13 questions over one long document was 12.2× cheaper and 10× faster than 13 sequential calls on Jev 1.12. Running the 13 calls concurrently would narrow the wall-clock speed gap, not eliminate the duplicate input-token cost.

OpenAI's announcement explicitly says a "set of user-defined questions". That rules out treating "one question per call" as an established limitation. What remains unknown is how the set is represented, whether questions run together in one request, and what the response contains. Do not assign Jev a five-to-one call-count advantage without testing the released API.

5. Inputs: images are the Decisions API's clear win

The Decisions API's stated image input is a clear workflow advantage for image-first tasks. Jev is text only. If your decision depends on a screenshot, a scanned form, or a product photo, Jev requires OCR or captioning first; the announced Decisions API accepts the image directly. We still need real accuracy and cost tests on those tasks.

For image-heavy tasks, compare the entire workflow, including any OCR or captioning step.

Two footnotes. The typesafe-computer-use project shows an OCR/accessibility-tree preprocessing path before Jev, with a self-reported cost of about $0.0002 per decision for its tested loop; that figure does not include every possible preprocessing cost or workload. Image support in Decisions is still part of a preview, so test it on your own images before choosing a provider.

6. Honesty about failure

TypeSafe publishes a page called "jaggedness" for Jev 1.13: it reads instructions literally, does not count reliably, treats dates as text rather than ordered quantities, loses accuracy when state is padded with irrelevant material, and can be steered by adversarial content in state. Keep arithmetic and date comparison in code, and test hostile inputs.

OpenAI has published nothing comparable for the Decisions API yet. Which is fair for a preview. It is also why you cannot evaluate the two on equal terms today.

And a caveat on Jev, because the same standard should apply to both: its four-workflow evaluation reports 67.8% for Jev against consensus labels generated by GPT-6 Astra and Claude Fable 5.1, not independent ground truth. TypeSafe's "193.6× faster, 444.6× cheaper" headline comes from its own workflow comparisons, not a Jev-versus-Decisions-API test. These results are self-run; evaluate both products on your own traffic when you can access them.

Why OpenAI launching this is good news for Jev

The launches were 14 days apart: TypeSafe announced Jev on September 15, and OpenAI announced the Decisions API on September 29. That timing makes the comparison interesting. It does not establish when OpenAI started building the API or why it chose its launch date.

My read: OpenAI's announcement is a sign that bounded decisions inside software are an important product category. Jev currently gives builders more public detail to work with: pricing, question types, failure modes, and an explicit multi-question SDK contract. The Decisions API accepts images in preview, which Jev does not. There is not yet enough public evidence to say which API performs better overall.

There is a technical distinction in how the vendors describe their systems: OpenAI says the Decisions API focuses Luna's intelligence on bounded answers; TypeSafe says Jev is a distinct model class trained to return decisions and probabilities rather than to generate text. We do not have an independent, matched evaluation showing that one architecture wins on speed, calibration, or accuracy.

OpenAI may also have a distribution advantage for teams that already use its platform; images are a confirmed input in the Decisions preview. But the regular Luna model's context window and API contract should not be assumed to carry over to Decisions until OpenAI publishes those details.

Which should you use?

Try Jev now if you can get early access or use a supported gateway, want published input pricing and a documented typed interface, or need to ask several questions about one text input. Test TypeSafe's calibration claims on your own data before using probability thresholds for consequential actions.

Watch or test the Decisions API if your decisions depend on images, your company already buys through OpenAI, or you have limited-preview access. Put it through the same evaluation harness once the request schema and pricing are public.

Keep your decision layer replaceable either way. Give your application a stable, bounded answer type, record the provider and model version, and treat probabilities or confidence as optional provider-specific fields until you have tested their meaning. My Jev app, Jevinik, supports Jev and a GPT-5 fallback behind one classification contract, selected with a decisionEngine field. That is an example of an application-level adapter, not evidence that the unreleased Decisions API is already a drop-in replacement.


The part neither of them does: knowing things

Here is what gets lost in the speed race. Neither Jev's documented decision endpoint nor the announced Decisions API describes built-in retrieval. Their decision task is to judge supplied context. This does not mean the regular GPT-6 Luna model cannot search: Luna supports web search via the Responses API, a different interface. Jev's own docs are blunt about state quality: "retrieve and filter in code first, and send only the fields the question needs."

Read that carefully and the conclusion is sharp. A fast decision on stale or wrong evidence can still be a fast wrong answer. The context you assemble constrains the quality of the decision made from it.

That is the layer we build at Valyu: search for AI knowledge work. Valyu's data catalog covers the web, SEC filings, PubMed, clinical trials, patents, academic sources, and market and macro data, with source URLs in Search results. Coverage and access vary by dataset and plan; some sources are only available inside DeepResearch. The pattern is two layers: retrieve relevant evidence, then make a bounded judgment over it.

Retrieval can remain a separate layer even if the decision provider's input/output adapter changes. Give either model concise, current source material first.

Python
from valyu import Valyu
from typesafe_sdk import Choice, Noul, TypeSafeClient
 
valyu = Valyu() # reads VALYU_API_KEY
jev = TypeSafeClient() # reads TYPESAFE_API_KEY
 
# 1. Retrieve relevant PubMed results; abstracts may supplement full-text chunks
hits = valyu.search(
"GLP-1 receptor agonists cardiovascular outcomes",
included_sources=["valyu/valyu-pubmed"],
include_abstracts=True,
start_date="2024-01-01",
max_num_results=20,
)
 
# 2. Judge a short excerpt; review the original paper before using the result
for doc in hits.results:
a = jev.system_one(
state={"title": doc.title, "excerpt": doc.content[:3000]},
questions={
"study_type": Choice(
instructions="Which study design is described in the excerpt?",
criteria={
"rct": "A randomised controlled trial",
"review": "A systematic review or meta-analysis",
"observational": "An observational study",
"other": "Another design or not enough information",
},
),
"cv_outcomes": Noul(
instructions="The excerpt reports cardiovascular outcomes"
),
},
).answers
if (a["study_type"].choice == "rct"
and a["study_type"].confidence > 0.8
and a["cv_outcomes"].noul > 0.8):
print(doc.title, doc.url) # shortlist, not a verified clinical conclusion

Jevinik applies the same broad pattern to stocks: its README documents five parallel Valyu searches (market data, company news, industry news, analyst views, macro risk), followed by a compact evidence state and a Jev judgment about whether a stock trades higher in 30 days. It currently offers a GPT-5 fallback, not a Decisions API integration. If you later switch decision engines, check each engine's output contract and measure it on the same evidence. The retrieval layer remains a separate part of the system.

To try the evidence layer, create an API key at platform.valyu.ai and read the Search documentation. Valyu's pricing page currently lists $10 in free credits for new accounts ($20 with a work email); check source entitlements before running the clinical-trials or other subscription datasets.

FAQ

What is the difference between Jev and the OpenAI Decisions API?

Both are designed to return bounded decisions rather than free-form prose. Jev is TypeSafe AI's text-only System One model with published pricing, question types, and multi-question calls; TypeSafe says its probabilities are calibrated. OpenAI's Luna-based Decisions API accepts text or images and is in limited preview; its public request schema, confidence semantics, and price are not yet published.

Is the Decisions API faster than Jev?

We do not have a matched head-to-head benchmark. Launch coverage reports about 150ms for Decisions API, while TypeSafe reports 70–500ms for Jev. The "10× faster" comparison is against Luna via OpenAI's regular API, not against Jev.

How much does the Decisions API cost?

OpenAI has not published a Decisions API price as of October 1, 2026. Its standard short-context GPT-6 Luna text prices are $0.10 per million input tokens and $0.50 per million output tokens, but those cannot be assumed to apply to Decisions.

How much does Jev cost?

$0.042 per million input tokens. Output tokens are free.

Does the Decisions API return confidence scores?

Some launch coverage says yes, but OpenAI has not documented the field, its meaning, or its calibration as of October 1, 2026. For comparison, Jev returns a probability distribution and derived confidence on Choice and Score; Noul returns only a yes probability. A confidence value is not itself a measured accuracy percentage.

Can Jev read images?

No. Jev is text only. Transcribe, OCR or caption images first. The Decisions API accepts images.

Is the Decisions API available now?

OpenAI announced a limited preview on September 29 and said broader release was planned "in the coming days." As of October 1, there is no public Decisions API integration guide or pricing entry in its developer documentation.

Is OpenAI's Decisions API a response to Jev?

OpenAI has not said it built or timed the Decisions API in response to Jev. Its announcement came 14 days after Jev's, and The New Stack framed the two as competitors; the product-development timeline remains unknown.

Can Jev or the Decisions API search the web?

No built-in search is described for Jev's decision endpoint or OpenAI's announced Decisions API. Pair a decision call with retrieval when the answer needs current evidence. GPT-6 Luna via the separate Responses API does support a web-search tool.

Should I replace my LLM with Jev or the Decisions API?

No. Use a bounded decision model when the answer space is known; use code for exact calculations and a generative model when you need text or open-ended reasoning. A routing layer can send hard cases to a larger model or a person, but test the whole workflow before assuming a cost or speed saving.

Valyu Add

Join 12,000+ professionals and knowledge workers.

Valyu Add is a free weekly research briefing for builders, investors and operators. Every issue is sourced, cited and verified with Valyu DeepResearch.