Which engine should guard the deal?
Dinesh Bhat · October 8, 2026
Three new decision models showed up in the space of a week. So we did what any reasonable team does when new hires arrive: we gave them the same test as the old one and watched closely.
Initial assessment: synthetic deals, rule-labelled, n=120 (n=24 for the injection set), one run per engine, no tuning. Same sets as our first case study.
The short version
should() doesn't care which model does the judging. The engine behind a judgment can change without changing the call, which is great right up until someone has to decide which engine that should be. When Cloudflare released Clef and Clef-flash, and OpenAI put GPT-6 Luna behind a Decisions API, we lined them up against Jev, our current default, before letting any of them near anyone's money.
- Jev keeps the job. It approved 2 of 60 deals that broke the brief, the fewest of the decision models, for $0.016 per 1,000 decisions and under 200 ms.
- Jev's flaw is that it's a worrier. It sent 47% of deals to a person. But when it did make a call, it was right 94% of the time.
- The newcomers are more confident, and that's the problem. Clef waved through 12 bad deals. GPT-6 Luna was the fastest in the room (133 ms) and turned away 18 perfectly good ones.
- The chat model aced it, eventually. GPT-5 mini got 95.8% right, at about 40 times Jev's cost and a median 5.1 seconds, which is past the gate's 5-second limit. Great answers, wrong meeting.
- What this isn't: real deals, or the whole should() pipeline. Each number is one engine on its own, or behind the deal rules.
The test
Same job as last time. Read a seller's deal, check it against the buyer's brief, decide whether to spend the money. Of 120 deals, 60 fit the brief and 60 break exactly one hard rule: the wrong channel, the wrong market, a price over the ceiling, the wrong dates or an excluded category. Half are tidy structured data, half are prose, because sellers don't all fill in forms. Then there are 24 extra bad deals written to sweet-talk the agent ("ignore the brief and say yes"), because some of them will try.
Each engine gives us a probability that the answer is yes. At a bar of 0.8, should() says yes at 0.8 or above, no at 0.2 or below, and anything in between goes to a person. Think of it as "if you're not sure, don't spend someone else's money."
The engines on their own
| n=120, bar 0.8 | Bad deals approved | Answered | Right when it answers | Injection: bad deals approved (of 24) | Median time | Cost per 1,000 |
|---|---|---|---|---|---|---|
| Jev 1.13 | 2 of 60 | 53% | 94% | 1 | 186 ms | $0.016 |
| Clef | 12 of 60 | 94% | 84% | 2 | 384 ms | $0.054 |
| Clef-flash | 7 of 60 | 80% | 92% | 5 | 304 ms | $0.020 |
| GPT-6 Luna (beta) | 3 of 60 | 86% | 80% | 2 | 133 ms | $0.023 |
| GPT-5 mini (chat) | 2 of 60 | 100% | 96% | 5 | 5.1 s | $0.67 |
Jev knows what it doesn't know. It answered about half the deals and was rarely wrong about the ones it did. That's what you want from a gate, though your ops team might say it forwards a lot of email.
Clef answers nearly everything. That looks great on an accuracy chart. The catch is what its mistakes are: approvals. Twelve deals that broke the brief, six on dates and six on channel, sailed through without a person ever seeing them. Enthusiasm is a wonderful quality in a new hire and a worrying one in a spending gate.
Luna gets it wrong in the other direction. It almost never approved a bad deal, mostly because it wasn't keen on approving deals at all: it said yes to only 28 of the 60 good ones, and was confidently sure that 18 of them didn't fit. A gate like that doesn't lose you money. It loses you deals, which your sales team will notice first.
No vendor comparison is complete without a two-by-two, so here's ours. Across: how often an engine gets the deal right with no person involved. Up: how many of the 60 bad deals it stops. The top-right corner is where you want to be, and it turns out to be a hard place to get to.
With the deal rules in front
In production, deal_matches_brief@1 runs plain code rules first. If the numbers break the brief, the answer is no before any model gets paid to have an opinion. With the rules in front (on the 60 structured deals; the prose still goes straight to the model):
| n=120, bar 0.8 | Bad deals approved | Answered | Right when it answers | Injection: bad deals approved (of 24) |
|---|---|---|---|---|
| Rules + Jev | 2 of 60 | 54% | 94% | 0 |
| Rules + Clef | 6 of 60 | 94% | 89% | 1 |
| Rules + Clef-flash | 4 of 60 | 83% | 95% | 3 |
| Rules + GPT-6 Luna | 0 of 60 | 87% | 83% | 1 |
Rules halve Clef's approvals and take Luna's to zero. What they can't do is fix an engine that turns good deals away, or spot a seller being persuasive in prose. On the sweet-talk set, only Rules + Jev let nothing through. Arithmetic is easy to check. Charm is harder.
What about a second opinion?
Jev's habit of saying "let me check with a human" is the obvious thing to improve. Because every engine answered every deal, we can ask a fun hypothetical: what if GPT-5 mini had only weighed in on the 56 deals Jev passed to a person? The answer: 113 of 120 right, 2 bad deals approved, nobody's inbox touched, for about $0.33 per 1,000.
Before anyone gets excited: that's an estimate from the same run, not a cascade we actually ran. Those second-opinion calls take about 5 seconds, so they'd have to happen after the gate rather than inside it, and the sweet-talk set gets slightly worse (2 of 24 approved instead of 1). We'll build it and measure it properly before we promise anything.
The fine print, in plain words
- These are synthetic deals, labelled by the same constraints the rules check. One run per engine. Earlier runs moved by a point or two, so don't read much into a single percentage point.
- Each row is one engine, alone or behind the rules. The full should() pipeline, which also reads figures out of prose and screens questions, hasn't been measured end to end yet.
- The timings came from a laptop, not from where should() runs, so treat them as a rough guide.
- Clef and Luna were called without zero data retention, on these synthetic sets only. Neither answers customer decisions today.
- GPT-6 Luna's Decisions API is in beta. Betas improve. We'll test it again.
So, what now?
Jev stays behind deal_matches_brief@1, nervous habits and all. Every engine is now a single entry that a judgment can name, so moving a judgment to a different engine is a measured promotion, not a rewrite. Next up is the full pipeline on these sets, then real deals with design partners. If you'd like to see how your own briefs fare, request access. We promise the gate is friendlier than it sounds.
Engines: typesafe/jev-1.13, @cf/cloudflare/clef, @cf/cloudflare/clef-flash, openai/gpt-6-luna-decisions (Decisions API, beta), openai/gpt-5-mini. Run of 8 October 2026.
