should() · judgment as software
Developer previewShip judgments you can prove.
Agentic software runs on recurring semantic judgments — does this satisfy the policy, does this product match the brief, should this escalate? should() started as a probabilistic if statement. It has grown into where those judgments are defined, tested, run, measured, and evolved — powered by Jev, the decision engine underneath. The durable artifact isn't the model call. It's the judgment, and the evidence of how it performs.
await should(
"Does this seller product satisfy the buyer brief?",
{ brief, product }
){
objective: "Reach sports fans",
channel: "ctv",
maxCpm: 32.00,
geo: "US"
}{
channel: "ctv",
geo: "US",
offerCpm: 28.00
}{ judgment: "deal-fit:v3",
yes: true, confidence: 0.97,
decision_id: "dec_7f29a" }LLMs reason. Decision models decide. Governance permits. Agents act.
The idea
A large share of agent steps have a small answer space.
Yes/no. Approve/reject. Route/escalate. Match/no-match. A general-purpose LLM is built to interpret, reason, and generate — and for these steps, that's more machine than the job needs. A decision model is optimized for bounded output and returns a confidence you can route on: proceed above the threshold, escalate below it. The wedge is the recurring judgment — high-volume, consequential, measurable, frequent enough that scattered prompts and ad-hoc evaluation create real engineering pain. One-off judgments are not the point.
Tool authorization
“Should this action execute?” The step every agent loop repeats hundreds of times — a yes/no with the context already in hand.
Policy checks
“Does this creative satisfy the relevant policy?” The policy text and the candidate are both present; the answer space is two values.
Match decisions
“Does this seller product satisfy the buyer brief?” Relevance judgments between two structured objects — no prose required.
Escalation gates
“Is confidence high enough to skip human review?” The decision about the decision — routed by threshold, not by vibes.
What it is
One horizontal primitive.
should(question, context) → yes/no + confidence. That's the whole developer surface — but the call is an instance of a named, versioned judgment, not an ad-hoc question. Jev makes the judgment. should() makes the judgment manageable as software: defined, tested against evidence, run in production, measured against outcomes. The call you write never changes; the judgment underneath it versions and improves.
What it is not
- A multi-model router, prompt manager, or general context engine — Jev is the V1 engine, and model-independence is earned, not claimed: a judgment passes that test only when it can move engines without losing its definition, tests, history, or outcomes
- A new advertising protocol, bidder, or RTB decision engine
- A replacement for LLM reasoning — open-ended questions stay with the frontier model
- A replacement for deterministic business rules — exact conditions stay exact
- Automatic execution because a model reports high confidence
The product
The Judgment Record is the product.
Everything else is in service of making it exist, stay current, and stay model-independent. A judgment isn't a prompt someone remembers — it's a record:
What the judgment means and the question being answered
How the definition, context contract, or threshold changes over time
Validated examples, edge cases, expected decisions, and holdouts
What information the judgment expects and what was supplied
Production result and confidence for each call
Where the judgment does not perform as expected
Human disagreement or explicit correction
What happened after the decision, when observable
Where outcomes are delayed, ambiguous, or missing, should() exposes that fact rather than manufacturing certainty — a judgment with weak outcome coverage is visibly marked as having insufficient production evidence.
support-escalation:v3
test quality 96.1%
production decisions 18,421
known outcomes 12,814 (69.6%)
unknown outcomes 5,607 (30.4%)
measured prod quality 94.8%
evidence status HEALTHYHow improvement works
Judgments don't mutate. They version.
should() can recommend improvements — but every change is a candidate, never a silent mutation. Candidates are versioned, evaluated against historical and holdout evidence, reversible, and logged. A context change is a judgment change, so it forces a new version and re-calibration — a moving input distribution can't silently invalidate prior evidence.
current: support-escalation:v3
│ should() identifies a possible improvement
▼
candidate: support-escalation:v4
│ historical + holdout + production evidence
▼
better by predefined criteria? → promote v4
not better? → discard candidate
v3 stays available for rollbackThe record compounds: as a judgment accumulates validated examples, decisions, failures, corrections, and outcomes, a better Jev release can be tested against it — and one day, another engine can be too. The judgment survives the model.
How the skill fits
Claude should know when Claude is overqualified.
An SDK waits to be discovered. The skill is behavioral guidance: when a step is a bounded decision with sufficient context, Claude considers delegating it. The MCP interface is the capability it calls; the underlying should() API stays portable beyond Claude.
CLAUDE recognizes a bounded decision
│
▼
SHOULD SKILL the behavioral trigger — this page's skill
│
▼
should() MCP the capability Claude calls
│
▼
DECISION RUNTIME rules · decision models · LLMs · humans
│
▼
YES · 0.98 + decision_id, logged
│
▼
CLAUDE CONTINUES frontier reasoning where it's actually neededThe same recognition works in reverse: point Claude Code at a project and ask it to find the places where a frontier model is being used for a bounded decision. The skill turns that into a candidate-by-candidate audit — which calls are really a yes/no, whether their context is self-contained, and what the substitution would look like.
> find places where I'm using a frontier model for a bounded decision
Scanning call sites…
agent/router.ts:41 "execute this tool?" yes/no → should()
support/escalate.ts "route or escalate?" enum(4) → decide()
policy/check.ts:88 "creative meets policy?" yes/no → should()
12 of 214 call sites consume a bounded answer.
Review the candidates? (y/n)Illustration — the audit loop running on a sample project.
Where the line sits
The decision says what appears appropriate. Governance says what is permitted.
A confidence score is an input to policy — never an authorization. A probabilistic model can recommend an action; deterministic governance decides whether that recommendation is allowed to become one. Identity, authorization, policy, and human approval sit between every YES and every execution, and the decisions that do execute land as signed records.
Bounded, not blind
should() answers small questions with the context in hand. Open-ended judgment, incomplete context, and matters of strategy stay with reasoning — the skill encodes when not to delegate.
Confidence routes, policy decides
High confidence proceeds under policy; low confidence escalates to deeper reasoning or a human. The threshold is the caller's policy decision, never the model's.
Every decision on the record
Each call returns a decision_id and lands in a decision log — so the judgments inside an agent become as auditable as the transactions it executes.
First proving ground
Advertising is the first test, not the product boundary.
We're proving the primitive where we already operate — where machine decisions lead to real commercial actions, under real governance. The API itself carries no advertising-specific fields. The simulator at the top of this page is the first demo, verbatim: one deal-to-brief question, and a single moved term — CPM beyond the authorized range — flips the decision.
Counteroffers, IO-change checks, creative policy, relevance filtering, review gating — the same shape, over and over. Each one today is either a frontier call or a hand-written heuristic.
