HyperMindZ

should() · judgment as software

Developer preview

Ship judgments you can prove.

Agentic software runs on recurring semantic judgments — does this satisfy the policy, does this product match the brief, should this escalate? should() started as a probabilistic if statement. It has grown into where those judgments are defined, tested, run, measured, and evolved — powered by Jev, the decision engine underneath. The durable artifact isn't the model call. It's the judgment, and the evidence of how it performs.

should() · judgment deal-fit:v3Simulation — deterministic, no model called
await should(
  "Does this seller product satisfy the buyer brief?",
  { brief, product }
)
brief
{
  objective: "Reach sports fans",
  channel: "ctv",
  maxCpm: 32.00,
  geo: "US"
}
product
{
  channel: "ctv",
  geo: "US",
  offerCpm: 28.00
}
Change the offer — the decision follows
decision
YES· 0.97
caller threshold 0.85
→ proceed, under policy
{ judgment: "deal-fit:v3",
  yes: true, confidence: 0.97,
  decision_id: "dec_7f29a" }

LLMs reason. Decision models decide. Governance permits. Agents act.

The idea

A large share of agent steps have a small answer space.

Yes/no. Approve/reject. Route/escalate. Match/no-match. A general-purpose LLM is built to interpret, reason, and generate — and for these steps, that's more machine than the job needs. A decision model is optimized for bounded output and returns a confidence you can route on: proceed above the threshold, escalate below it. The wedge is the recurring judgment — high-volume, consequential, measurable, frequent enough that scattered prompts and ad-hoc evaluation create real engineering pain. One-off judgments are not the point.

01

Tool authorization

“Should this action execute?” The step every agent loop repeats hundreds of times — a yes/no with the context already in hand.

02

Policy checks

“Does this creative satisfy the relevant policy?” The policy text and the candidate are both present; the answer space is two values.

03

Match decisions

“Does this seller product satisfy the buyer brief?” Relevance judgments between two structured objects — no prose required.

04

Escalation gates

“Is confidence high enough to skip human review?” The decision about the decision — routed by threshold, not by vibes.

What it is

One horizontal primitive.

should(question, context) → yes/no + confidence. That's the whole developer surface — but the call is an instance of a named, versioned judgment, not an ad-hoc question. Jev makes the judgment. should() makes the judgment manageable as software: defined, tested against evidence, run in production, measured against outcomes. The call you write never changes; the judgment underneath it versions and improves.

What it is not

  • A multi-model router, prompt manager, or general context engine — Jev is the V1 engine, and model-independence is earned, not claimed: a judgment passes that test only when it can move engines without losing its definition, tests, history, or outcomes
  • A new advertising protocol, bidder, or RTB decision engine
  • A replacement for LLM reasoning — open-ended questions stay with the frontier model
  • A replacement for deterministic business rules — exact conditions stay exact
  • Automatic execution because a model reports high confidence

The product

The Judgment Record is the product.

Everything else is in service of making it exist, stay current, and stay model-independent. A judgment isn't a prompt someone remembers — it's a record:

Definition

What the judgment means and the question being answered

Versions

How the definition, context contract, or threshold changes over time

Test evidence

Validated examples, edge cases, expected decisions, and holdouts

Context

What information the judgment expects and what was supplied

Decisions

Production result and confidence for each call

Failures

Where the judgment does not perform as expected

Corrections

Human disagreement or explicit correction

Outcomes

What happened after the decision, when observable

Where outcomes are delayed, ambiguous, or missing, should() exposes that fact rather than manufacturing certainty — a judgment with weak outcome coverage is visibly marked as having insufficient production evidence.

Judgment healthillustration
support-escalation:v3 test quality 96.1% production decisions 18,421 known outcomes 12,814 (69.6%) unknown outcomes 5,607 (30.4%) measured prod quality 94.8% evidence status HEALTHY

How improvement works

Judgments don't mutate. They version.

should() can recommend improvements — but every change is a candidate, never a silent mutation. Candidates are versioned, evaluated against historical and holdout evidence, reversible, and logged. A context change is a judgment change, so it forces a new version and re-calibration — a moving input distribution can't silently invalidate prior evidence.

current: support-escalation:v3
  │  should() identifies a possible improvement
  ▼
candidate: support-escalation:v4
  │  historical + holdout + production evidence
  ▼
better by predefined criteria?  → promote v4
not better?                     → discard candidate
                                  v3 stays available for rollback

The record compounds: as a judgment accumulates validated examples, decisions, failures, corrections, and outcomes, a better Jev release can be tested against it — and one day, another engine can be too. The judgment survives the model.

How the skill fits

Claude should know when Claude is overqualified.

An SDK waits to be discovered. The skill is behavioral guidance: when a step is a bounded decision with sufficient context, Claude considers delegating it. The MCP interface is the capability it calls; the underlying should() API stays portable beyond Claude.

CLAUDE                      recognizes a bounded decision
  │
  ▼
SHOULD SKILL                the behavioral trigger — this page's skill
  │
  ▼
should() MCP                the capability Claude calls
  │
  ▼
DECISION RUNTIME            rules · decision models · LLMs · humans
  │
  ▼
YES · 0.98                  + decision_id, logged
  │
  ▼
CLAUDE CONTINUES            frontier reasoning where it's actually needed

The same recognition works in reverse: point Claude Code at a project and ask it to find the places where a frontier model is being used for a bounded decision. The skill turns that into a candidate-by-candidate audit — which calls are really a yes/no, whether their context is self-contained, and what the substitution would look like.

claude code — ~/your-agent
> find places where I'm using a frontier model for a bounded decision

Scanning call sites…

agent/router.ts:41     "execute this tool?"        yes/no    → should()
support/escalate.ts    "route or escalate?"        enum(4)   → decide()
policy/check.ts:88     "creative meets policy?"    yes/no    → should()

12 of 214 call sites consume a bounded answer.
Review the candidates? (y/n)

Illustration — the audit loop running on a sample project.

Where the line sits

The decision says what appears appropriate. Governance says what is permitted.

A confidence score is an input to policy — never an authorization. A probabilistic model can recommend an action; deterministic governance decides whether that recommendation is allowed to become one. Identity, authorization, policy, and human approval sit between every YES and every execution, and the decisions that do execute land as signed records.

Bounded, not blind

should() answers small questions with the context in hand. Open-ended judgment, incomplete context, and matters of strategy stay with reasoning — the skill encodes when not to delegate.

Confidence routes, policy decides

High confidence proceeds under policy; low confidence escalates to deeper reasoning or a human. The threshold is the caller's policy decision, never the model's.

Every decision on the record

Each call returns a decision_id and lands in a decision log — so the judgments inside an agent become as auditable as the transactions it executes.

First proving ground

Advertising is the first test, not the product boundary.

We're proving the primitive where we already operate — where machine decisions lead to real commercial actions, under real governance. The API itself carries no advertising-specific fields. The simulator at the top of this page is the first demo, verbatim: one deal-to-brief question, and a single moved term — CPM beyond the authorized range — flips the decision.

Counteroffers, IO-change checks, creative policy, relevance filtering, review gating — the same shape, over and over. Each one today is either a frontier call or a hand-written heuristic.