Introducing should() — a probabilistic if statement for AI agents
HyperMindZ · September 24, 2026
Today we're releasing should(), an open developer primitive for the decisions inside AI agents: should(question, context) → yes/no + confidence. It ships first as a Claude Skill you can install now, with the API and MCP surface opening in developer preview.
HyperMindZ — Thursday, September 24, 2026
Update, September 25: the thesis has sharpened since this launch — the Judgment Record is the product. Read the follow-up: should() v1: the Judgment Record.
Traditional software evaluates exact conditions. if (cpm <= maxCpm) has run the world for seventy years, and it isn't going anywhere.
But agentic software increasingly has to evaluate semantic conditions: should I call this tool? Does this satisfy the policy? Does this product match the brief? Should a human review this? Today those judgments are made in one of two unsatisfying ways — a frontier LLM call, which brings the full weight of a reasoning engine to a yes/no question, or a hand-written heuristic, which brings none of the judgment.
We think this class of step deserves its own primitive.
should()
const d = await should(
"Should this action execute?",
context
)
// → { yes: true, confidence: 0.97, decision_id: "dec_…" }
The mental model is a probabilistic if statement. A large share of agent steps have a small answer space — yes/no, approve/reject, route/escalate, match/no-match — with all the context needed to decide already in hand. should() makes that judgment explicit, bounded, and measurable: one question, one structured answer, one confidence score you can route on, one decision id in the log.
The one-line thesis: LLMs reason. Decision models decide. Governance permits. Agents act.
Why a primitive, and not another framework
A general-purpose LLM is built to interpret, reason, and generate. Decision models are optimized for bounded output — and for repetitive decisions they can be evaluated the way software should be: for accuracy, calibration, latency, and cost, against a labeled set.
should() is deliberately narrow. It is not a new agent framework, not a replacement for LLM reasoning, and not a replacement for deterministic business rules. Exact conditions stay exact; open-ended questions stay with the frontier model. The primitive covers the layer in between — and behind the one call you write, a runtime routes each decision to the cheapest reliable provider: a deterministic rule, a specialized decision model, a frontier LLM, or a human. The call never changes; the routing is free to improve.
One boundary we consider non-negotiable: a confidence score is an input to policy, never an authorization. A probabilistic model can recommend an action; deterministic governance — identity, authorization, policy, human approval — decides whether that recommendation is allowed to become one. High confidence proceeds under policy. Low confidence escalates. Nothing executes because a model felt sure.
What we're releasing today
The should Claude Skill. A single open file that teaches Claude to recognize when a step is a bounded decision, make the decision explicit rather than burying it in a chain of reasoning, and delegate it wherever a should() decision tool is available in the session. It also runs in reverse: point Claude Code at a project and ask it to find the places where a frontier model is being used for a bounded decision, and the skill turns that into a candidate-by-candidate audit.
mkdir -p .claude/skills/should
curl -o .claude/skills/should/SKILL.md \
https://www.hypermindz.ai/claude-skills/should/SKILL.md
A home page with a live simulator. The skill lives at hypermindz.ai/developers/skills/should — including an interactive, deterministic simulation of the API shape: one deal-to-brief question, and a single moved term flips the decision from a high-confidence yes to a high-confidence no.
The developer preview. The should() API and MCP surface are opening to a working cohort: early access, a labeled evaluation set for your decision class, and direct feedback into the provider interface. Request access here.
We're being deliberate about what this is: a preview of a primitive, not a platform. The interface is small on purpose, provider-independent by design, and the measurements — accuracy, calibration, latency, cost — will decide what gets built on top of it.
Why advertising first
We're proving the primitive where we already operate: the transaction infrastructure for agentic advertising, where machine decisions lead to real commercial actions under real governance. Deal-to-brief matching, counteroffer acceptance, IO-change checks, creative policy, review gating — the same bounded shape, hundreds of times a day, with every executed decision landing as a signed record.
But advertising is the first test, not the product boundary. The API carries no advertising-specific fields, and the skill knows nothing about media. If bounded decisions are a real layer of agentic software — and we think they are — the primitive should be useful anywhere an agent stands between reasoning and action.
The idea we want to test in the open
An SDK waits to be discovered. A skill is behavioral: it teaches an agent to recognize the moment a capability applies. That's the distribution idea we're most curious about — that the agent itself can know when the step in front of it doesn't need frontier reasoning.
Claude should know when Claude is overqualified.
If you're building agents — in advertising or anywhere else — install the skill, read the raw SKILL.md, and tell us where the primitive holds up and where it breaks. That feedback, not our roadmap, is what the preview is for.
