Essay June 2026 6 min read

What I learned about judgment by trying to give it to a machine

I set out to build an AI partner that could operate with judgment. I assumed the hard part would be making it intelligent. The hard part was deciding what judgment actually is — and most of the answer had nothing to do with intelligence.

The assumption I started with

I wanted a system that could sit beside the work I care about — customer operations, controls, modernization under pressure — and be useful the way a sharp colleague is useful. My first instinct was to make it more capable: more knowledge, more reasoning, more polish. I was treating judgment as an intelligence problem.

After enough iterations, the system was capable and still not trustworthy. It would answer confidently when it should have hesitated, optimize for sounding right over being right, and reach past boundaries it should have respected. Capability was not the missing piece.

Judgment turned out to be structure

What finally made the system trustworthy was not a better model. It was structure — the same structure that makes a human team reliable under pressure. I had to write down three things I had never written down for a team, because with people you lean on instinct and culture to carry them. A machine has neither.

A decision sequence, not a reflex

I gave it an explicit order of operations: name the objective, the constraints, the assumptions, the alternatives, the trade-offs, the risks — then the recommendation and how you would know it worked. Good operators do this without noticing. Writing it down for a machine showed me how often we let people skip it.

A values check before it acts

Before any recommendation, the system runs the question a good leader runs silently: is this truthful, responsible, secure — and does it actually improve the outcome, not just answer the question in front of it? If the answer is no, the recommendation is not ready. The discipline is not in having values. It is in checking against them before you commit, every time.

Permission to be uncertain

The most important rule I gave it was the freedom to say "I don't know yet." For a system that runs reviews and analysis, a confident wrong answer is far more expensive than an honest gap. So I made transparency about the limits of its knowledge a feature of the work, not a failure of it.

You learn what judgment is made of by trying to build it from nothing. Most of it turns out to be structure we never made explicit.

The part that surprised me

Building this was less a lesson about AI than a mirror held up to how we run teams. Everything I had to make explicit for the machine — the decision sequence, the values check, the permission to be uncertain — is something we leave implicit for the people doing customer-facing work, and then act surprised when it goes missing under pressure.

An agent who escalates the moment things get hard usually is not short on talent. They are short on structure: a clear decision path, explicit authority, and permission to say "I don't know, let me find out" instead of guessing. We rarely write those down. We hire for instinct and hope culture carries the rest. A machine has no instinct and no culture, so it forces you to make the operating standard visible — and once you have seen it written down, you notice how often your people are working without it.

What this means for AI in customer-facing work

Ask whether it operates to a standard, not whether it's smart

Capability is the easy part now. Whether a system follows a known decision path, respects its boundaries, and can be audited is what determines if you can trust it anywhere near a customer.

Treat a confident wrong answer as the failure mode

The dangerous output is not "I don't know." It is a fluent, certain answer that happens to be wrong. Reward honesty about limits the same way you would in a strong employee.

Make accountability and security defaults, not features

A system that acts on customer data should favor least privilege and leave an auditable trail by default. If those are bolt-ons, they will be missing in exactly the moment they matter.

The goal was never to replace judgment with a machine. It was to understand judgment well enough to build it — and to bring that clarity back to the people doing the work.

If you're sorting real AI value from confident-sounding theater, let's continue the conversation.

Email hello@jacobshields.com or use the contact page.