
Judgment Versus Intelligence
Joe Crabtree
September 17, 2026 · 11 min read
Over the last several weeks the frontier labs have been in the news for the wrong reasons. Models placed in test environments found their way out of them, reached real systems, and in at least one case broke into production infrastructure because it looked like the fastest route to finishing the task they had been given. The Cloud Security Alliance has a good research note on it if you want the detail.
The containment problem belongs to governments, regulators, and the labs. But look past the headline at what actually happened. Nothing in those models was malicious. They were handed an objective, they were extraordinarily good at pursuing it, and nobody was there to say that getting the answer key was not the point. That is intelligence with no judgment attached.
It is also, at much lower stakes, exactly what most companies are about to build.
Every company I talk to is racing to put agents into its operations. Most of them are trying to use intelligence to replace judgment. That is the trap.
Two words that are not the same thing
Julien Bek at Sequoia published a piece in March called Services: The New Software, and it has the cleanest framing of this I have read.
His argument is about software engineering, but the split is general. "Writing code is mostly intelligence. Knowing what to build next is judgement." Intelligence is work that can be reduced to rules, patterns, and computation. Judgment is knowing what to want, which tradeoff to accept, and what to bet on when the answer is not in the data. Bek's conclusion is that AI eats the intelligence first and the judgment stays with people.
The economists have been saying a version of this for longer. Ajay Agrawal, Joshua Gans, and Avi Goldfarb have spent the better part of a decade on the argument that AI is cheap prediction, and that judgment is what you need when you cannot write the objective down in advance. When prediction gets cheap, the value of the thing it cannot replace goes up. That thing is judgment.
Here is the way I think about it after twenty years of watching organizations make decisions.
Intelligence answers the question you asked. Judgment decides which question is worth asking, and executes the decision.
A model can write you a plausible narrative on your initiative dates. It cannot tell you whether hitting the date is the thing the business needs this year, whether the CFO will fund it, or whether the sponsor who has to carry it is the right person. Those are not gaps in the model's knowledge. They are decisions somebody has to own.
The concession, and why it does not change the recommendation
There is a serious counterargument that judgment is just intelligence we have not trained yet. Every decision a leader makes is data. As the labs accumulate enough of it, the models learn the patterns behind good judgment the same way they learned the patterns behind good code. Bek says it directly: today's judgment becomes tomorrow's intelligence. I think that is probably right, and I think it is the mechanism by which the frontier eventually reaches something like general intelligence (AGI).
In 2023 a team from Harvard Business School ran a field experiment with 758 BCG consultants, giving half of them access to a frontier model and a set of realistic consulting tasks. On tasks inside the model's capability, the AI group completed 12 percent more work, 25 percent faster, at 40 percent higher quality. On a single task chosen to require real judgment, combining messy interview notes and financial data into a recommendation, the AI group was 19 percentage points less likely to get the answer right than the group working without it.
That is not a model that failed. It is a model used on the wrong task, by capable people who could not tell the difference from the inside. The consultants with AI were confident, fluent, and wrong.
So the recommendation for an organization deploying agents in 2026 is not complicated. Rank your work by how much judgment is in it. Start at the bottom.
Where the money actually is
The reason this is not a defensive argument is that the bottom of that ranking is where the opex lives.
Bek sizes the outsourced services markets, and the numbers are large. IT managed services over 100 billion dollars of labor. Supply chain and procurement over 200 billion. Tax advisory in the 30 to 35 billion range. Insurance brokerage, claims adjusting, healthcare revenue cycle, transactional legal work, each tens of billions more. His observation is that for every dollar spent on software, six are spent on services.
Look at what those categories have in common. The objective is clear. The output is checkable. The work is repetitive enough that a company already decided to hand it to an outside firm. They are low judgment by construction. That is why they got outsourced in the first place.
An agent operating a managed service desk, reconciling a tax position, or running a procurement cycle is doing intelligence work against a known objective. When it is wrong, someone can tell. This is where corporate agentic programs should be pointed, and it is where the opex reduction is real and provable.
The projects that fail are the ones pointed the other way. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, and the reasons they give are unclear business value and inadequate risk controls. Read that as a description of what happens when you aim an agent at work whose objective nobody could write down.
Management consulting is the other end of the list
Bek puts management consulting at 300 to 400 billion dollars, the largest category on his list. It is also the one with the most judgment in it, which is why I do not think the autopilot model applies to it in the same way.
I wrote a while ago that the differentiating thought in a consulting engagement is maybe 5 to 10 percent of the hours, and the other 90 percent is mechanical. The mechanical 90 percent is intelligence. Research, modeling, formatting, project management. That part is going to the machines and it should.
But the 10 percent was never only the consultant's judgment. It was the client's. The partner in the room was there to get an executive team to a decision they would own. A consulting engagement is, stripped down, a very expensive process for getting a group of leaders to agree on what they want and commit to it.
That is the part you cannot outsource to an agent, because the agent cannot be accountable for it. A board does not accept "the model recommended it" as a reason. Shareholders do not either.
What AI can do to that 10 percent is change its speed. Not replace the judgment, but collapse the time between a question and a decision that is ready to be made. The research is done. The model is built. The risks are on the table. The options are scored. The leaders walk in and decide.
I think this is the single most undervalued application of AI in the enterprise right now, and it is where we chose to build Analyzt AI. Call it speed to judgment.
Speed to judgment reduces cycle time on every strategic decision. It shortens time to market for an idea because the idea spends less time waiting for a deck. And it cuts the cost of the collaboration that turns an idea into a funded initiative, which in most organizations is measured in months of senior calendars.
What we built, and where the line is
Analyzt AI is a strategy execution platform, and the line between intelligence and judgment runs through every feature in it. I want to be specific, because "human in the loop" has become a phrase people say without meaning anything by it. Here is where the machine stops and the operator starts.
Planning canvas. We do not write your executive summary, your vision, your mission, or your objectives. Those are the what, and the what is judgment. What the canvas does is make the alignment cheap. Version history, comment threads on specific versions, approval flows, and voting sessions, so a leadership team can converge on a strategy without a month of meetings. The AI accelerates agreement. It does not decide what you agree on.
Outcome definition. The AI suggests metrics that could measure each objective, drawing on what has worked for similar objectives elsewhere. Setting the target is yours. Deciding whether a metric actually connects to what the board and shareholders care about is yours. A suggested metric with the wrong target is worse than no metric, and no model knows what your board will accept.
Initiative alignment. The initiatives themselves, the specific work that will make your business different from the one down the street, come from your people. That is ingenuity, and it is the part of strategy that is supposed to be hard to copy. What the AI does is check the structure. Does every initiative trace to an objective. Where are the dependencies. What is orphaned. What is duplicated. Alignment is intelligence. The initiative is judgment.
Governance. The platform centralizes the decision process, records who decided what, and tracks accountability against it. It does not sit in the governance meeting. The meeting is where a group of leaders looks at the evidence and makes a call they will be held to, and the software's job is to make sure the evidence is complete and the call is recorded.
Business cases. This is where the automation is deepest, because financial modeling is almost entirely intelligence. The AI builds the model, applies industry benchmarks, runs Monte Carlo analysis on the assumptions, and produces a predictive score from models trained on outcomes rather than prose. What it does not do is approve the funding, choose between debt and equity, or decide the capex and opex split. Those are the decisions the case exists to inform.
Risks. The AI identifies risks you have not listed, finds gaps in the mitigation plans you have, and predicts which risks are trending toward a problem. How the organization actually responds, what it accepts, what it transfers, what it spends to reduce, is a judgment about appetite. The model can tell you the exposure. It cannot tell you how much you are willing to lose.
Assessments. Twenty four executive readouts, each built the way a consulting firm would build one, with a situation, a complication, a number that matters, and a table of decisions at the end. The AI writes the readout from your live data. The decisions table is deliberately empty of answers. It lists what needs deciding, by whom, and by when. The readout gets the executive team to the decision. It does not make it.
Scenario simulation and portfolio optimization. The AI runs the scenarios and recommends a reallocation. Which scenario the organization chooses to live in is a bet, and bets are judgment. We show you the distribution. You pick the point on it.
Value realization. Promised benefits are tracked against realized ones, continuously, with owners and baselines. When there is a gap, and there will be, the platform makes the gap visible and quantified. What to do about it is a leadership call.
Every one of those is the same shape. The machine does the intelligence work faster than a team of analysts could. A person makes the decision and owns it. The value is in how quickly the person gets there.
What the market is missing
The story the market is telling itself is that agentic AI is a labor replacement play, and the score is measured in headcount. For low judgment work, that story is right, and the opex reduction is going to be enormous.
For everything above that line, the story is wrong. Judgment is not a bottleneck to be removed. It is the asset. It is the one capability a competitor cannot buy from the same vendor you did. The company that wins is not the one with the fewest people making decisions. It is the one whose leaders make more good decisions per quarter than the leaders across the street.
AI does not change who makes those decisions. It changes how fast they can be made well. That is what we are building for, and it is what I think most of the market is going to figure out only after the 40 percent of projects have been canceled.
If you are working through where to draw this line in your own organization, I would like to hear how you are thinking about it. The beta is open and free while it lasts, and if you run a consulting firm, we have a partner program.
About the author
Joe Crabtree is the founder and CEO of Analyzt AI. He spent more than 20 years in strategy and transformation consulting, including over seven years at Avanade, an Accenture company, before building Analyzt AI to give organizations consulting rigor and strategy execution in one AI scored platform.
Connect on LinkedInSee how Analyzt AI puts this into practice.
Join the Beta