The first question leadership teams ask me is almost always some version of “which AI use case should we start with?” It is a reasonable question, and it is premature. Use-case selection is the second decision. The first is an honest reading of what the organization can currently metabolize — because the same use case that compounds inside a ready organization becomes an undead pilot inside an unready one.

So every engagement starts the same way: two weeks, five dimensions, each scored 1 to 5. The composite — I call it the AI Readiness Score — predicts outcomes better than any technology choice, and more usefully, the lowest dimension tells you exactly what to do first. (A three-question quick version lives in the Lab; this essay is the full instrument behind it.)

The five dimensions

1. Data foundations

Not “do you have a data lake” — most companies scored highly by their own data teams still fail this dimension where AI actually needs it. The question is whether the data that describes your core workflows is accessible, accurate, and fresh enough to act on. A 2 looks like: the answer to any operational question requires an analyst and a week. A 4 looks like: the systems of record agree with each other, someone owns each critical dataset, and a new integration is a task, not a project. You cannot retrieve your way around records that are wrong.

2. Workflow clarity

AI automates workflows, and you cannot automate what you cannot describe. A 2 is an organization where the process lives in the heads of three tenured operators and every walkthrough produces a different diagram. A 4 is an organization that knows its volumes, exception rates, and handoffs — where someone can say “step four is 30% of cycle time and 70% of errors” with numbers behind it. Low workflow clarity is the most common silent killer: the pilot gets built against the imagined process and dies against the real one.

3. Decision rights and governance

When an AI system produces a wrong answer with consequences, who finds out, who decides what changes, and how fast? A 2 is governance by escalation and improvisation — every incident convenes an ad-hoc meeting of everyone. A 4 has boring, written answers: quality bars set in advance, an owner per deployed system, a defined path from “the model is wrong” to “the model is fixed.” Teams over-invest in AI ethics theater and under-invest in this mundane operational version, and it is the mundane version that determines whether anything reaches production.

4. Capability density

Not headcount with AI titles — fluency in the teams that own the work. Can the product trio closest to your core workflow read an eval, reason about failure modes, and estimate cost per task without filing a ticket to a specialist group? A 2 is an organization where all AI knowledge sits in one team that everything queues behind. A 4 is one where the median product team has shipped and operated something probabilistic, however small. Density beats specialization because compounding happens where the work is, not where the expertise is warehoused.

5. Adoption and trust posture

The dimension every technical assessment skips, and the one that most often decides the outcome. What happened to the last three tools rolled out to the frontline — adopted, tolerated, or quietly worked around? Do operators own the errors of tools they are told to use? Is there a correction path that costs nothing socially? A 2 is a workforce with scar tissue from transformation programs past; a 4 is one where frontline teams have visibly shaped a rollout before and trust that a fallback path exists. AI lands on whatever trust surface already exists. It does not create one.

The score's real output is not the number. It's the sequence: your lowest dimension is your first project — and it is almost never a model.

How the two weeks run

The method is deliberately unglamorous. Week one: read the artifacts the organization already produces — roadmaps, incident channels, ops dashboards, the last three pilot post-mortems if they exist (their absence is itself a data point) — and walk the top two workflows end to end, sitting with the people who run them. Week two: structured interviews across the leadership seam — product, engineering, ops, finance, frontline management — scoring each dimension from evidence, not self-report. Self-assessment inflates every dimension by roughly a point; the artifacts and the workflow walk are the correction.

Reading the composite

  • Below 2.5 — build foundations, don't buy capability. An organization here that funds an ambitious AI program is pre-paying for a purgatory pilot. The right spend is on the lowest one or two dimensions: instrument the workflow, fix the record systems, name the owners. Six months of unglamorous work moves the score a full point and changes what every later dollar buys.
  • 2.5 to 3.5 — one narrow wedge, run properly. Ready enough to learn from production, not ready to parallelize. Pick a single high-clarity workflow, run the disciplined 90-day pattern with a hard scale-or-kill gate, and use the pilot deliberately as a readiness-building instrument — it should raise governance and adoption scores whether or not it scales.
  • Above 3.5 — the constraint is ambition, not readiness. Organizations here usually under-invest out of habit, running cautious pilots when they could be re-architecting workflows. The work at this level is portfolio construction and pace: multiple concurrent bets, explicit horizons, kill discipline.

A pattern worth naming: scores cluster tighter than teams expect. The modal mid-market organization I assess lands between 2.0 and 2.5 — genuinely capable people, real data assets, and one or two foundational gaps that would silently cap any AI investment made on top of them. Telling a leadership team eager for an AI roadmap that their first project is workflow instrumentation is not the fun version of this job. It is the useful version. The score exists to make that conversation a matter of evidence instead of opinion — and to make the next assessment, a year later, show a number that moved.