Re-writing the data-driven playbook for 2026
Give a smart new hire full access to your systems and nobody to ask — how far do they get? That's roughly how far your first agent gets, except the new hire tells you when they're confused. What AI actually needs from your data that BI never did, why an unopenable vendor is a procurement problem rather than a data problem, and a 30-day sequence that ends with something you can hand a builder.
Here is a test you can run on your own business this afternoon, without buying anything.
Give a smart new hire full access to your systems. Don't let them ask anyone a question. Could they do the job?
Wherever they get stuck is roughly where your first agent gets stuck. Same systems, same gaps, same ambiguities. But there is one difference, and it matters more than any of the rest: the new hire stops and asks which record is right. The agent picks one, confidently, and you find out three steps downstream. Confusion in a person is visible and self-correcting. In an agent it is silent, and it propagates.
That is why the readiness you built for BI is not the readiness you need now. A dashboard is a thing a human looks at, brings judgement to, and quietly corrects for. An agent has no judgement to bring and nothing to correct with. You can have every dashboard in the building, a warehouse, a data team, and a quarterly metrics review — and still not be able to ship a single agent.
What follows is what AI actually needs from your data, in the order it bites. Each one has a move you can make this month and an artifact you are left holding at the end of it.
1. Reachability: can a program get there at all?
Before quality, before structure, before anything: can software reach the record without a human in the middle?
There is a ladder, and every system you own sits on a rung of it:
Paper. A binder, a clipboard, a signed form in a filing cabinet.
Screen-only. A system where the data exists but the only way in is a person looking at a screen and typing. No export worth using, no API.
Scheduled export. A CSV that lands somewhere on a timer. Awkward, stale by design, but it is a start.
Direct database read. You can query the underlying store. Considerably better than it sounds — a read replica solves a lot of problems.
A real API. Documented, stable, covers the operations you need rather than the three the vendor found easy.
MCP or native agentic access. The system was built expecting a program to use it on your behalf.
The move: take one job. List every system it touches. Put each on a rung. The artifact: a reachability map — one page, and usually the most clarifying page anyone in the company has seen that quarter.
Most companies find something uncomfortable when they do this. The system holding the data that matters most is the one lowest on the ladder.
When the vendor is the ceiling
Here is the part that turns a data exercise into a business decision.
If your system of record sits at the bottom of that ladder and the vendor has no roadmap off it, the thing blocking you is a supplier contract, and it needs to be handled like one.
A great deal of legacy SaaS has no API worth the name, no MCP story, and no intention of building one. This is not incompetence. The business model was a locked-in seat and a locked-in database, and opening the data threatens both. Every year that stays true, that vendor becomes more expensive to you than the invoice suggests, because they are now a hard ceiling on everything you can automate.
So treat it like a supplier decision, on the timeline supplier decisions run on. Put API and agentic access into your renewal criteria. Ask every incumbent, in writing, what their roadmap for programmatic and agent access is. Note who answers with a date, who answers with a paragraph, and who does not answer. That last group is telling you whether they will still be your supplier in three years.
Three ways out, cheapest first
What you should not conclude is that you must replace the system before you can start. That conclusion is how this gets deferred for two years. There are three doors, and they cost wildly different amounts:
Route around it — days. Most companies discover the locked system gates far fewer jobs than they assumed. The billing platform with no API does not block the supplier correspondence, the quality reporting, the intake triage, or the tribal knowledge capture. Re-run your candidate list with that system marked unreachable and look at what is still standing. It is usually a lot.
Bridge to it — weeks. An export, a read replica, a report you pull on a schedule, occasionally a scraper you accept as explicitly temporary and budget to kill. Not elegant. Frequently correct.
Replace it — years. Sometimes genuinely the right answer. Never the first move, and never a prerequisite for starting.
The rule for choosing between the three doors matters more than the doors do. You do not pick your first job from what happens to be reachable. You generate candidates from your goals and the constraints blocking them, rank them by value, then apply reachability as a filter, and take the highest-value survivor. That ordering does most of the work, and it is the scoping funnel applied to this specific question.
Invert it and you get a pile of automations built because the data was handy. Each one demos well. None of them touches a number anyone reports on. That is agent sprawl arriving from a different direction, and it is the most common way "start small" goes wrong.
One more argument for routing around, and it is political rather than technical. A shipped win in an adjacent process is the thing that buys you the integration budget, or the vendor fight, later. Nobody funds a system migration off a slide deck. They fund it after something worked.
All of it points the same way. AI adoption works as a sequence of strategic increments rather than as a programme. The big-bang version — clean everything, replace the ERP, stand up the platform, and then begin — fails the way every big bang fails. It spends two years before the first piece of real evidence arrives, and by the time it does, the assumptions it was built on have gone stale. The value is in the transitions, and you only get to make one by making the last one.
2. Cleanliness: would any of it make sense to someone reading it?
Reachable is not the same as usable. This is where the new-hire test earns its keep, because the failures are exactly the ones that confuse a person:
Duplicates. Three customer records that are one customer, differing by a trailing space and a legal suffix.
Contradictions. The CRM says one thing, the spreadsheet says another, and there is no rule anywhere for which one wins. Everyone who has been here five years knows to trust the spreadsheet. Nobody wrote that down.
No sense of what is current. Fields that were abandoned in 2021 sitting beside fields still in use, with nothing to distinguish them.
Free text where structure was needed. Status captured in a notes field, in eleven different phrasings.
Drives where the filename is the only metadata, and the folder contains
four documents called some variant of final_v2.
A new hire hits any of these, frowns, and asks someone. Your agent hits them and resolves them — instantly, plausibly, and sometimes wrongly, with no frown to warn you.
The move: take five real, recent cases. Try to answer them using only what is in the systems. Ask nobody. The artifact: a written list of every point where you had to guess.
That list is your cleanup backlog, and notice what it is not. It is not an enterprise data quality programme. It is scoped to one job, it is usually a couple of weeks of work, and it is specific enough that someone can actually finish it.
3. Decision logs, not just outcome data
Your BI stack records what happened. An agent needs something your BI stack almost certainly threw away: what a good decision looked like.
That means three things together — the inputs that were available at the time, the choice somebody made, and what followed. Outcome data alone cannot teach it, because outcomes do not distinguish a good decision from a lucky one.
Most companies do capture this. It is just sitting in an email thread, a Slack channel, or a conversation nobody recorded. The buyer who decided not to expedite that order had reasons, weighed them, and was right. None of that is in the ERP.
The move: find where your target decision currently gets recorded, however informally, and start capturing it deliberately. The artifact: a decision record — inputs, choice, outcome — accumulating from today forward.
4. Ground truth, or you cannot grade anything
This is the one almost nobody has, and the one with the highest return.
You cannot ship what you cannot score. Without a set of cases where you know the right answer, every question about the agent becomes a matter of opinion: is it good enough, is it better than last week, did that change help or hurt, can we let it run without review. Teams without ground truth answer those questions with vibes, and vibes do not survive contact with a customer complaint.
With it, all of those become measurements. It is also what makes a model migration a routine afternoon instead of a leap of faith, and what lets you put a real threshold in a contract.
The move: assemble 50 to 100 historical cases with known-correct answers. Real inputs, real outputs, drawn from work your team already did. Do this before anything gets built. The artifact: an evaluation set.
If you build only one thing on this list, build this one. It is the most valuable data asset a company can create for itself, it does not expire, and it is worth more than the agent it grades.
5. Tribal knowledge, written down
The last category is the knowledge that never became a record at all — why this customer gets a different tolerance, why that job runs on machine three at reduced feed, which supplier's lead times you quietly pad. It lives in people, and every retirement is a data loss event.
The move: take the top five exception cases for your chosen job. Shadow the person who handles them, or interview them and write it down. The artifact: five written procedures.
Do not attempt this as a standalone programme. Documentation for its own sake fails everywhere it is tried. It works as a byproduct of an agent already in the daily flow of work — the knowledge surfaces because someone needed it to answer a real question, and gets retained on the way past.
What data-driven actually means now
None of the above makes you data-driven. Having data never did.
You are data-driven when a decision changes because of a measurement, on a named cadence, owned by a named person. All three parts. A measurement nobody looks at is decoration. A cadence nobody owns silently stops. An owner without a cadence means it happens when there is time, which is never.
Applied to AI specifically, that means: the eval runs on a schedule, someone reads the result, and a number moving the wrong way causes something to happen. That loop is the difference between an agent that improves for three years and one that quietly decays for three years while everyone assumes it is fine.
What to stop doing
Three things reliably delay a first agent by years, and all three feel responsible while they are happening.
The data lake first. You do not need a warehouse to ship your first agent. You need one job's data to be reachable and unambiguous.
The two-year master data management project. Sometimes necessary. Never a prerequisite. If it is genuinely required for your first use case, you have picked the wrong first use case.
"We need to clean everything first." This is where the apparent contradiction with section two resolves: clean the slice the job touches, not the enterprise.
Enterprise-wide cleanliness as a precondition is how this gets deferred forever, because it is never finished. Job-scoped cleanliness is a two-week task with an end state you can point at. Same discipline, three orders of magnitude apart in cost, and only one of them ships anything.
The 30-day version
If you want to actually do this rather than agree with it, here is the month. Pick one job first — highest value, per the rule in section one.
Week one. Build the reachability map for that job. Run the new-hire test on it. You now know what a program can reach and where a person would get stuck.
Week two. Write the guess list from the test, and set up the decision log so it starts accumulating from today. The guess list is your cleanup backlog.
Weeks three and four. Assemble the evaluation set. 50 to 100 historical cases with known-correct answers. This is the slowest and most valuable part, which is why it gets half the month.
At the end you are holding four things: a reachability map, a cleanup backlog, a decision log that is filling up, and an evaluation set.
That is not preparation for a project. That is most of what a competent builder would have asked you for in week one anyway — and you can now hand it to one, or use it to work out that the job is not worth building at all. Either answer is worth a month.
