Back to all articles
6 min read

Why Your AI Pilot Stalls Long Before Insight

Most AI pilots in investment firms die in the data room, not the demo. Here is the audit of data, workflows and decisions to run before you buy a tool.

Clean Data Is a Relationship, Not a State

A vendor demo takes forty minutes. The pilot that follows takes nine months and produces nothing anyone can trade on. That gap is the quietest story in investment technology right now, and the reason for it is almost never the reason people give.

Ask an operations lead why the pilot stalled and you will usually hear two words: dirty data. It sounds honest. It is also, most of the time, wrong. The data is often fine. What is missing is agreement on what the data is supposed to answer.

Where the Nine Months Actually Go

Here is a scene that repeats at firm after firm. A team buys a tool to surface signals from earnings calls and internal research notes. Before it can run, someone decides the note archive needs tidying. Tags get rewritten. Duplicates get merged. File names get standardized. Six months later the archive is beautiful, the tool is running, and the output is a list of themes every analyst already knew.

It also helps to spend a minute on az names worth knowing before settling on an approach.

Nothing was wasted on the cleanup, exactly. It just never touched the thing that mattered. Clean is not a property that data owns on its own. Clean is a relationship between a dataset and a specific question. A holdings file can be perfect for reconciliation and useless for attribution, because the sector tags were set at purchase and never refreshed. Same file. Same accuracy. Two completely different verdicts.

This is why "get the data ready" is such a dangerous project charter. Ready for what? Without that second half of the sentence, teams default to tidiness, because tidiness is easy to measure and easy to show a committee. Meanwhile the actual blocker sits untouched in a workflow diagram nobody drew.

Signs You Are Cleaning the Wrong Thing

A few tells show up early. None of them look like failure at the time, which is exactly the problem.

Firms that recognize these patterns early tend to move faster than better funded competitors, and it shows in how a few ai-driven wealth managers have quietly outpaced peers with far larger technology budgets.

<<>>

Audit the Decision Before the Dataset

Flip the order. Instead of asking what data you have, ask what choices your firm makes with real money attached. Then trace backward. This single reversal saves more pilots than any data platform ever has.

Map the Decisions That Move Money

Sit down with the investment team and list every recurring decision that changes a position, a weight, or a client's plan. Most firms land on somewhere between five and twelve. For each one, write down four things in plain language:

  • Who decides, and who can overrule them.
  • What they look at in the last hour before deciding, not the last month.
  • What they wish they could see but currently guess at.
  • How often the decision is wrong, and how you would know.

That fourth item stings, and it is the one that separates a real audit from a slide deck. If you cannot describe what a bad call looks like after the fact, no model will be able to either. The output has nothing to learn against.

What the Map Tends to Expose

Once the decisions are on paper, three things usually surface at once. First, several decisions run on data nobody stores. A portfolio manager's read on management tone, a client's off-hand comment about a coming liquidity need, the reason a name was passed on last year. All of it lives in heads and inboxes. Second, one or two datasets carry enormous weight and nobody has audited them in years. Third, the tool being evaluated touches none of the above.

"The firms that get value early are the ones that pick one decision, not one dataset, and rebuild everything around it," says Marisol Vega, director of portfolio operations at Zenith Investment Management. "We stopped asking whether our data was clean and started asking whether it was decisive. Very different question, very different project."

There is a second benefit to mapping decisions first. It gives you a natural place to measure results later, which matters enormously once metric most firms track starts shaping how leadership judges the whole program. A pilot tied to a named decision can be scored. A pilot tied to a dataset can only be admired.

The Pre-Purchase Checklist Nobody Runs

Before any contract gets signed, run a short audit on the one decision you selected. Two weeks is usually enough. It costs almost nothing and it kills bad projects while they are still cheap to kill.

Six Questions Worth Two Weeks

  • Lineage: for the three fields that matter most, can someone show where each value came from and when it last changed?
  • Disagreement: when two systems report the same holding differently, which one wins today, and who made that rule?
  • Timing: does the data arrive before the decision or after it? Late truth is not truth.
  • Context: is the reasoning behind past calls written down anywhere, or only the outcomes?
  • Access: can the people testing the tool reach the data without a ticket queue?
  • Reversibility: if the output is wrong on a Tuesday, what breaks and how fast can you unwind it?

Question four is the sleeper. Most firms hold years of outcomes and almost no reasoning. Models trained on outcomes alone learn what happened but never why, and the gap between those two is where judgment lives. Capturing the why is unglamorous work, but it compounds. It also changes how your material gets interpreted by machines generally, a dynamic that shows up anywhere machines read your content and decide what deserves attention.

Build the Smallest Thing That Could Work

With the audit done, resist the urge to scope big. Take one decision, two data sources, one reviewer, and a four week window. Define in advance what a useful result looks like and what result would make you walk away. Write both down before you start, because memory is generous after the fact.

Keep a running log of every moment a human overrides the output, along with the reason. That log becomes the most valuable asset the pilot produces, more valuable than the output itself. It tells you exactly where the model is blind and exactly where your people are anchored to habit. Which brings up the next uncomfortable truth, because the way the smartest analyst interacts with these systems can quietly degrade them, and that is the subject waiting in the next part of this series.

The Unsexy Work That Buys the Edge

The firms winning with AI in investment management are not the ones with the tidiest warehouses. They are the ones who were honest early about which decisions actually move performance, then went and looked hard at the handful of inputs feeding those decisions. Cleaning everything is a way of avoiding that honesty. It feels like progress, it fills status reports, and it postpones the moment someone has to admit that a core process runs on assumptions nobody has tested since the last cycle.

Start narrow and start with a choice, not a spreadsheet. Name the decision, trace what feeds it, write down what wrong looks like, then buy the smallest tool that could plausibly help. The edge is not in the model. It is in knowing precisely what you want the model to be right about, and being able to prove afterward whether it was.

Argue with this in person

Barcamp Boston runs on hallway disagreements and half finished ideas. Pitch a session, grab a slot on the board, and take the conversation further than a comment box ever will.

See the schedule