Back to all articles
6 min read

Adoption Theater And The AI Metrics That Actually Pay

Usage stats make AI look like a win while alpha stays flat. Here is the attribution and governance setup that ties investment AI to real results.

The Adoption Trap Hiding In Plain Sight

Part two looked at the people side of this: how a sharp analyst can quietly steer a model toward the answer they already liked. Now comes the part that decides whether any of the work lasts. How you keep score.

Walk into almost any firm running an AI pilot and ask how it is going. You will get a number. Seats activated. Prompts run last month. Hours saved. Documents summarized. The chart points up and to the right, the committee nods, and the budget gets renewed for another year.

Not one of those numbers has anything to do with returns.

Usage is an input. It tells you the tool is being touched. It says nothing about whether a position size got better, whether a weak name got caught two weeks earlier, or whether the team reached a conviction call faster than it used to. A firm can post 90 percent adoption and completely flat alpha. That happens far more often than anyone admits on a conference panel.

“The moment a firm tells me their AI is working because everyone logs in, I know nobody has tied it to a decision yet,” says Marcus Feld, director of portfolio technology at Zenith Investment Management. “Usage is the easiest thing to measure and the least useful thing to report.”

Why The Usage Chart Is So Seductive

This is not a story about careless firms. Usage metrics win because they behave well, and well behaved numbers get reported.

  • They move fast. Alpha needs quarters or years to speak. Login counts move in a week.
  • They are clean. No benchmark argument, no attribution fight, no awkward questions about luck.
  • They flatter everyone. The vendor looks good, the sponsor looks good, the team looks busy.
  • They shift blame. Weak usage becomes a training problem, never a tool problem.

So the pilot gets graded on effort instead of output. Effort metrics stick around because they let a firm avoid the honest conversation about whether judgment actually got sharper, which is the exact ground the smartest analyst pulls apart.

Attribution Rules That Survive An Audit

The fix is simple to describe and uncomfortable to run. You stop logging activity and start logging decisions.

Build A Decision Ledger First

Pick a small set of repeating decisions the model touches. Screening a new name. Sizing a position. Flagging a covenant risk. Drafting a client review. For each one, capture a few fields at the moment the decision happens, never afterward from memory:

  • The decision, the owner, and the date
  • What the model produced, in one plain line
  • What the human changed, and the stated reason
  • How long the same task used to take
  • The expected outcome and the date it will be checked

Six months of that gives you something no dashboard can fake: a paired record of machine output, human edit, and result. That is attribution. Everything else is a feeling.

You also need a baseline, and this is where most teams cut corners. Keep a control set. Research a slice of names the old way on purpose. Compare error rates, speed, and coverage between the two groups. Firms that skipped the foundation work behind the clean data delusion usually find their ledger unreadable, because nothing was time stamped or sourced properly in the first place.

Five Numbers Worth Bringing To The Investment Committee

  • Decision hit rate delta. Same decision type, model assisted versus not, measured against outcomes twelve months out.
  • Time to conviction. Days from first look to a sized recommendation. Speed matters when it is real.
  • Coverage per analyst. Names properly researched, not names skimmed.
  • Losses avoided. Risks flagged early, with a note on what the old process would likely have missed.
  • Override quality. How often a human overruled the model, and how often that override was right.

Override quality is the one nobody expects. If your team overrides the model half the time and is wrong on most of those calls, your problem is not the tool. If they almost never override, the tool is being trusted without scrutiny, which is worse. A healthy range sits somewhere in between, and watching it move over a year tells you more about the partnership than any usage report ever will.

From One Pilot To Firmwide Muscle

A well measured pilot still dies if nobody owns it. Scaling is mostly a governance question, not a technology one.

Governance That Fits On A Single Page

  • A named owner for each use case, with the mandate to shut it down
  • Written kill criteria agreed before launch, not after the first bad quarter
  • A quarterly drift review comparing current output against samples from month one
  • A change log for every model, prompt library, and data feed update
  • A plain disclosure line for clients and consultants explaining where AI sits in the process

Kill criteria deserve special attention. Write down the exact conditions that end the project: override quality below a set level, no measurable time gain after two quarters, error patterns that compliance cannot explain. Teams that skip this step end up defending tools out of pride, and sunk cost quietly becomes strategy.

Client mix changes what you should measure, too. A desk running concentrated, illiquid, asset heavy portfolios cares about very different signals than a diversified equity book, which is why firms for property owners tend to build their screens around cash flow durability and holding period rather than quarterly earnings surprise.

Scale The Framework, Not The Prompts

When the first use case clears its bar, resist the urge to copy the prompts into a new domain. Copy the ledger, the baseline, the governance page, and the review cadence. Those transfer. The prompts almost never do, because the second decision has different data, different failure modes, and different people arguing about it.

Three or four measured use cases beat twelve unmeasured ones every time. And the same discipline pays off outside the portfolio, since the way a firm structures its published research now decides whether machines quote it at all, a shift worth understanding beyond backlinks and traditional search thinking.

The Scoreboard That Earns Its Keep

Across this series the same pattern kept showing up in different clothes. Firms start with the data they think they have and find out it was never decision ready. They hand the tool to their best thinker and watch that thinker's confidence bleed into every output. Then they grade the whole effort on how often people clicked. Three different traps, one shared root: measuring the easy thing instead of the true thing.

The firms that pull ahead are not the ones with the largest model budget. They are the ones that can point to a specific decision, show what the machine said, show what the human changed, and show what happened next. That record is slow to build and impossible to fake, and it is what turns a promising pilot into an edge that survives a bad quarter, a staff change, and a due diligence questionnaire. Start logging decisions this month. In a year you will have the only AI scorecard that ever mattered.

Argue with this in person

Barcamp Boston runs on hallway disagreements and half finished ideas. Pitch a session, grab a slot on the board, and take the conversation further than a comment box ever will.

See the schedule