Back to all articles
6 min read

Your Best Analyst May Be Teaching Your AI Bad Habits

Senior judgment can quietly poison AI research output. Here is how to set roles, prompts and review loops so your best people sharpen the model.

The Expertise Paradox On A Research Desk

Part two of three. The problem is rarely that analysts refuse to use AI. The problem is that your sharpest one knows exactly how to make it agree with them, and nobody on the desk can tell the difference.

Part one made the case that pilots stall in the plumbing, and that the "clean data" delusion that quietly is what quietly sinks a program before a single insight arrives. Assume you fixed that. Assume the data flows, the workflow is mapped, and the tool is live. There is a second failure mode waiting, and it does not look like failure at all. It looks like productivity.

Here is the pattern. Your most experienced analyst has spent fifteen years building pattern recognition. She reads a filing and knows within a page where the risk sits. When she starts prompting a model, she does not ask an open question. She describes the answer and asks the model to support it. The output comes back fast, well written, and aligned with her view. She approves it. It goes into the note. It goes into the committee deck. Nobody argues, because the analysis now carries two signatures instead of one, and one of them looks objective.

“The most dangerous prompt on a research desk is the one that already contains the answer,” says Marcus Devlin, head of quantitative research at Zenith Investment Management. “You are not testing a thesis anymore. You are hiring a very fluent assistant to agree with you, and then reporting the agreement as evidence.”

Three Ways Expertise Leaks Into The Output

  • Loaded framing. The prompt names the conclusion, the tone, or the comparison set. The model fills in the gaps politely.
  • Silent editing. The analyst deletes the parts of the output that contradict the thesis before anyone else reads it, so the disagreement never gets logged.
  • Selective re-rolling. Run the query six times, keep the version that fits. This is not research. It is shopping.

None of this is dishonest. It is what expertise feels like from the inside. The fix is not more training on prompt tricks. The fix is structural, and it starts with pulling two jobs apart that most desks treat as one.

Splitting The Question From The Operator

On most desks, one person defines the question, runs the model, reads the output, and writes the recommendation. That is four jobs in one seat, and it removes every point where an error could get caught, which is why the most capable analyst often does the most damage. Break the chain in two places and the same person becomes far more useful.

The Question Setter And The Operator

The question setter decides what needs to be known and writes it as a neutral question with a testable answer. Not "show me why this name is mispriced" but "what evidence exists on both sides of the pricing case for this name over the last eight quarters." The operator runs it, unchanged, and returns the raw output before anyone edits it. The senior analyst can hold either role. She just cannot hold both on the same question. That single rule kills most of the contamination, because a leading prompt looks obviously leading the moment a second person has to read it out loud.

Why Juniors Often Run The Model Better

Analysts two years into their career tend to produce better AI-assisted first drafts, and the reason is uncomfortable. They have less to defend. They ask the model what it actually found, sit with an answer that surprises them, and bring it upstairs intact. Firms that lean into this pattern, including several of the arizona wealth management picks now publishing their process openly, tend to route the raw generation work down and the challenge work up, rather than the reverse.

Role rules that survive a deadline:

  • The prompt is a document, not a keystroke. It gets saved with the output, every time.
  • Whoever writes the thesis does not run the query that tests it.
  • Raw output goes into the file before any human edit. Edits are tracked separately.
  • Every model run ends with a required question: what would have to be true for this to be wrong.
  • Senior review happens on the raw output plus the edit trail, not on the polished note.

Portfolio construction needs the same discipline, applied to weights instead of words. When a model proposes an allocation and the portfolio manager overrides it, that override should cost something visible. Give the desk an override budget: a fixed number of manual adjustments per rebalance, each with a written reason. Overrides stop being casual and start being arguments. The reason writing matters more than people expect, because notes that are structured and specific stay retrievable later, and much of the thinking in beyond backlinks applies just as well inside your own research archive as it does on the open web.

Review Loops That Catch Contamination Early

Roles set the boundaries. Review loops are what tell you whether the boundaries held. Three loops cover almost everything, and none of them require new software.

The Pre-Commitment Note

Before the model runs, the analyst writes two or three sentences stating what she expects to find and how confident she is. Timestamped, thirty seconds of work. Now the output can be compared against a fixed prior instead of a memory that quietly reshapes itself. When the model disagrees and the analyst still wins the argument, that is judgment doing its job. When the analyst's expectation and the final note match perfectly every single time, you have a problem, and the pre-commitment note is the only artifact that shows it.

The Disagreement Log

Keep a running record of every case where the model and the human landed in different places, what was decided, and what happened afterward. Read it quarterly. Most desks discover two things fast. First, disagreements cluster around a handful of sectors or data types, which usually means the underlying inputs are weaker there than anyone admitted, a problem that loops right back to the data readiness groundwork that should have happened before launch. Second, one or two people win nearly every disagreement, and their win rate has nothing to do with being right.

Calibration Instead Of Confidence

Score the humans and the model on the same scale. Not on tone, not on thoroughness, on whether the stated confidence matched the eventual outcome. An analyst who says she is 70 percent sure should be right about 70 percent of the time. Track it for a year and you learn something no performance review will tell you: which people sharpen the model, which people mostly reformat it, and which people are quietly overriding it into worse decisions while sounding excellent in meetings.

All three loops produce something valuable beyond quality control. They produce a paper trail that connects a specific decision to a specific process, which is exactly what you need when the question shifts from "are we using AI" to whether any of it reached the return stream. That question is where most programs go soft, and it is the subject of the metric most firms track in part three.

The Desk That Argues Better Wins

The uncomfortable truth in AI-assisted investment management is that experience cuts both ways. The same pattern recognition that makes a senior analyst valuable also makes her the most efficient possible source of confirmation bias, because she can shape a prompt, read the output, and file the result without ever noticing that she taught the machine her answer. Skill does not protect you here. Structure does. Separate the question from the operator, make prompts and raw outputs part of the permanent record, put a price on overrides, and keep a log of every disagreement instead of resolving it in a hallway.

Do that and something shifts. The model stops being a mirror and starts being a counterparty, and the best analyst on the team goes back to doing what she was actually hired for: judging a genuinely independent second opinion. That is the whole point of the pairing. Not faster notes, not more coverage, but a desk where two different kinds of reasoning collide often enough to catch each other's mistakes, and where you can prove afterward which one was right.

Argue with this in person

Barcamp Boston runs on hallway disagreements and half finished ideas. Pitch a session, grab a slot on the board, and take the conversation further than a comment box ever will.

See the schedule