One of the more surprising problems I encountered while building and operating an AI decision system was this:
The system could find the right evidence and still reach the wrong conclusion.
At first, that felt like a retrieval problem. Maybe the search needed to improve. Maybe the knowledge needed better organization. Maybe the system needed more context. Sometimes those things helped. But over time, a different problem became clearer:
Finding information and understanding information are not the same capability.
A system can retrieve a relevant case, metric, document, or example and still misunderstand what that evidence proves. It can find a projected benefit and treat it like a measured outcome. It can find independent work and interpret it as professional client experience. It can find experience in one context and assume it transfers directly to another. It can find no documentation for something and conclude that the experience never happened. The evidence may be relevant. The interpretation can still be wrong.
That changed the question for me. Instead of asking only: Can the system find the right information?
I began asking: Can the system understand what that information means well enough to use it responsibly?
Relevant Does Not Mean Sufficient
Search systems are very good at finding similarity. A question contains certain ideas, and the system finds documents containing related ideas. That is useful. But relevance does not tell us how much a piece of evidence actually proves.
Suppose someone is trying to determine whether a person has experience taking a product from strategy into implementation. A search might return a case containing a roadmap, prototypes, requirements, and a delivery plan. That case is clearly relevant.
But important questions remain:
- Was the roadmap approved?
- Did delivery begin?
- Did the person remain through implementation?
- Was the capability launched?
- Were outcomes measured?
Those are different levels of evidence. A source can be highly relevant while supporting only part of the conclusion someone wants to draw.
I had learned a version of this long before applying it to AI. A customer interview tells you what someone said in a particular context. Analytics tell you what happened, but not necessarily why. A stakeholder request tells you what someone wants, not automatically what should be built. A competitive feature shows what another organization offers, not whether you should copy it.
Research rarely arrives with its meaning attached.
Evidence becomes useful through interpretation. AI does not remove that requirement.
Different Evidence Answers Different Questions
Prudential provides a useful professional example. The work involved modernizing PruXpress, a portal used by financial professionals and insurance brokers, while the organization was also navigating a content-platform migration.
There was no single source of evidence that could define the future direction. We looked at analytics, heatmaps, help-desk data, competitive platforms, card-sorting results, stakeholder interviews, and workshops with leaders, managers, and brokers.
Each source answered a different kind of question:
- analytics could show where people were going;
- help-desk data could reveal recurring access problems;
- card sorting could help explain how people expected information to be organized;
- stakeholders could explain business needs and constraints;
- competitive research could establish what alternatives existed in the market.
None of those sources, by itself, was the strategy. The useful work came from understanding how they related. The resulting direction included capabilities around search, task-based navigation, personalization, entitlement visibility, and content governance, shaped by several forms of evidence rather than one dominant input. Following delivery, Prudential measured a 47% improvement in content findability and a 30% reduction in help-desk tickets related to content access.
Those outcomes are meaningful. They also show why evidence boundaries matter. They are downstream measured outcomes of the delivered platform. They are not the same thing as the research findings that informed the strategy, and they should not be written as though one person independently caused them.
Evidence has meaning partly because of where it sits in the story.
Provenance Changes What Evidence Can Prove
This became especially important in the AI decision system I built independently.
The system works across several kinds of evidence:
- professional client work;
- independent applied work;
- portfolio cases;
- professional learning;
- market observations;
- system observations.
They can all be useful. They are not interchangeable. This is what I mean by provenance: where the evidence came from and what kind of claim it can reasonably support.
- An independently developed AI governance model may demonstrate current strategic thinking. It does not become evidence of years spent running AI governance inside a client enterprise.
- A professional course may demonstrate learning. It does not establish implementation experience.
- A portfolio case may document strong strategic work. It does not prove ownership of every downstream result. Without those distinctions, AI can make evidence sound stronger simply by combining it.
Three weaker signals do not automatically become one strong fact. A conceptual model, a course, and an adjacent project might together suggest familiarity with a subject.They should not be combined to manufacture experience that never occurred. That is not only an AI problem.
People do it too. AI simply makes the synthesis faster.
Missing Evidence Is Not Negative Evidence
Another issue surfaced as I worked with larger amounts of structured professional information. Sometimes the system could not find evidence for a claim. It is tempting to treat that as the answer: No evidence found. Therefore the capability is absent. But those are not the same statement.
Documentation can be incomplete. A source can be unavailable. An older engagement may not yet be represented. The evidence may exist under different terminology. Or the experience may genuinely not exist. Those conditions require different conclusions.
This matters anywhere AI is expected to reason over incomplete information: a compliance review, a customer record, an internal investigation, a knowledge search, or a due-diligence process.“Not found” is not automatically the same as “did not happen.”
Sometimes the most accurate answer is simply: We do not have enough evidence to know.
That may feel unsatisfying because AI systems are often expected to produce conclusions. But preserving uncertainty is sometimes the more accurate response. Knowing when the evidence is incomplete can be more useful than confidently filling in the gap.
Terminology Can Hide Relevant Evidence
Evidence can also be missed even when it exists. Different organizations use different language for similar capabilities. One company may call something product strategy. Another calls it transformation. Another calls it strategic initiatives. One team may describe workflow redesign as operating-model work. Another may describe similar work as service transformation. The words matter, but they are not the capability itself.
If matching becomes too literal, strong evidence can be overlooked because the terminology differs. If interpretation becomes too flexible, almost anything can be translated into almost anything else. The goal is accurate translation.
Useful questions include:
- Does the earlier work actually involve the same underlying capability?
- Are the performance expectations comparable?
- Is the difference in context material?
- Does the evidence support the level of responsibility being claimed?
Those questions help distinguish a terminology gap from an experience gap. That has become especially important across Product, transformation, Experience, AI, and strategy, where overlapping work is often described differently depending on the organization. New domains sometimes introduce new vocabulary faster than they introduce entirely new kinds of enterprise work. But sometimes the context really is different.
Good interpretation has to recognize both.
More Knowledge Can Create a Harder Problem
There is an irony in knowledge systems. At first, the obvious problem is not having enough information. As the knowledge grows, the problem changes. Now there are more sources, more overlap, more context, more potential conflicts, and more material that is related but not equally useful. Retrieval may improve while decision quality does not. The system can find ten relevant things instead of two. Someone, or something, still has to decide which ones matter most.
That raises a different set of questions:
- Is the source current?
- Is it authoritative for this question?
- Is it direct evidence or interpretation?
- Does another source carry more weight?
- Is the outcome measured, estimated, targeted, or observed?
The constraint shifts from: Do we have the information? to: Can we make sense of the information we have?
Evidence Strength Depends on the Question
There is no single universal ranking of evidence. The strongest evidence depends partly on what you are trying to determine.
Professional implementation experience may be stronger than independent work if the question is whether someone has delivered inside a client enterprise. Independent applied work may be more relevant if the question concerns a newer AI problem explored through current hands-on work.
A measured outcome is powerful when discussing results. It may tell you much less about how the original strategy was developed. That leaves two useful questions:
- How strong is the evidence?
- How relevant is it to this question?
Those are not always the same thing. Good decision support has to consider both.
Good Evidence Systems Should Be Able to Say “Not Enough”
There is a natural pressure in AI systems to complete the pattern. A question gets asked. Relevant information is found. The system is expected to generate an answer. But the evidence does not always support one.
Sometimes the strongest conclusion is:
- the evidence supports part of the claim;
- the evidence is transferable but indirect;
- the source is relevant, but implementation is not established;
- the result was projected, not measured;
- the documentation is incomplete;
- more evidence is needed.
Those are not failures. They are distinctions that protect decision quality.
As AI systems influence more important work, that discipline becomes increasingly valuable. A confident answer built on weak interpretation is not necessarily better than a cautious answer grounded in stronger evidence.
The aim should be the strongest conclusion the evidence can honestly support.
Retrieval Is the Beginning of Reasoning
Enterprise AI has made retrieval a major technical focus. That makes sense. AI systems become more useful when they can work from relevant organizational information rather than relying only on general knowledge. But finding the right document does not solve the whole problem.
The evidence still has to be interpreted:
- What kind of source is this?
- How direct is it?
- How current is it?
- What context produced it?
- What does it prove?
- What does it not prove?
- Where is the uncertainty?
Those questions are where retrieval becomes reasoning. The more important the decision, the less comfortable I am with the idea that relevance alone is enough. A useful evidence system should not simply make information easier to find.
It should make the boundaries of that information easier to understand.
Because the real value of evidence is not that we can retrieve it. It is that we can interpret it well enough to make a better decision.
Related Work
Defining an Advisor Productivity Platform Under Migration Constraint
The professional case behind the article’s main example, showing how analytics, research, stakeholder evidence, platform constraints, and measured downstream outcomes played different roles in shaping and evaluating the work.
Building a Governed Intelligence Operating System
The independent applied system in which evidence provenance, source authority, interpretation, uncertainty, and claim boundaries became practical design problems rather than abstract principles.
Creating Value with AI Takes More Than a Model
The broader system-level companion to this article, focused on why reliable AI-assisted work depends on a trustworthy information foundation as well as the model itself.
