Thinklytics

Knowledge Retrieval · 10 min read · October 2026

Why nobody can find what your company already knows

By Sean Majidi, Founder, Thinklytics

Asking a colleague wins because it solves four problems at once: the content is unstructured, nothing indexes it usefully, nobody knows which copy is current, and the asker may not be cleared to see it. Any fix that solves one loses to the colleague. Also, the statistic everyone quotes for this is from 2012.

Somebody asks where the current version of the renewal process is. Three people answer with three links. One is from 2024.

The usual reading is that the search is bad. The more useful reading is that asking a colleague solved four problems at once, and search solves one of them.

The four reasons

Four reasons the answer cannot be found, and what each one needs

Asking a colleague wins because it solves all four at once. Any fix that solves only one loses to the colleague.

ReasonWhat it looks likeWhat it actually needs
The content is unstructuredThe answer is a sentence inside a PDF, a thread, or a recorded callExtraction into records with fields, not a better search box
Nothing indexes it usefullySearch matches the words and misses the meaning, or matches a superseded versionRetrieval that cites what it used, so the reader can check which version they got
Nobody knows which copy is currentFour near-duplicates, two contradicting, none datedAn owner per document set, which is a governance decision rather than a tool
The person asking may not be allowed to see itThe honest reason a lot of content was never indexed at allPermissions resolved per user at query time, which is the hard part

A colleague answers correctly, with context, from the current version, and only tells you what you are cleared to hear. That is the bar, and it explains why most internal search projects lose to a Slack message.

Source: Thinklytics engagement pattern across the document processing and knowledge retrieval engagements in the case library.

The content is unstructured. The answer is a sentence inside a PDF, a thread, or a recorded call. No amount of ranking turns a paragraph into a field you can filter on.

Nothing indexes it usefully. Search matches the words and misses the meaning, or matches a version that was superseded two quarters ago and looks identical.

Nobody knows which copy is current. Four near-duplicates, two of them contradicting, none dated and none owned.

The person asking may not be allowed to see it. This is the honest reason a lot of content was never indexed in the first place, and it is the one that gets left out of the business case.

A colleague answers correctly, with context, from the version they know is current, and only tells you what you are cleared to hear. That is the bar. It explains why an internal search project can be technically successful and still lose to a Slack message.

The statistic worth dating

The statistic everyone quotes for this problem

  • What gets cited. 1.8 hrs/day. Or 9.3 hours a week, or 20% of the working week, depending on the deck. Presented as a current finding about knowledge work.
  • What it actually is. 2012. A McKinsey Global Institute figure about interaction workers, from 2012. The hourly versions are other people's arithmetic on that 20%, not McKinsey numbers.

Worth dating deliberately rather than dropping. This is not a problem AI created an opportunity to solve, it is a problem organisations have failed to solve for fourteen years, which should set expectations for how much a new tool changes on its own.

Source: McKinsey Global Institute, 2012, as traced in our own piece on internal knowledge assistants.

Before the business case gets written, one number to handle carefully.

The figure everyone cites comes from McKinsey Global Institute and it is from 2012: interaction workers spend nearly 20% of the working week looking for internal information or tracking down colleagues who can help. The versions circulating as 1.8 hours a day or 9.3 hours a week are someone else's arithmetic on that 20%, not McKinsey figures.

We date it deliberately rather than dropping it. This is not a problem that AI created an opportunity to solve. It is a problem organisations have been failing to solve for fourteen years, and that should inform how confident anyone is that a new tool fixes it. Our own earlier analysis traces the lineage in full, along with a benchmark worth reading before any purchase: internal knowledge assistants, what breaks.

That benchmark is the uncomfortable part for the obvious fix. Over roughly 500,000 documents across nine real enterprise platforms, seeded with misfiled documents, near-duplicates and conflicting information, BM25 keyword ranking beat vector search by more than 17 points on correctness. The semantic-search pitch does not survive contact with a real corpus, which means the second of the four reasons is not even a settled technology question.

Why the extraction number matters more than the model

What it looks like when finding the answer gets fixed

Published outcomes. In each case the work was structuring the content and settling who owned it, not installing a search tool.

EstateBeforeAfter
Adverse event intake, pharmaceutical18 days to process a case, 6 unconnected sources4 hours for standard cases, 6 sources unified, full on-time FDA submission
Statutory information requests, federal agency34 days per request across 6 regional offices8 days, 2,400 requests a year handled, $1.2M of annual penalties eliminated
Invoice and document intake, benchmarkedVendors quote field-level accuracy in the mid 90sWhole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%
Public health reporting, county14 manual Excel reports assembled by handDashboards used by 180 staff, $620K a year of reporting labour removed

The third row is the one to carry into a vendor conversation. A per-field accuracy figure and a per-document accuracy figure are different numbers, and only the second one decides whether a person still has to read everything.

Source: Thinklytics case library, published delivery outcomes; whole-document accuracy benchmark as cited in our document processing cost analysis.

If the answer has to come out of documents, the accuracy figure you are quoted and the accuracy figure that decides your case are different numbers.

Vendors quote field-level accuracy, which lands in the mid 90s and sounds like a solved problem. Whole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%. The gap is not vendor dishonesty, it is a different measurement: getting 19 of 20 fields right on a document means getting the document wrong.

That distinction decides whether a person still has to read everything that comes through, which is the entire economics of the project. The detail is in what AI document processing costs, and the design consequence, which is that the escalation path matters more than the accuracy, is in document processing: design the escalation path.

What it looks like when it gets fixed

A pharmaceutical client was processing adverse events across six unconnected sources, taking 18 days per case. Unifying the sources and automating the pipeline took that to four hours for standard cases and six hours for serious ones, with full on-time FDA submission compliance and five of eight pharmacovigilance specialists' time freed. See the pharmacovigilance pipeline engagement.

A federal agency across six regional offices had the same problem in its statutory form: requests arriving for information that existed somewhere, with a legal obligation to release what was releasable and withhold what was not. Response time went from 34 days to 8 across 2,400 requests a year, eliminating $1.2M of annual penalties. See the federal agency governance engagement.

In both cases the work was structuring the content and settling who owned it. Neither was a search installation.

The reason this gets deferred

The fourth reason is the expensive one, and it is also the one nobody wants to raise in the kickoff.

An assistant that answers from a document the asker could not open has not leaked a file. It has leaked the contents, which is worse, because nothing in a file access log shows that it happened. There is no event to find afterwards.

Resolving permissions per user against the live source at query time is the only model that stays correct. It is also slower and more expensive than copying the permission graph into the index at build time, so the cheaper option tends to get chosen by default rather than by decision. The difference does not show up in a demo. It shows up the first time somebody's access is revoked and the assistant keeps answering.

That decision is the subject of retrieval that respects who is asking.

What we would do first

Take the ten questions your team actually asks most often. For each one, write down where the answer lives, how many versions exist, who owns the current one, and who is allowed to see it.

Four columns, ten rows, one afternoon.

Most teams find the blocker is the fourth column. That is a governance decision rather than a procurement one, and discovering it before shortlisting software rather than during an implementation is worth a quarter.

Delivery sits in RAG consulting for retrieval that cites what it used, data governance consulting for the ownership that makes currency answerable, and AI security consulting for the access model and the adversarial testing behind it. If the goal is being found by answer engines rather than by colleagues, that is a different problem and it sits in answer engine optimization. The full set of work in this area sits under nobody can find what we already know.

Frequently asked questions

Why can't anyone find information that already exists internally?

Four things are wrong at once. The content is unstructured, so the answer is a sentence inside a PDF or a thread rather than a field in a record. Nothing indexes it usefully, so search matches words and misses meaning or matches a superseded version. Nobody knows which copy is current, because there are four near-duplicates and none is dated or owned. And the person asking may not be cleared to see it, which is the honest reason a lot of content was never indexed at all. Asking a colleague solves all four at once, which is why it wins.

Is it true that employees spend 1.8 hours a day searching for information?

That figure is other people's arithmetic on a McKinsey Global Institute finding from 2012 that interaction workers spend nearly 20% of the working week looking for internal information or tracking down colleagues. The hourly and weekly versions are not McKinsey numbers. Worth dating deliberately rather than dropping: this is not a problem AI created an opportunity to solve, it is one organisations have failed to solve for fourteen years, which should set expectations for how much a new tool changes by itself.

Will better search fix it?

It addresses the second of the four reasons and leaves the other three. A search box cannot tell you which of two permitted copies is current, cannot turn a sentence in a recorded call into a field you can filter, and cannot decide whether you are allowed to see the answer. In a 2026 benchmark over roughly 500,000 documents across nine real enterprise platforms, BM25 keyword ranking beat vector search by more than 17 points on correctness, so even the search half is not a settled technology question.

What does fixing it actually involve?

Extraction into structured records where the answer is a fact rather than a paragraph, retrieval that cites what it used so a reader can check which version they got, an owner per document set so currency has an answer, and permissions resolved per user at query time. The first three are data and governance work. The fourth is the one that decides whether the result can be approved.

What is the hard part?

Permission-aware retrieval. An assistant that answers from a document the asker could not open has not leaked a file, it has leaked the contents, and nothing in a file access log records that it happened. Resolving permissions per user against the live source at query time is the only model that stays correct, and it is more expensive and slower than copying the permission graph into the index, which is why the cheaper option gets chosen by default.

What results have you measured on this?

A pharmaceutical client cut adverse event processing from 18 days to 4 hours for standard cases by unifying six unconnected sources, with full on-time FDA submission compliance. A federal agency across six regional offices took statutory information requests from 34 days to 8, handling 2,400 a year and eliminating $1.2M of annual penalties. A county health department replaced 14 manually assembled Excel reports with dashboards used by 180 staff, removing $620K a year of reporting labour.

How accurate is document extraction in practice?

Lower than the number you will be quoted, and the difference is the one that matters. Vendors quote field-level accuracy in the mid 90s. Whole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%. A per-field figure and a per-document figure are different measurements, and only the second decides whether a person still has to read everything that comes through.

Where should we start?

Take the ten questions your team actually asks most often and write down, for each, where the answer lives, how many versions exist, who owns the current one, and who is allowed to see it. Four columns, ten rows, one afternoon. Most teams find the blocker is the fourth column, which is a governance decision rather than a procurement one, and finding that before shortlisting software saves a quarter.

The work behind this

Six engagements in the case library carry document processing, and the retrieval work sits alongside the governance engagements that settle who owns what. The published outcomes include adverse event intake from 18 days to 4 hours and statutory information requests from 34 days to 8.

Document processing, 6 engagements.

Topics covered

  • enterprise search
  • knowledge retrieval
  • internal documents AI
  • find company information
  • unstructured content
  • document search
  • knowledge management

Frequently asked questions

Why can't anyone find information that already exists internally?

Four things are wrong at once. The content is unstructured, so the answer is a sentence inside a PDF or a thread rather than a field in a record. Nothing indexes it usefully, so search matches words and misses meaning or matches a superseded version. Nobody knows which copy is current, because there are four near-duplicates and none is dated or owned. And the person asking may not be cleared to see it, which is the honest reason a lot of content was never indexed at all. Asking a colleague solves all four at once, which is why it wins.

Is it true that employees spend 1.8 hours a day searching for information?

That figure is other people's arithmetic on a McKinsey Global Institute finding from 2012 that interaction workers spend nearly 20% of the working week looking for internal information or tracking down colleagues. The hourly and weekly versions are not McKinsey numbers. Worth dating deliberately rather than dropping: this is not a problem AI created an opportunity to solve, it is one organisations have failed to solve for fourteen years, which should set expectations for how much a new tool changes by itself.

Will better search fix it?

It addresses the second of the four reasons and leaves the other three. A search box cannot tell you which of two permitted copies is current, cannot turn a sentence in a recorded call into a field you can filter, and cannot decide whether you are allowed to see the answer. In a 2026 benchmark over roughly 500,000 documents across nine real enterprise platforms, BM25 keyword ranking beat vector search by more than 17 points on correctness, so even the search half is not a settled technology question.

What does fixing it actually involve?

Extraction into structured records where the answer is a fact rather than a paragraph, retrieval that cites what it used so a reader can check which version they got, an owner per document set so currency has an answer, and permissions resolved per user at query time. The first three are data and governance work. The fourth is the one that decides whether the result can be approved.

What is the hard part?

Permission-aware retrieval. An assistant that answers from a document the asker could not open has not leaked a file, it has leaked the contents, and nothing in a file access log records that it happened. Resolving permissions per user against the live source at query time is the only model that stays correct, and it is more expensive and slower than copying the permission graph into the index, which is why the cheaper option gets chosen by default.

What results have you measured on this?

A pharmaceutical client cut adverse event processing from 18 days to 4 hours for standard cases by unifying six unconnected sources, with full on-time FDA submission compliance. A federal agency across six regional offices took statutory information requests from 34 days to 8, handling 2,400 a year and eliminating $1.2M of annual penalties. A county health department replaced 14 manually assembled Excel reports with dashboards used by 180 staff, removing $620K a year of reporting labour.

How accurate is document extraction in practice?

Lower than the number you will be quoted, and the difference is the one that matters. Vendors quote field-level accuracy in the mid 90s. Whole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%. A per-field figure and a per-document figure are different measurements, and only the second decides whether a person still has to read everything that comes through.

Where should we start?

Take the ten questions your team actually asks most often and write down, for each, where the answer lives, how many versions exist, who owns the current one, and who is allowed to see it. Four columns, ten rows, one afternoon. Most teams find the blocker is the fourth column, which is a governance decision rather than a procurement one, and finding that before shortlisting software saves a quarter.

Related reading

If this is the problem you have