Skip to content
GenAIDocument Intelligence

Enterprise Knowledge Assistants: Getting Answers With Sources From Your Documents

INFIQON Insights7 min read

Almost every organisation that has tried generative AI has built the same prototype: point a language model at a folder of documents, ask it questions, and watch it answer fluently. The demo takes days. Getting from that demo to an assistant that support engineers, analysts or relationship managers rely on every day takes considerably longer, and the work that fills the gap is mostly not about the model.

The reason is simple. An answer is only useful at work if the person reading it can decide whether to act on it. That requires knowing where the answer came from, whether the source is current, and whether the person asking was allowed to see it in the first place. An enterprise knowledge assistant is, in practice, a system for answering those three questions reliably — with a language model doing the summarising at the end.

Why the citation is the product

The architecture most teams use is retrieval-augmented generation: find the passages most relevant to a question, hand them to the model, and ask it to answer from those passages only. Done properly, every sentence of the answer can be traced back to a specific document, section and version. That traceability is not a nice-to-have feature layered on top. It is what turns a plausible paragraph into something an engineer can verify in thirty seconds before changing a customer's configuration.

Assistants without sources tend to follow a predictable arc. Early usage is enthusiastic, someone acts on a confident answer that turns out to be outdated or invented, the story travels, and usage quietly collapses back to searching five systems and messaging a senior colleague. With sources shown by default, the same wrong answer is caught at the point of use — the reader clicks through, sees the document is from 2021, and goes elsewhere. The system is no less fallible, but the failure is visible and cheap.

  • Cite at passage level, not document level. A link to a 200-page manual is not a source; a link to the paragraph is.
  • Show the document's owner and last-updated date alongside the citation, so staleness is obvious without opening it.
  • Instruct the model to say it does not know when retrieval returns nothing relevant — and test that it actually does.
  • Distinguish quoting from synthesising. When the answer combines three sources, say so.

The content problem nobody budgets for

Retrieval can only return what exists, and most enterprise knowledge bases contain several conflicting versions of the truth. The 2019 onboarding guide, the 2023 revision and the Slack thread that superseded both will all be retrieved for the same question, and the model will do its best to reconcile them. It usually shouldn't have to.

The most valuable early work is therefore unglamorous: inventory the sources, decide which are authoritative for which topics, assign an owner to each, and set freshness rules — what gets archived, what gets re-reviewed, and how often. This does not need to be a year-long content programme. A few weeks spent on the sources that answer the most frequent questions covers the majority of real traffic, and the assistant itself will tell you where to go next.

That last point is underrated. Once the assistant is live, every question it answers badly is a signal: either the content is missing, the content exists but is wrong, or retrieval failed to find it. Logging questions against the documents used to answer them gives enablement and documentation teams something they have rarely had — evidence of which content earns its keep and which gaps cost people time every day.

Permissions are a hard requirement, not a phase two

The fastest way to end a knowledge-assistant programme is for it to surface a salary band, a customer's contract terms or a draft board paper to someone who should never have seen it. Language models do not respect access controls; retrieval has to. The assistant must only retrieve passages the asking user could already open in the source system, evaluated at query time rather than copied once at indexing time.

  • Carry source-system permissions into the index as metadata, and filter on them before anything reaches the model.
  • Re-sync permissions frequently. A person who changed teams last week should lose access this week, not at the next full re-index.
  • Treat customer-specific material as its own boundary, so content from one account can never inform an answer about another.
  • Log what was retrieved for each answer, not just what was generated. It is the audit trail you will want the first time someone asks.

In a risk-tiered governance model, a well-built internal assistant that cites sources and respects permissions usually sits in a low tier: its output is a recommendation that a person reads, checks and decides on. Skipping permission-aware retrieval moves it up several tiers at once.

Retrieval quality decides answer quality

When an assistant gives a poor answer, the instinct is to blame the model or reach for a larger one. In most cases we examine, the model was handed the wrong passages. Improving retrieval is where the accuracy actually comes from, and the techniques are well understood.

  • Chunk documents along their structure — sections, headings, table rows — rather than fixed character counts that split a procedure in half.
  • Combine semantic search with keyword search. Product codes, error numbers and clause references are exactly what embeddings handle worst and people ask about most.
  • Add a reranking step, which reorders a broad candidate set by relevance before the model sees it and is often the single largest quality gain.
  • Prefer recent and authoritative sources explicitly, using the ownership and freshness metadata from the content work.

Measure answers, not usage

Weekly active users is a poor measure of whether an assistant works; people will keep trying a tool for a while even when it disappoints them. A more honest scorecard combines a fixed evaluation set with production signals. Build a few hundred real questions with known correct answers and sources, drawn from tickets and internal channels, and run them on every change to content, retrieval or model. Then watch, in production, the share of answers with a source clicked through, the thumbs-down rate by topic, the rate of 'I don't know' responses, and — most importantly — the downstream operational number the assistant was meant to move.

For an enterprise SaaS company we worked with, that number was internal support load. Fifteen years of product knowledge sat across documentation, resolved tickets, release notes and chat threads. A permission-aware assistant over more than 150,000 documents, embedded in the support console and in Slack, with every answer cited, brought internal support requests down 38% and shortened new-hire ramp-up by 40%. The gains came as much from the content clean-up and the feedback loop as from the retrieval stack itself.

Put it where people already work

A separate assistant portal is one more tab people forget to open. Adoption follows placement: inside the support console next to the ticket, in the chat tool where questions are already being asked, in the CRM beside the account. The same retrieval layer can serve all of them, and it can also power agent-facing copilots that draft a reply with the policy clause and order history attached — the pattern behind the 65% faster first response we saw with a retail group's support team.

A sensible first ninety days

  • Weeks 1–3: pick one team with a measurable pain — repeated questions, slow onboarding, escalations to experts — and collect a few hundred real questions they ask.
  • Weeks 3–6: clean and assign owners to the sources that answer those questions; build retrieval with permissions and passage-level citation; stand up the evaluation set.
  • Weeks 6–9: release to a pilot group inside their existing tools, with feedback captured on every answer.
  • Weeks 9–12: fix the top content and retrieval failures, compare the operational metric against baseline, and decide on the next team from evidence rather than enthusiasm.

The organisations that get lasting value from knowledge assistants treat them as a knowledge-management programme with a very good interface, not as a model deployment. If you want to know how ready your own content, permissions and data estate are before you start, our AI readiness assessment is a practical place to begin.

Talk to the people behind the writing

If this resonates, the conversation is better — book a session and bring your hardest questions.