---
title: "Verifiable AI Answers From Company Documents"
metaTitle: "Verifiable AI Answers From Documents"
description: "How to build document answers with verifiable sources, access controls and a clear response when the required information is missing."
date: 2026-09-23
tags: [agentic-ai, knowledge-base]
cover: cover.webp
coverAlt: "Illustrated woman and colleague cross-checking a document against a screen in an office archive"
draft: false
updated: 2026-09-23
---

When an employee asks an AI agent about a policy and gets a confident answer with no reference to any file, that is not an assistant - it is a plausible-text generator. A verifiable answer works differently: it names a specific document, section, and version date; it separates a direct quote from a conclusion; and it says so when no source was found. That is the kind of system worth building.

## Why the model has no idea what is in your folder

A language model is trained on a large corpus of text, but your internal policies were never part of it. The model cannot see your document folder in real time, and it does not update automatically when you change a policy.

One approach is called RAG - Retrieval-Augmented Generation. Before the model composes an answer, the system searches an index of your documents for relevant passages and passes them to the model together with the question. In a production setup, the model is constrained to those passages and the answer is checked to confirm it actually rests on them.

The important caveat: RAG alone does not guarantee correct citations or a clear chain of evidence. That requires separate design work on retrieval, citation formatting, and passage-grounding checks. If a document contains an error, the model may reproduce it. A citation makes an answer checkable - it does not prove the answer is right.

## Three things to settle before you go live

Before you switch on search over a knowledge base, three questions need answers.

The first is version currency. A folder often holds several files with similar names. If the system cannot tell which one is current, it may blend an old policy with a new one, and the employee gets a contradictory answer. For every document, record: effective date, status (current / archived / draft), and owner.

The second is document ownership. This is not a formality. When the agent finds no answer, or finds conflicting versions, it needs to know who to direct the question to. Without a named owner, the chain breaks.

The third is access rights. Not every employee should see every document. This is resolved before indexing, not after.

## Access rights: the index must not bypass restrictions

A hidden problem comes up here often. Suppose a sales manager has no access to HR documents. If the agent builds a single index across all files and then simply suppresses results at display time, that is not reliable protection. A fragment from a restricted document can leak into an answer indirectly.

A sound architecture: the agent searches only the documents the requesting user is permitted to see. The index is either segmented by role, or filtering is applied at the retrieval stage before any passages reach the model. Testing this is straightforward - log in with a restricted account and ask about a restricted document directly. The agent must neither answer from it nor hint at its contents.

Write access and the ability to trigger actions are a separate matter. The right to read a document and the right to initiate an action based on it are not the same permission level.

For a deeper look at structuring access controls in a knowledge base, see the article on [company AI permissions](/blog/company-brain-permissions/).

## A worked example: two versions of the same policy

Consider a fictional scenario. An employee asks: "What is the current policy for handling a customer request?"

The folder contains two files:
- `request-policy-v1.pdf` - an older file with no status marked
- `request-policy-v2.pdf` - the current version, with a later effective date

If the system does not distinguish their statuses, it may treat both as relevant and either blend their content or answer from the older one. The problem is not that the old file is "indexed better" - it is the absence of an explicit status and effective date in the metadata.

What the system should do in this case:

1. At indexing time, check metadata: status and effective date.
2. Either exclude the archived document from search, or flag it so the agent will not use it as a primary source.
3. In the answer, state explicitly: "Source - `request-policy-v2.pdf`, section 3.2, effective [date from metadata]."

## What a good answer looks like, and what a bad one looks like

For illustration, assume the current fictional policy contains the line: "For a non-standard request, the employee escalates the matter to the department head." An answer might then read:

> Source: "Request Handling Policy," version 2, section 3.2. Quote: "For a non-standard request, the employee escalates the matter to the department head." Conclusion: this request should be escalated to the department head. The processing deadline is not mentioned in the cited passage; check the full document for that detail.

The source and version are named, the quote is separated from the conclusion, and the missing deadline is not invented. This is an illustration using a fictional document - it is not an excerpt from any Majento or client policy.

A bad answer looks like this:

> Requests are processed within two business days. Contact the service department to proceed.

No file, no version, no section. It is impossible to tell whether this reflects a current policy or the model's general knowledge. There is nothing to verify and nothing to challenge.

## When no source is found, that is also an answer

An important scenario: the agent finds no relevant document. A weak system responds by generating plausible text from general knowledge. A good one says clearly: "No documents matching your question were found in the knowledge base. I recommend contacting the owner of this area or checking whether the relevant policy has been uploaded."

This is not a system failure. The employee knows what to do next, rather than walking away with an answer nobody can verify.

For guidance on building an AI rollout that accounts for scenarios like this, see the [AI implementation section](/ai-implementation/).

## Acceptance criteria for your system

Before handing the agent over for real use, check four things:

- Every answer that draws on a found document includes the file name and the precise location of the passage, where the source structure allows it.
- The quote and the agent's conclusion are visually or textually separated.
- A request from a user without access to a document returns none of its content - neither directly nor indirectly.
- For a question where no documents exist, the agent says so and names a next step.

If any one of these fails, the system is not ready for production.

## FAQ

### Can you fully trust an answer if the agent cited a source?

No. The agent reproduces a passage from a document, but it may pick the wrong passage, interpret it imprecisely, or miss important context from an adjacent section. A citation is an invitation to verify - not a guarantee of correctness. Check critical answers manually.

### How often should the index be updated?

It depends on how frequently documents change. Index updates should be tied to document changes and publication events, not left to an infrequent scheduled run. Otherwise the system may answer from an outdated version even when the date in the response looks convincing.

### What about documents that are still being drafted or awaiting approval?

Drafts and documents under review are best excluded from the index entirely, or kept in a separate space with explicitly restricted access. The agent must not treat a draft as a current policy.

### Does each department need its own agent?

Not necessarily. Properly segmented access rights are usually enough: one agent can serve the whole company, with each user seeing only what they are permitted to see. Separate agents make sense when departments have fundamentally different document workflows.

If you want to work through a specific setup for your knowledge base, reach out on Telegram at [@shimaoz](https://t.me/shimaoz) or by email at [hello@majento.ai](mailto:hello@majento.ai).

## Source

Majento consultations and training courses.
