Knowledge sources and citations
Before an AI Agent Searches Internal Files, Decide What It May Cite
A knowledge base is not a folder dumped into a model. Owners, versions, expiry, and citations keep an Agent from answering confidently from stale rules.
“Let the AI read our SOP and product notes” sounds like the obvious first step. The files exist, people search slowly, and an Agent should help. The obstacle is that those files are rarely in a citable state: a three-year-old Word doc, a deck on someone’s desktop, a screenshot from a group chat, a price list that disagrees with the website. People hesitate. A model does not. It completes the sentence. For a knowledge Agent, the main work is often deciding which sources may be cited and who unpublishes them when they expire—not the embedding technique itself.
Split what may be said externally from what is only internal context
Approved service descriptions, public processes, and standard FAQs can be citation sources. Drafts, unpublished pricing strategy, and one-off customer exceptions should not become customer-facing text.
Every citable document needs an owner. Unowned files become “probably still right,” then get copied into more conversations by the Agent.
A typical conflict is a seven-day website policy, a warehouse SOP that refuses opened goods, and a verbal exception in a staff group. People ask. An Agent may blend all three into a sentence nobody approved.
- Citable: approved, dated, owned documents and system fields
- Internal only: drafts, negotiation room, unpublished cost, one-off promises
- Excluded: unsourced screenshots, expired forms, personal notes
Citations have to be visible or reviewers will not release the answer
A fluent paragraph is hard to check. Better output is the answer, the document, its update date, and anything that could not be sourced.
Live facts should not be read from PDFs. Shipment status, remaining sessions, or full bookings belong in an authorized system query. Documents hold stable rules; systems hold current state.
AgentTech treats retrieval as a tool the Agent calls, not as unbounded context. Policy search hits the approved set; order search hits the system. Sources can then be replaced without rebuilding the whole workflow.
The update path matters more than the first import
The first cleanup can be a project. Next month’s price change will not. If refresh depends on someone remembering to re-upload, the Agent will use old rules at the worst time.
Keep the effective rule in one place and sync it. Who changes the official content should trigger an update and a small sample check.
If a stale citation appears, unpublish the file before tuning the prompt. Otherwise the model will work harder to make the old text sound right.
A knowledge Agent can draft without speaking for the company
Commitments still need a person even when the source is clean. The Agent’s job is to find the relevant passage and mark gaps so a decision takes less time.
This is different from low-risk auto-replies that quote fully approved text. They may share sources; they should not automatically share send authority.
The first collection can be small if it matches real questions
Do not start by vectorizing the entire drive. Take the last month’s real questions that already have official answers, make those documents citable, and search only that set.
Test with customer wording, not document titles. If the Agent cannot find a source, report the gap. That list is more useful than a confident guess.
Service for this topic
AI Agent system integration
AI AGENT
Connect AI Agents to custom systems, apps, data, and third-party tools, with explicit data access, allowed actions, human approval, logs, and exception handling.