Skip to content
Stratalytic

DATA DRIVEN DECISIONS

AI & Machine Learning

Building an Internal Knowledge Base Chatbot: From Company Data to Smart Assistant

Published:

Internal chatbot interface with knowledge base documents

Key Takeaways: An internal knowledge base chatbot with RAG architecture makes company documents searchable via natural language, saving employees an average of 45 minutes per day searching for information. Custom implementations cost 10,000-30,000 euros, SaaS alternatives 500-2,000 euros per month. This article describes the technical architecture, compares build-versus-buy options, addresses privacy requirements and shows how the WBSO subsidy covers up to 40% of development costs.

Employees spend an average of 19% of their working time searching for information, and an internal AI chatbot reduces this by 60-75%

The problem is universal and costly: company information is scattered across emails, SharePoint folders, Notion pages, manuals, CRM notes and the heads of experienced colleagues. McKinsey estimates that knowledge workers spend an average of 19% of their working time searching for and gathering information. For a company with 50 employees at an average salary of 55,000 euros per year, that is a hidden cost of more than 520,000 euros annually.

Traditional solutions like SharePoint search functions or structured wikis only partially solve the problem. They require employees to know which search terms to use and where the information resides. A RAG chatbot solves this fundamentally differently: employees ask questions in natural language and receive answers synthesised from all available company documents, with source references so the information is verifiable.

Companies implementing an internal knowledge base chatbot consistently report positive results. Deloitte documents an average time saving of 45 minutes per employee per day, a 60-75% reduction in search time. Additionally, the chatbot reduces dependence on specific individuals for knowledge transfer, which increases organisational resilience and accelerates new employee onboarding by an average of 35%.

The technology is mature enough in 2025-2026 to deliver reliable results. Hallucination rates of RAG systems have dropped to 3-8% when correctly implemented, compared to 15-25% with direct LLM interaction without a knowledge base. This makes the technology suitable for professional use, provided you make the right architectural choices.

RAG architecture: how it works technically

Retrieval-Augmented Generation, or RAG, is the architecture that combines an LLM with your company documents so the model generates answers based on your specific information rather than only its training data. The process runs in four steps: document ingestion, embedding creation, retrieval and generation.

Document ingestion is the first step where your company documents are collected, cleaned and split into manageable pieces. PDFs, Word documents, emails, wiki pages and spreadsheets are each processed with specific parsers. Splitting documents into chunks of typically 500-1000 tokens is crucial: chunks too large contain too much irrelevant information, chunks too small miss context. Overlap of 10-20% between consecutive chunks prevents information loss at split points.

Embedding creation translates each chunk into a numerical vector capturing semantic meaning. Embedding models like OpenAI's text-embedding-3-large, Cohere's Embed or open-source alternatives like BGE or E5 produce vectors of 768-3072 dimensions. Semantically similar texts receive vectors that lie close together in vector space. This makes it possible to match questions with relevant document fragments based on meaning rather than exact word matches. The cost for embedding is negligible: approximately 0.13 dollars per million tokens for OpenAI's most advanced model.

Retrieval is the process where the chatbot, given a user question, fetches the most relevant chunks from the vector database. The question is first transformed into an embedding vector and then compared with all stored document vectors via cosine similarity. The top-k most relevant chunks, typically 3-10, are selected as context. Hybrid search, combining vector search with keyword search via BM25, improves retrieval quality by 15-25% according to Pinecone benchmarks.

Generation is the final step where an LLM like GPT-4, Claude or Llama combines the selected chunks with the user question to formulate a coherent answer. A carefully designed system prompt instructs the model to answer only based on provided context, cite sources, and honestly indicate when available information is insufficient. This minimises hallucinations and maximises answer reliability.

Document ingestion: from chaos to structured knowledge base

The quality of your RAG system depends 70% on the quality of the document ingestion pipeline, not on the LLM model. Insufficient attention to this phase is the most common cause of disappointing chatbot results.

Start with an inventory of all document sources and their formats. Typical sources for an SME are: product documentation and manuals in PDF, process descriptions in Word or Notion, customer communication in the CRM, technical specifications in spreadsheets, and FAQs on the website. Prioritise sources based on the search volume they generate and the impact of quick access to that information.

PDF processing deserves special attention because it is by far the most common document format in business environments and simultaneously the most difficult to parse. Scanned PDFs require OCR processing. PDFs with tables, images and complex layouts require specialised parsers like Unstructured.io or LlamaParse that preserve document structure. Investing in a robust PDF pipeline pays back through significantly better answers.

Metadata enrichment substantially increases retrieval quality. Add metadata to each chunk such as document title, author, creation date, department and document type. This enables filtering retrieval to relevant sources. A question about HR policy need not search through technical documentation. In practice, metadata filtering improves answer relevance by 20-30%.

The update strategy determines how current your chatbot stays. Implement automated synchronisation with your document sources so new or modified documents are automatically indexed. Ready-made connectors exist for SharePoint, Google Drive and Notion. Frequency depends on how often your documents change: daily synchronisation suffices for most companies, but rapidly changing knowledge bases may require real-time synchronisation.

Build versus buy: cost comparison

The choice between building yourself and buying a SaaS solution depends on your specific requirements around customisation, privacy and budget. Both routes have clear advantages and disadvantages expressed in concrete figures.

SaaS platforms like Glean, Guru, Slite AI or Notion AI offer internal knowledge base chatbots as a service. Costs are 500 to 2,000 euros per month depending on user count and document volume. Implementation time is typically one to three weeks. The advantage is speed and convenience: you connect your document sources, the platform handles embedding, retrieval and generation. The disadvantage is limited customisation, dependence on the vendor for data processing, and ongoing costs that over three years accumulate to 18,000-72,000 euros.

A custom RAG implementation with open-source components costs 10,000 to 30,000 euros in initial development. The technical stack typically consists of LangChain or LlamaIndex as orchestration framework, a vector database like Pinecone, Weaviate, Qdrant or ChromaDB, and an LLM via API (OpenAI, Anthropic) or self-hosted (Llama, Mistral). Ongoing costs are 200 to 800 euros per month for LLM API calls, vector database hosting and infrastructure. Over three years total investment is 17,200 to 58,800 euros, comparable to SaaS but with full control.

An enterprise-grade implementation with extended features like permission management, multi-language support, feedback loops and analytics costs 30,000 to 80,000 euros initially and 500 to 2,000 euros per month ongoing. This is the route for companies with more than 100 employees, strict compliance requirements or complex document landscapes.

The break-even calculation is favourable for every variant. With 50 employees and a time saving of 45 minutes per day at an average hourly rate of 35 euros, the annual saving is approximately 205,000 euros. Even if only 30% of employees actively use the chatbot, the saving is 61,500 euros per year, more than enough to justify any implementation variant.

Privacy and security: essential for internal data

Privacy considerations for internal knowledge base chatbots are not optional but determinative for architectural choices. Business-sensitive information such as financial data, customer details, contracts and strategic plans requires adequate protection.

Data residency is the first consideration: where are your documents and embeddings stored? With SaaS platforms this depends on the vendor, often in the US. For European companies subject to GDPR, EU data residency is a hard requirement. Azure OpenAI Service offers EU-hosted LLM endpoints, and vector databases like Weaviate and Qdrant offer EU hosting options. With a custom implementation you have full control over data location.

Document-level access control prevents employees from reaching information via the chatbot they would not normally have access to. An HR employee asking the chatbot about salary structures should receive answers from HR documents but not from financial reports intended only for management. This requires your RAG system to adopt and apply the existing permission structure from your document management system during retrieval.

LLM provider choice directly affects your data privacy. When using cloud LLM APIs, your document fragments are sent to the provider as context. OpenAI and Anthropic offer enterprise contracts with data processing agreements excluding use of your data for model training. Self-hosted models like Llama 3 or Mistral eliminate this concern entirely but require dedicated GPU infrastructure costing 500 to 3,000 euros per month.

Logging and audit trails are essential for compliance. Log which questions are asked, which documents are used as sources and which answers are generated. This enables detecting misuse, monitoring answer quality and reconstructing which information was shared in case of incidents. Implement data retention policies aligned with your privacy policy.

Choosing embedding models and vector databases

The choice of embedding model and vector database significantly affects the quality and cost of your RAG system. Dozens of options now exist, but for most implementations it comes down to three decisions.

For the embedding model you choose between cloud-hosted models and self-hosted alternatives. OpenAI's text-embedding-3-large delivers the best quality available today at 0.13 dollars per million tokens and is the simplest option. Cohere's Embed v3 performs comparably and offers multilingual support relevant for companies with documentation in multiple languages. Open-source models like BGE-large or E5-mistral-7b-instruct run on your own hardware and eliminate data transfer to external parties but require GPU resources.

For multilingual company documents, which is the case at most Dutch companies, a multilingual embedding model like Cohere Embed v3 or multilingual-e5-large performs significantly better than an English-optimised model. Tests show 15-25% better retrieval quality for Dutch documents when using a multilingual model.

The vector database choice concerns scalability, hosting and cost. Pinecone is the most popular managed option: easy to use, scalable and available from 70 dollars per month. Weaviate and Qdrant are open-source alternatives you can self-host for maximum control. ChromaDB is ideal for prototyping and small implementations: free, locally runnable and easy to integrate. For production with more than 100,000 documents, Pinecone, Weaviate or Qdrant are the more reliable choices.

WBSO subsidy for chatbot development

Building an internal knowledge base chatbot qualifies in many cases for the WBSO subsidy, which reimburses up to 40% of wage costs and outsourced R&D. The technical innovation lies in multiple aspects of the project.

Developing a document ingestion pipeline that processes company-specific document formats, terminology and structures is technically innovative when it exceeds standard solutions. Building a hybrid retrieval system that combines vector search with keyword search and metadata filtering for optimal results also qualifies. Implementing domain-specific evaluation and feedback systems that continuously improve answer quality is a third innovation axis.

For a custom chatbot project of 25,000 euros in development costs, the WBSO yields approximately 8,000 to 10,000 euros in subsidy. The AIP scheme can additionally be relevant when the project contains machine learning components, such as fine-tuning embedding models on your domain-specific data or training a classification model for document routing.

The application procedure is straightforward: you submit a WBSO application to RVO before the project starts with a technical description of the innovative aspects. Processing time is six to eight weeks. A specialised subsidy consultancy can handle the application for 1,500-3,000 euros, an investment that more than pays for itself upon approval.

Conclusion: the ROI is beyond question

An internal knowledge base chatbot is one of the AI applications with the most compelling ROI for SMEs. The combination of significant time savings, improved knowledge sharing and accelerated onboarding makes the payback period short and the impact broad.

Start with an inventory of your document sources and employee search behaviour. Evaluate whether a SaaS solution meets your privacy and customisation needs, or whether a custom implementation is justified. Begin with a pilot on a bounded document set, validate answer quality with end users, and expand gradually. With the WBSO and AIP scheme you reduce the investment by up to 40%, making the business case viable even for companies with 20 employees.

Get the AI-subsidy radar

1 email per month. New subsidies, deadlines, and what changed for SMEs. 5-minute read.

Unsubscribe with one click. No spam, ever.

Let's talk business

Do you want to know how we can help you grow your business? Schedule free consultation with one of our experts and discover the possibilities.

Rutger Geerlings, founder of Stratalytic

Rutger Geerlings

Solution Architect

Discover what data and AI can concretely deliver

Latest cases

All cases