Answers about private AI.
What local and private AI is, what it does and does not solve, how a project runs and what you need to get started.
What is local AI?
What is local or private AI?
With local AI, the language models and the application around them run on infrastructure you manage or that is ring-fenced for you, rather than through a public model API. That way you keep control over which data is processed and who can access it.
Does private AI always run offline?
No. Fully offline operation is an option where the process calls for it, but private and hybrid setups can also be appropriate. In a hybrid setup we combine local processing with external services on purpose, with clear boundaries.
What exactly does Gold IT Services do?
We build complete private and local AI systems for businesses and organisations, from models and infrastructure to document processing, knowledge assistants, automation, user interfaces and monitoring. Your workflows, documents and data shape the solution.
What kinds of applications do you build?
Among others: knowledge assistants on your own documents (RAG), document translation, audio transcription, automated analyses and AI workflows with room for human review. The projects page shows current projects and anonymous cases.
Data, security and GDPR
Does local deployment make a system GDPR-compliant?
No. The environment you choose helps with data control, but legal basis, access, retention and operational practice still matter. Local processing can help keep sensitive content away from public model APIs, but it does not guarantee compliance.
Do my documents and data leave my environment?
That depends on the environment you choose: your own infrastructure, a ring-fenced private environment, or hybrid. Models and the application can run on your own servers under your organisation’s control. We agree in advance where processing takes place.
How do I know an answer from the AI is correct?
For knowledge assistants we show the sources next to answers, so you can read the document or passage yourself. When files or recordings are processed there is room for human review. A language model can be wrong, which is why we test with real data before anything goes into production.
Costs, hardware and technology
Does local AI remove the cost of using AI?
Not entirely. Local inference can reduce or avoid external API token fees, but hardware, power, operations and maintenance remain part of the business case. Running your own inference environment can reduce dependence on public model providers and their token pricing.
What hardware do I need?
It follows from the task, the number of concurrent users, the context length you need and the response time you can accept. As a guide: an entry system with one 24 GB GPU suits smaller models and one to three users, while production with more users needs several GPUs. We first define what the AI must do and then choose the hardware.
Which models and software do you use?
We choose open models per task; for Dutch documents and translation, for example, Qwen, Mistral and Gemma often work well, and for speech we use Whisper. As inference runtime we use vLLM for many parallel requests or llama.cpp for flexibility and smaller servers. The choice depends on context length, VRAM, latency and required quality.
What is RAG?
RAG stands for retrieval-augmented generation. Documents are split up and indexed; when someone asks a question, the application retrieves the relevant passages and passes them to the model as context. The model then answers from your own sources instead of general model knowledge, and can point to those sources.
Approach and getting started
How does a project work?
We start with the problem and the process, then review data, access rights and constraints and test a design or prototype on representative examples. We then build and connect the application, test it with real data and put operations, monitoring and documentation in place.
Can it work with our existing systems?
Often, yes. We connect sources such as documents and shared folders to existing applications and APIs. We first review interfaces, data formats and access rights.
What does a project cost and how long does it take?
That varies per application and depends on scope, data, integrations and the chosen environment, so we do not quote a fixed price or timeline up front. In a first conversation we assess feasibility and a sensible first step.
How do I get started?
Tell us which task you want to improve, what data it uses and where the solution should run. Email info@golditservices.nl or call 0657677128 and we will discuss the options.
Could private AI help with a process in your organisation?
Tell us how the process works today, what data is involved and what a useful outcome would look like.
Mark GouderjaanGold IT Services