A private LLM for the files you cannot paste into a public AI.
Contracts, HR files, source code, figures before publication: the documents your teams most need help with are the ones nobody may paste into a public AI. We install a language model on your servers or in a European cloud, and those files can finally go through it.
With a public AI service
- Your prompts and files are processed on the provider’s servers.
- Retention and usage terms are the provider’s to change.
- You cannot audit what is kept, or for how long.
- Sensitive documents stay banned from every AI tool.
With your private LLM
- It runs on your servers or in an EU cloud you choose.
- You choose the model, its version and when it changes.
- Every request is logged on your side, under your policy.
- Contracts and HR files can go through it too.
What is a private LLM?
A private LLM is a large language model that runs in an environment you control: your own servers, a private cloud or a dedicated instance in Europe. Prompts, documents and answers stay inside that perimeter and are never used to train anyone else’s model. Self-hosted or on-premise, you choose the model, its version and who can reach it.
An on-premise LLM is the strictest form: the model sits in your data center, and can even run with no internet connection at all. A private cloud deployment is often the better balance: dedicated resources in an EU region, with the hardware looked after by someone else.
The models come from open-weight families such as Mistral, Llama or Qwen, and we check each license against your use. We do not quote benchmark scores. Candidate models run on your own tasks, you see the results, and we recommend the smallest one that does the job well.
Public API, EU-hosted API, private cloud or on-premise: where should your LLM run?
Four ways to run a language model, from the most convenient to the most controlled. The right one depends on what your data allows.
| Criterion | Public AI API | EU-hosted API | Private cloud (EU) | On-premise |
|---|---|---|---|---|
| Where your data is processed | Provider’s servers, possibly outside the EU | Provider’s servers in the EU | Dedicated resources in an EU region | Your own data center |
| Who can reach it | The provider, under its terms | The provider, under its terms | You and your hosting partner | You alone |
| Choice of model | The provider’s catalog | The provider’s catalog | Any open-weight model | Any open-weight model |
| Works without internet | No | No | No | Yes, if required |
| How costs behave | Pay per use | Pay per use | Reserved capacity, monthly | Hardware upfront, then running costs |
| Best for | Non-sensitive data, fast start | Most business data | Sensitive data, steady use | Regulated or confidential data |
What your security team can verify for itself.
Security and dataThe perimeter is written down.
Network rules state exactly what may go out, and your team can test them.
No training on your data.
The model is used as it is, or tuned only for you, and the result stays yours.
Logs on your side.
Every request and answer can be traced, under the retention policy you decide.
Your access rules, mirrored.
People only reach through the AI the documents they could already open.
Documented for your DPO.
Data flows, hosting, retention and AI Act obligations are written up for review before go-live.
The confidential work your teams can now hand to AI.
Six typical cases. Confidentiality kept AI away from each of them, and once the model runs on your side, they are often the ones with the most to gain.
Contract review
BeforeLawyers read every clause by hand, because contracts cannot go to a public service.
AfterThe model flags unusual clauses and missing terms, on your servers, and a lawyer decides.
HR files
BeforeHR answers the same policy questions by hand, keeping employee data well away from any AI.
AfterAn assistant answers from your policies and files, within each person’s access rights.
Source code
BeforeDevelopers are asked not to paste code into AI tools, so they go without the help.
AfterA coding assistant runs inside your network, on your repositories, with nothing sent out.
Research and patents
BeforeR&D notes and drafts stay out of every AI tool, because a leak would cost the lead.
AfterResearchers search, summarize and compare their own work with a model that never leaves the lab.
Finance before publication
BeforeFigures under embargo are handled by a few people, with no help and no margin for error.
AfterDrafts, checks and commentary are prepared by the model, inside the finance perimeter.
Public sector records
BeforeCitizen files and internal notes cannot be sent to a service outside your control.
AfterStaff draft replies and find past decisions with an AI hosted under your jurisdiction.
From perimeter to first users, step by step.
Week 1
Set the perimeter
Which data, which users, where the model must run, and what the law and your own policies require. The week ends with a written scope and a fixed price.
Weeks 2 to 3
Test models on your tasks
Candidate open-weight models are run on a set of your real tasks. You see the results side by side and choose.
Weeks 3 to 5
Install and connect
The model is deployed on your servers or private cloud, linked to your sign-in and your documents. If hardware has to be bought, its delivery sets the pace.
Week 6
Open to your teams
A first group starts using it, with monitoring of quality, load and usage, then the circle widens.
Where the model runs sets most of the price.
Your own hardware, a private cloud or a dedicated EU instance: that choice weighs the most. The rest depends on how many people use the model and what for. We fix the price in writing once the perimeter is set, before anything is installed.
See market ranges in the AI agent cost guide- 01Where it runs: your own hardware, a private cloud or a dedicated EU instance.
- 02The size of the model, which sets the computing power it needs.
- 03The number of users and how many use it at the same time.
- 04The documents and systems it has to be connected to.
- 05Security, audit and certification requirements in your sector.
- 06Who operates and updates it once it is live.
Private LLM: the questions buyers ask us
Is a private LLM as good as the big public models?
It depends on the task. For summarizing, extracting, drafting and answering from your documents, open-weight models are often more than enough. For the hardest reasoning, the largest cloud models may keep an edge. We test on your own cases before recommending, and a mix of both is possible.
What hardware do we need to run an LLM on premise?
It depends on the model size and the number of simultaneous users. A small model can run on a single server with one graphics card; larger models need more. We size the hardware from your real load, or propose a private cloud if buying is not worth it.
Which models do you deploy?
Open-weight families such as Mistral, Llama and Qwen, among others. We choose on your tasks, your languages and the license terms, which differ from one model to another and are checked for your use before anything is installed.
Can the model run with no internet connection?
Yes. An on-premise installation can run fully disconnected from the internet, which suits the most sensitive environments. Updates to the model or the software are then delivered through a controlled procedure that your team reviews and approves.
Does a private LLM make us GDPR and AI Act compliant?
It removes a large part of the risk, because the data no longer leaves your control. Compliance also depends on the purpose, retention and information given to people. We document all of it so your DPO can review and sign off.
Can we start in the cloud and move on premise later?
Yes. We build the application layer so that the model can be moved without rewriting it. A common path is to start in a private EU cloud to prove the value, then bring the model in-house once usage is known.
Often built together
- Knowledge assistant To ask your private model about your documents.
- AI agents To give your private model real tasks.
- AI integration To bring it into the tools your team uses.
- Sovereign AI, explained What the term covers, and what it takes in practice.
- EU AI Act compliance What the regulation asks of your AI systems.
- AI governance Who decides which data and which uses the model is allowed.
- GDPR and AI What the GDPR asks of a company that uses AI, article by article.
Which documents would you finally trust to an AI?
Tell us which data must stay private and where it sits today. Within one business day, we come back with the setup we would recommend for it. The first 30-minute assessment is free.