Back to blog
ProductivityData Sovereignty

Can you use AI on confidential documents without sending data to third-party servers?

Private AILuxembourgGDPRSME Packages

Builder & Founder

Dirigeante d'une PME luxembourgeoise relit un contrat confidentiel traduit par l'assistant IA privé de son entreprise

Short answer: yes, if the architecture is chosen before the tool

Yes: a company can use AI on confidential documents without those documents leaving its perimeter, provided it chooses the architecture before the tool. Three architectures make this possible: a model deployed on your own infrastructure, a model hosted in a dedicated European environment governed by contract, and, as a complement to either one, an internal document base that the model consults without absorbing. The criterion is not the brand of the tool, it is the path the data takes: who hosts it, who accesses it, what is retained (a US provider remains subject to the CLOUD Act of March 23, 2018, wherever the data is stored).

This article puts three uses under the microscope: translation, PDF analysis and the internal assistant. For each one, it shows what a private AI document processing agent executes and where the data goes. It ends with the seven questions to ask a provider.

1. Where a document goes when sent to a consumer AI tool

A German-language contract pasted into a free translation tool has just left the company.

The path is always the same: upload to the provider's server, often outside the European Union, retention under its terms of use, sometimes reuse to improve the service. A US provider remains subject to the CLOUD Act of March 23, 2018, which allows United States authorities to request data from it regardless of the country where that data is stored.

What sets apart a ChatGPT alternative for a Luxembourg company comes down to this single path.

2. Three architectures that keep the data inside your perimeter

A sovereign AI processes the data in an environment you control or have governed by contract. Three architectures make this possible. The first two answer the question "where does the model run"; the third is added to either one and answers "how does it read your documents".

Architecture 1: the model on your infrastructure. A model you install yourself runs on a server you administer. Nothing leaves; it is the most demanding option in terms of hardware and operations.

Architecture 2: the dedicated European environment. The model runs in an instance reserved for your company, with a European host, under a data processing agreement compliant with Article 28 of the GDPR: no reuse, no training, a fixed retention period. This is the most common option for an SME.

Architecture 3: the internal document base (RAG). RAG, or retrieval-augmented generation, indexes your documents in a base you control. For each question, the model receives only the relevant passages, answers, and retains none of it.

Three architectures for processing confidential documents with AI

Architecture

Who hosts

Who accesses

What is retained

When to choose it

Model on your infrastructure

You or your managed service provider

Your teams

What you decide

Technical team in place, very high-risk data

Dedicated European environment

European host, reserved instance

Your teams; the provider on written instruction

Nothing beyond the term of the contract

Starting point for most SMEs

Internal document base (RAG)

You, in your DMS (document management system)

Each employee, with their own access rights

The index stays with you; nothing in the model

To combine as soon as an assistant queries your documents

3. Three uses under the microscope: translation, PDFs, internal assistant

3.1 Translating documents across four languages

As of January 1, 2026, 46.6% of Luxembourg residents are foreign nationals (Statec, STATNEWS No. 15/2026): translation is a daily use. The agent receives the document in the dedicated environment, splits it up, and translates each segment with your glossary. The data makes a single trip, from your DMS to the dedicated instance and back, with no copy held by a third-party vendor. This case features among the nine use cases for AI agents in Luxembourg.

3.2 PDF analysis: extraction, checking, summary

A custom-built business AI agent reads the PDF in your index, extracts a structured record from it (parties, amounts, deadlines), and writes the summary in the requested language. The PDF stays in your DMS; only the passages needed for the question travel, to the model you host or have governed by contract.

3.3 Internal assistant on the document base

This is not the chatbot on a website. This assistant works for your teams, on your procedures, standard contracts and precedents. It answers from the internal index and inherits the existing access rights: an employee only sees what they would be allowed to open themselves. Every exchange is logged on your side, nothing feeds a shared model.

💡 Worth knowing: a Luxembourg para-public organization has deployed a private sovereign LLM, hosted in Europe, whose answers draw on an internal document base.

4. What the law requires, and what it does not

The GDPR, applicable since May 25, 2018, does not mention AI. Article 28 requires a written contract with every processor; its paragraph 2 makes the use of another processor subject to your prior written authorization, general or specific.

Article 32 requires security appropriate to the risk. Articles 44 and following govern transfers outside the European Union. The full framework is in our article on AI and the GDPR in business.

Some sectors add professional secrecy: Article 35 of the Law of August 10, 1991 for lawyers, Article 300 of the Law of December 7, 2015 for the insurance sector.

Article 4 of the AI Act, applicable since February 2, 2025, asks companies to promote AI literacy among the people who use these systems. The employee who pastes a contract into a free tool falls within this obligation. The regulation has been generally applicable since August 2, 2026; its Article 50 requires, from that same date, that a person be informed when they are interacting with an AI.

What the law does not require: a server on your premises. On-site hosting is a choice of risk level, not a legal obligation.

5. Seven questions to ask a provider before entrusting it with a document

  1. Where are the model and my documents hosted? Country, host, dedicated or shared instance.
  2. Which providers other than you touch my data? Each one requires your prior written authorization, general or specific (Article 28, paragraph 2, of the GDPR).
  3. How long are my queries and my documents retained? A duration, a location, a purge.
  4. Is my data used to train a model? The only acceptable answer: a written commitment of non-reuse.
  5. Who accesses my data on the provider's side, and is access logged? A list of named roles, and a log you can consult.
  6. Is the data encrypted at rest and in transit? Yes to both, with the encryption standard named.
  7. How do I retrieve and erase everything at the end of the contract? Open format, timeframe, certificate of erasure.

A vague answer to any one of the seven questions counts as a no.

6. The Luxembourg case: four languages, regulated sectors, state aid

Luxembourg combines four working languages with a high density of regulated sectors: accounting firms, lawyers, brokers, healthcare, entities supervised by the CSSF, the para-public sector. An aid scheme makes private AI accessible to an SME. The SME Packages Digital and AI cover up to 70% of eligible costs for a project between EUR 3,000 and EUR 25,000 excluding VAT (source: guichet.lu, "SME Packages Digital" and "SME Packages AI" factsheets, consulted on September 8, 2026). Our guide to choosing a sovereign AI partner in Luxembourg takes up the criteria above; our selection of AI tools to choose when your data is sensitive completes them.

FAQ: your questions on AI and confidential documents

1. Can AI translation tools be deployed without sending sensitive data to third-party servers?

Yes, provided the architecture is chosen before the tool: the engine runs on your infrastructure or in a dedicated European environment under a contract compliant with Article 28 of the GDPR, and nothing is retained once the document has been returned. A consumer translation tool guarantees neither the hosting nor the erasure by default. In Luxembourg, where 46.6% of residents are foreign nationals (Statec, January 1, 2026), translation is a daily use.

2. Which "chat with PDF" tools are reliable for analyzing confidential documents?

Reliability does not depend on the brand but on three answers: where the PDF is stored, who accesses it, how long it is retained. A service whose provider is American remains subject to the CLOUD Act of March 23, 2018, regardless of the country where the PDF is stored. An agent that reads the PDF in your own index and retains nothing passes all three.

3. Does an AI assistant connected to internal data improve productivity without exposing that data?

Yes, if the index it consults is hosted on your side (internal RAG) and inherits your access rights. Three checks are enough: the index stays inside your perimeter, each employee only sees what they are allowed to open, the exchanges are logged on your side. Article 4 of the AI Act, applicable since February 2, 2025, also asks that AI literacy be promoted among the employees who use such an assistant.

4. Which AI tools are suitable for highly confidential transaction data?

Those that deploy on your infrastructure or in a dedicated European environment, not a consumer tool. For an entity supervised by the CSSF, the hosting, the data processing agreement (Article 28 of the GDPR) and the security measures (Article 32) are documented before the first test.

5. Do you need a server on your premises to comply with the GDPR?

No. The GDPR, applicable since May 25, 2018, imposes no hosting location: it governs processing (Article 28) and requires security appropriate to the risk (Article 32). An on-site server is a choice of risk level, not an obligation.

The path of the data decides, not the tool

A company can translate, analyze and query confidential documents with AI without them leaving its perimeter, by choosing the architecture before the tool: the model on your side or in a dedicated European environment, plus an internal document base if an assistant queries your documents. The GDPR governs this choice without imposing an on-site server, and seven questions to the provider are enough to verify that the data does not leave.

Do you have a type of document and a doubt about its path?

📞 Discuss your use case