Private AI Appliance
Your organisation's knowledge becomes a private AI platform, installed inside the organisation
RAG Enterprise is the application layer running on top of EuLLM and the I3K Local AI appliances: one local AI platform instead of multiple disconnected AI SaaS subscriptions.
This is not an argument about hardware. The question this page answers is what your organisation can actually do once the models, the documents and the index all sit on a machine you control — and what changes the moment you stop metering it from someone else's data centre.
Early Access · planned configurations
Three layers, one machine
The appliance is not a monolith: it is a stack where each layer does one job and can be replaced without redoing the others.
RAG Enterprise
The application layer. The organisation's knowledge: document ingestion, search, answers with sources, roles and permissions, backups.
EuLLM
The inference engine. Loads the models, holds them in video memory, generates. You choose which model, and you change it without touching the layer above.
I3K Local AI appliance
The machine. GPU, memory, storage, a prepared operating system. The same software runs on every configuration: what changes is how much fits and how fast it goes.
The practical consequence of that separation is that the hardware does not lock you into the software. Start with the open-source Community edition on a machine you already have, move to an appliance when the load calls for it, and swap the model when a better one appears — without rebuilding the index or reinstalling anything.
What you can do with it
The workloads a Local AI appliance can host, grouped by the problem they solve. They are enabled module by module: nobody has to turn on all of them.
Knowledge and documents
- Private company AI chat
- Semantic search and hybrid retrieval
- RAG with traceable sources and documents
- Knowledge assistant over the internal archive
- Analysis of PDFs, Office files, contracts, procedures and technical documentation
- Document intelligence and structured extraction
- OCR and vision over scans and images
- Translation, summarisation and document generation
Models and local engine
- Local inference, with no per-token cost
- A choice between different models, replaceable over time
- Embeddings and reranking computed locally
- OpenAI-compatible internal APIs for your own applications
Agents and automation
- AI agents over internal procedures
- Workflows and automation
- Governed access to company tools and applications
- Coding assistant for technical teams
Voice and meetings
- Meeting transcription
- Meeting intelligence: summaries, decisions, actions
- Internal voice assistants and voicebots
Integration with the rest of the business
- CRM and ERP
- Help desk and ticketing
- File servers and document archives
- Databases and line-of-business systems
Read this as a list of possibilities, not of switches already flipped. RAG Enterprise covers the document layer — chat over your documents, semantic and hybrid search, answers with cited sources, OCR, local embeddings and reranking. The rest are workloads the appliance is sized to host and that are enabled as modules, according to what is actually needed: which ones enter your installation is decided during analysis, not assumed here.
One local platform instead of several separate subscriptions
For many organisations, AI spend is not one line item: it is a subscription for chat, one for transcription, one for the coding assistant, one for the tool that reads PDFs, plus metered usage on a couple of integrations. None of them talk to each other, and each one sees a slice of the company's data.
A local appliance attacks the problem from the other end: one place where the models run, one perimeter to govern, one contract. The same capabilities, served inside the organisation rather than bought piecemeal outside it.
We are not claiming it automatically replaces every service you use: some things the large providers do better, and some tools are wired into processes that do not move in a quarter. The honest formulation is that it can replace or reduce dependency on multiple external AI services — and which ones is something you work out by looking at what you actually use.
Why the arrangement holds up
Documents and data stay under the organisation's control
Not as a vendor's contractual promise, but as a consequence of where the software runs. What does not leave needs no justification, no mapping across processors, and does not depend on terms that can change.
It can run without sending content to external LLM providers
The models sit on the machine. An appliance can work without a single line of your documents reaching a third-party service.
Local inference has no per-token cost
The bill moves from consumption to hardware, which you buy once and keep. Past a certain intensity of internal use, that is the difference between a budget line that grows on its own and one that does not.
You choose the models, and change them when you want
You are not tied to whatever a provider decides to serve this month. A local model is a file: it stays what you installed until you replace it.
The same software stack on different hardware configurations
What separates the sizes is how much fits and how fast it goes, not which features are available. It scales from a small business to a large organisation, up to multi-node installations, without changing platform.
External AI services remain an option, not a requirement
If in some cases you also want a frontier model, that is possible — as a separate, governed channel you decide on and can audit. The system does not need it to work.
Three categories, not a GPU catalogue
The useful choice is not which card to fit: it is what you want the machine to be good at. Everything else follows from that.
Early Access: planned configurations, still being defined. Names, specifications and prices may change before general availability.
Capacity
Very large models and extensive knowledge bases
For holding substantial models in memory and indexing very large archives. It favours how much fits over how fast it goes.
Local AI Capacity 128
from €5,900 + VAT
Performance
High-speed RAG, chat and interactive agents
For day-to-day conversational use, where what counts is the time between the question and the first words of the answer. The right category for most offices.
Local AI Performance 32
from €5,900 + VAT
Local AI Performance 64
from €8,900 + VAT
Enterprise
More users, more workloads, multi-node installations
For large organisations running several workloads at once and serving many people concurrently, across more than one node where needed.
Local AI Enterprise 96
from €22,900 + VAT
Indicative starting prices, VAT excluded. The current list, the exact specifications and purchasing all live on the I3K Local AI landing page.
See configurations and pricing on I3KHow the licence works, briefly
The model is built to be predictable: what you pay depends on how large the organisation is, not on how much you use it.
- 12 months of the AI Suite are included with the appliance.
- From the second year onward, a subscription sized to the organisation applies.
- Local inference has no per-token cost: using it more does not raise the bill.
- Licences are sold in user bands — we do not sell the idea of “unlimited users”, because it would not be a serious promise.
- I3K support is provided to the technical contacts the organisation designates, not directly to every employee.
- New features within modules you already bought are included while the subscription is active.
- Any future new products or modules may require a separate licence.
Frequently asked questions
- What exactly is a private AI appliance?
- A machine installed on the organisation's premises, with the GPU and operating system already prepared, running the inference engine and the AI applications. The models, the documents and the index stay on that machine: traffic does not leave for an external provider. I3K Local AI is the appliance line; RAG Enterprise is the application layer that runs on it.
- Do we have to buy the hardware from you?
- No. The Community edition of RAG Enterprise is open-source under AGPL-3.0 and runs on your own server, GPU or not. The appliance is for when you want a machine already sized, prepared and supported, and when workloads beyond the document layer join in. The sensible order is almost always to install the software first on something you already have.
- Does it replace ChatGPT, Copilot and the other services we use?
- It can replace or reduce dependency on multiple external AI services, not necessarily all of them. On open-ended reasoning and creativity, frontier models remain ahead. Where an appliance wins clearly is work on the organisation's own content: there, retrieving the right passages matters far more than model size, and the documents do not have to leave.
- Which of these capabilities are available today?
- RAG Enterprise today covers the document layer: chat over documents, semantic and hybrid search, answers with cited sources, OCR, local embeddings and reranking. The other capabilities listed are workloads the appliance is sized to host and that are enabled as modules. Which ones enter your installation, and in what order, is decided during analysis.
- What is the difference between Capacity, Performance and Enterprise?
- Capacity favours how much fits: very large models and extensive knowledge bases. Performance favours responsiveness: chat, RAG and interactive agents with short answer times. Enterprise is for large organisations with several parallel workloads, more users and, where needed, more nodes. The software is the same across all three.
- Can we still use an external model when we need one?
- Yes, but as an explicit, governed choice rather than a requirement of the system. The appliance works without it. If for certain tasks you want to route a request to an external service, that is a separate channel you decide on, with its own rules about what may leave.
- When are the appliances available?
- The configurations shown are in Early Access: planned and still being defined, so names, specifications and prices may change before general availability. The I3K Local AI landing page is the reference for current status and for purchasing.
Where to start
The path we recommend is always the same: the software first, on real documents, to find out whether the answers are genuinely useful; then a measurement on your own hardware; then, only if the load calls for it, an appliance sized accordingly. Hardware is only bought well after you have seen some numbers.