AI on your own infrastructure
Local AI on existing hardware: an honest feasibility check
Starting questionA consulting firm with around 20 workstations asked: can an AI assistant run entirely on our own server?
Result
Works with conditions
Works only as a hybrid with a hardware upgrade. On the processor alone the language model was not usable interactively.
Not with the existing hardware alone: a language model running only on the processor was not usable interactively. We recommended a hybrid architecture with a targeted hardware upgrade. Confidential tasks stay with the local model; external models are used only after pseudonymization.
What we assessed
We assessed the existing hardware. It was sufficient for the application server. A language model running on the CPU alone, however, was not usable interactively: a few tokens per second and tight memory.
The recommendation was a hybrid architecture. A local model handles confidential tasks; external models are only used after pseudonymization. On top comes a targeted hardware upgrade.
What it means for business architecture
Whether AI runs on premises is decided by the infrastructure, not by intent. Checking up front prevents investments that do not hold up in daily use.
Lessons learned
- Without GPU acceleration, local inference for several users is not practical.
- Existing hardware can often be upgraded rather than replaced.
- Pseudonymization makes external models usable for non-confidential tasks.
- on-premises
- local AI
- hybrid architecture
- pseudonymization