Field report 12 · AI on your own infrastructure
Local AI on existing hardware: an honest feasibility check
Can an AI assistant run entirely on your own server? A feasibility check for a firm with around 20 workstations.
- Status
- Assessed Feasibility check for a client project, anonymised.
- Published
- Publisher
- SIMO GmbH, Aschaffenburg, Germany
Starting questionA consulting firm with around 20 workstations asked: can an AI assistant run entirely on our own server?
What we assessed
We assessed the existing hardware. It was sufficient for the application server. A language model running on the CPU alone, however, was not usable interactively: a few tokens per second and tight memory.
The recommendation was a hybrid architecture. A local model handles confidential tasks; external models are only used after pseudonymisation. On top comes a targeted hardware upgrade.
What it means for business architecture
Whether AI runs on premises is decided by the infrastructure, not by intent. Checking up front prevents investments that do not hold up in daily use.
Learnings
- Without GPU acceleration, local inference for several users is not practical.
- Existing hardware can often be upgraded rather than replaced.
- Pseudonymisation makes external models usable for non-confidential tasks.
- on-premises
- local AI
- hybrid architecture
- pseudonymisation