AI on your own infrastructure

Local AI on existing hardware: an honest feasibility check

AssessedPublished:

Starting questionA consulting firm with around 20 workstations asked: can an AI assistant run entirely on our own server?

Result

Works with conditions

Works only as a hybrid with a hardware upgrade. On the processor alone the language model was not usable interactively.

Not with the existing hardware alone: a language model running only on the processor was not usable interactively. We recommended a hybrid architecture with a targeted hardware upgrade. Confidential tasks stay with the local model; external models are used only after pseudonymization.

What we assessed

We assessed the existing hardware. It was sufficient for the application server. A language model running on the CPU alone, however, was not usable interactively: a few tokens per second and tight memory.

The recommendation was a hybrid architecture. A local model handles confidential tasks; external models are only used after pseudonymization. On top comes a targeted hardware upgrade.

What it means for business architecture

Whether AI runs on premises is decided by the infrastructure, not by intent. Checking up front prevents investments that do not hold up in daily use.

Lessons learned

  1. Without GPU acceleration, local inference for several users is not practical.
  2. Existing hardware can often be upgraded rather than replaced.
  3. Pseudonymization makes external models usable for non-confidential tasks.
  • on-premises
  • local AI
  • hybrid architecture
  • pseudonymization

More reports from the Lab

Consulting by SIMO GmbH

From the Lab to consulting.

This report shows what works technically. Whether it works in your company is a question of architecture. SIMO GmbH answers it with you, independent of vendors.

Decision Readiness Check: 12 questions, about 3 minutes, result at once. Initial call: 45 minutes, free of charge, by video. Entry with BISA 1 or BEIA. All on simo-online.com.