Field report 12 · AI on your own infrastructure

Local AI on existing hardware: an honest feasibility check

Can an AI assistant run entirely on your own server? A feasibility check for a firm with around 20 workstations.

Status
Assessed Feasibility check for a client project, anonymised.
Published
Publisher
SIMO GmbH, Aschaffenburg, Germany

Starting questionA consulting firm with around 20 workstations asked: can an AI assistant run entirely on our own server?

What we assessed

We assessed the existing hardware. It was sufficient for the application server. A language model running on the CPU alone, however, was not usable interactively: a few tokens per second and tight memory.

The recommendation was a hybrid architecture. A local model handles confidential tasks; external models are only used after pseudonymisation. On top comes a targeted hardware upgrade.

What it means for business architecture

Whether AI runs on premises is decided by the infrastructure, not by intent. Checking up front prevents investments that do not hold up in daily use.

Learnings

  1. Without GPU acceleration, local inference for several users is not practical.
  2. Existing hardware can often be upgraded rather than replaced.
  3. Pseudonymisation makes external models usable for non-confidential tasks.
  • on-premises
  • local AI
  • hybrid architecture
  • pseudonymisation

Matching consulting service

From the Lab into consulting.

What we trial in the Lab, SIMO GmbH puts into practice as consultants: Business Data Strategy & Architecture for AI, following the Zero Friction Data Flow principle. These services relate to this report:

45 minutes, free of charge, with the SIMO GmbH consultants. The form is on simo-online.com.