Learning

LLM orchestration:
who decides where a request goes?

No single language model is the best, the cheapest and the safest for every task at once. Orchestration means deciding deliberately for every request.

What it is about

In almost every project several models are on the table: a local one for confidential content, a large external one for demanding texts, a small one for routine work. The question is not which model is best but by which rules a request goes to which model.

In the Lab we tested these rules in three places: classification before the call, assessing local hardware and cutting the cost of tool output.

Learnings

What we learned

  • Classify first, then route

    Without rating confidentiality before the call, every split between local and external remains a matter of trust.

  • Confidential when in doubt

    The direction of the error has to be built in. That costs some quality and is worth it.

  • Cut costs in-house

    External compression services are ruled out for customer data. Optimisation has to run locally and prove that nothing essential is lost.

What it means for the architecture

Orchestration is an architecture decision: rules for confidentiality, quality and cost belong in one central layer, not in every single application.

Infrastructure sets the limits. Whether a local model holds up is decided by the hardware, not by intent. Checking up front prevents bad investments.

Related in the Lab

Field reports and articles

From the Lab into consulting

Does this fit your project?

SIMO GmbH puts into practice what holds up in the Lab: Business Data Strategy & Architecture for AI. The initial call takes 45 minutes and is free of charge.