Learning
LLM orchestration:
who decides where a request goes?
No single language model is the best, the cheapest and the safest for every task at once. Orchestration means deciding deliberately for every request.
What it is about
In almost every project several models are on the table: a local one for confidential content, a large external one for demanding texts, a small one for routine work. The question is not which model is best but by which rules a request goes to which model.
In the Lab we tested these rules in three places: classification before the call, assessing local hardware and cutting the cost of tool output.
Learnings
What we learned
Classify first, then route
Without rating confidentiality before the call, every split between local and external remains a matter of trust.
Confidential when in doubt
The direction of the error has to be built in. That costs some quality and is worth it.
Cut costs in-house
External compression services are ruled out for customer data. Optimisation has to run locally and prove that nothing essential is lost.
What it means for the architecture
Orchestration is an architecture decision: rules for confidentiality, quality and cost belong in one central layer, not in every single application.
Infrastructure sets the limits. Whether a local model holds up is decided by the hardware, not by intent. Checking up front prevents bad investments.
Related in the Lab
Field reports and articles
Classifying data before the AI call
A two-stage classifier, confidential when in doubt.
Local AI on existing hardware
An honest feasibility check.
AI orchestration in the mid-market
Blog: why the next tool does not solve the problem.
From the Lab into consulting
Does this fit your project?
SIMO GmbH puts into practice what holds up in the Lab: Business Data Strategy & Architecture for AI. The initial call takes 45 minutes and is free of charge.