Field report 3 · AI governance

Better too careful: classifying data before every AI call

How we decide, before calling a language model, whether an input is confidential and may leave the organisation.

Status
In trial Design in trial, not yet in production.
Published
Publisher
SIMO GmbH, Aschaffenburg, Germany

Starting questionBefore calling a language model, how do we decide whether an input is confidential and may leave the organisation?

What we trialled

We designed a two-stage classifier. A local model labels every input. When in doubt, it marks the input as confidential and does not forward it to the central review. Only clearly harmless borderline cases go there.

The bias towards “confidential when in doubt” is built in and cannot be relaxed through configuration.

What it means for business architecture

Classification is the precondition for splitting requests between a local and an external model. Without it, every hybrid architecture remains a matter of trust.

Learnings

  1. Erring on the side of confidentiality costs model quality. We accept that deliberately.
  2. Two classifiers have to stay aligned, otherwise users see contradictory results.
  3. A classification that only runs centrally would send out exactly the data it is meant to protect.
  • data classification
  • hybrid architecture
  • local AI
  • confidentiality

Matching consulting service

From the Lab into consulting.

What we trial in the Lab, SIMO GmbH puts into practice as consultants: Business Data Strategy & Architecture for AI, following the Zero Friction Data Flow principle. These services relate to this report:

45 minutes, free of charge, with the SIMO GmbH consultants. The form is on simo-online.com.

More reports on the same topic