---
title: "Better too careful: classifying data before every AI call"
description: "How we decide, before calling a language model, whether an input is confidential and may leave the organization."
canonical: "https://simosphereai.com/en/field-reports/classifying-data-before-ai-calls"
lang: en
---

# Better too careful: classifying data before every AI call

How we decide, before calling a language model, whether an input is confidential and may leave the organization.

- Published: 2026-10-06
- Updated: 2026-10-06
- Status: erprobung
- Publisher: SIMO GmbH

## Starting question

Before calling a language model, how do we decide whether an input is confidential and may leave the organization?

## Result: open

Design in testing: a local model classifies every input, as confidential when in doubt. That costs model quality, deliberately.

A local model classifies every input before it leaves the organization and treats it as confidential when in doubt. Only clearly uncritical borderline cases go on to central review. That direction is built in and cannot be relaxed by configuration.

## What we tested

We designed a two-stage classifier. A local model labels every input. When in doubt, it marks the input as confidential and does not forward it to the central review. Only clearly harmless borderline cases go there.

The bias toward “confidential when in doubt” is built in and cannot be relaxed through configuration.

## What it means for business architecture

Classification is the precondition for splitting requests between a local and an external model. Without it, every hybrid architecture remains a matter of trust.

## Lessons learned

- Erring on the side of confidentiality costs model quality. We accept that deliberately.
- Two classifiers have to stay aligned, otherwise users see contradictory results.
- A classification that only runs centrally would send out exactly the data it is meant to protect.

Rendered version: https://simosphereai.com/en/field-reports/classifying-data-before-ai-calls
