Field report 15 · Operations and observability

Fewer tokens without losing the line that matters

How to cut the cost of tool output in AI agents without handing customer data to external services.

Status
In trial Design in trial, not yet in production.
Published
Publisher
SIMO GmbH, Aschaffenburg, Germany

Starting questionHow do we cut the cost of tool output in AI agents without handing customer data to external compression services and without losing important information?

What we trialled

Until now, long output was simply truncated. The error message at the end of a test run got lost in the process.

The design calls for a local, deterministic filter. Error lines, paths, exit codes and identifiers stay untouched; structured data passes through unchanged. The filter first runs in measurement mode without altering any responses.

What it means for business architecture

Optimising AI costs is also a question of data sovereignty. External services are ruled out for customer data. The optimisation has to happen in-house and prove that nothing essential is lost.

Learnings

  1. Truncating is not compressing. It often removes the most important part.
  2. A measurement mode before switching over builds trust in the savings.
  3. Adopt the method from a tool, not its code. The tool does not know your requirements for tenants and logging.
  • tokens
  • AI agents
  • data sovereignty
  • cost optimisation

Matching consulting service

From the Lab into consulting.

What we trial in the Lab, SIMO GmbH puts into practice as consultants: Business Data Strategy & Architecture for AI, following the Zero Friction Data Flow principle. These services relate to this report:

45 minutes, free of charge, with the SIMO GmbH consultants. The form is on simo-online.com.

More reports on the same topic