Field report 15 · Operations and observability
Fewer tokens without losing the line that matters
How to cut the cost of tool output in AI agents without handing customer data to external services.
- Status
- In trial Design in trial, not yet in production.
- Published
- Publisher
- SIMO GmbH, Aschaffenburg, Germany
Starting questionHow do we cut the cost of tool output in AI agents without handing customer data to external compression services and without losing important information?
What we trialled
Until now, long output was simply truncated. The error message at the end of a test run got lost in the process.
The design calls for a local, deterministic filter. Error lines, paths, exit codes and identifiers stay untouched; structured data passes through unchanged. The filter first runs in measurement mode without altering any responses.
What it means for business architecture
Optimising AI costs is also a question of data sovereignty. External services are ruled out for customer data. The optimisation has to happen in-house and prove that nothing essential is lost.
Learnings
- Truncating is not compressing. It often removes the most important part.
- A measurement mode before switching over builds trust in the savings.
- Adopt the method from a tool, not its code. The tool does not know your requirements for tenants and logging.
- tokens
- AI agents
- data sovereignty
- cost optimisation