Learning
Data-driven AI:
data first, then AI.
Data-driven AI means artificial intelligence works systematically on your organisation’s own data, with control over provenance, quality and clearance.
What it is about
A general AI tool answers from world knowledge. It knows neither your price list nor the state of your contracts nor your customer history. AI only becomes useful once it works on your own data and names its sources.
That shifts the real work away from the model towards data model, permissions and governance. That is exactly what our trials in the Lab show.
Learnings
What we learned
Clearance belongs in the data model
Whether an AI system may use a piece of content is a field on the content, not a process afterwards. The safe default is “not cleared”.
Name the gaps
External and internal data only become an asset with proven provenance and coverage. An answer must say what is missing.
Need-to-know as architecture
Who may know what has to be enforced technically, at every hand-over between components.
What it means for the architecture
Data-driven AI starts with a business data strategy: which data supports which decision, who owns it, and under which rules may an AI use it?
The data architecture follows from that. Keep raw and prepared data apart, carry provenance along, enforce revocation all the way into the search index.
Related in the Lab
Field reports and articles
Knowledge with clearance
Content an AI may only use with permission.
Annual accounts as data
Making public registers machine-readable.
Sovereign AI is not a seal
Blog: four questions decide it.
From the Lab into consulting
Does this fit your project?
SIMO GmbH puts into practice what holds up in the Lab: Business Data Strategy & Architecture for AI. The initial call takes 45 minutes and is free of charge.