Lakehouse to Ontology in a Day
Two of your data sources and one decision your business makes every week, taken all the way from where the data lives to what it means: how the two sources actually join, who owns and can find what is in them, and the first ontology that turns both into objects, links and the actions people are allowed to take against them. By the end of the day you know whether your data can carry AI at all — which is the question that quietly decides whether any programme built on top of it will work.
- Fixed price
- €1,500
- Duration
- One day, hands on, with your data or a redacted extract
- Who it needs
- Four to eight people, engineers and one process owner
Most data platforms stall in the same three places, in this order — and the day walks them in that order, because you cannot skip one and land the next.
Integration
Do these two sources even join?
Keys, grain, freshness, and how the data would really move: batch, change capture or an API.
Without it, every downstream number has a caveat nobody can explain.
Discoverability
Can anyone find it, and would they trust it?
Who owns each source, what is documented, what lineage exists, and which definitions two systems disagree about.
Without it, the data exists and nobody outside the team that built it can use it.
Meaning
What does any of it actually mean?
The ontology: objects, links, and the actions a person or an agent may take against each.
Without it, tables stay tables — and an agent believes whichever version it is handed.
That last one is where programmes actually die, because tables are not the things the business talks about. An ontology is the layer where a shipment, a claim or a patient becomes one object with an owner, a lineage and a set of permitted actions.
- Hold a shipment
- Re-plan a route
- Escalate to a human
- Settle without review
Who should be in the room
- Whoever knows what the fields really mean
- One or two data or platform engineers
- The owner or steward of each source
- A business process owner
Four to eight people. The last one is the one people forget: the business process owner who makes the decision we are modelling. An ontology built without that person describes the database instead of the work.
What we need in advance
Two sources, agreed a week ahead
Two is deliberate: one teaches nothing about links, three turns the day into a workshop about scope.
One decision somebody makes repeatedly
Which repair crew goes where tomorrow, which claim can be settled without a human, which batch may ship.
Access, or an honest substitute
If you can share a redacted extract, we bring a sandbox and build on it live. If nothing may leave your perimeter — and for many of our clients nothing may — we work from schemas and synthetic data shaped like yours, and the modelling half of the day is unchanged.
How the day runs
- 09:30
The decision, not the schema
We map how that one decision is made today: who asks, what they look at, where they wait, what they override. This becomes the test everything else has to pass.
- 10:15
Integration: what the sources contain, and how they would ever meet
Keys, grain, cardinality, freshness, the fields that were repurposed in 2019 and never renamed — and then how each source would really be moved: batch, change capture or an API, how often, and what breaks when the volume is real. The gap between the data dictionary and the data is where most programmes lose their first year.
- 11:15
Discoverability: who owns it, and would anyone trust it
For each source: who owns it, what is documented, what lineage exists, and which definitions two systems disagree about — where "active customer", "delivered" or "case closed" mean two different things depending on who is asked. Data nobody can find or trust is not an asset, and an agent will believe whichever version it is handed.
- 13:30
Meaning: objects, links, actions
The modelling itself, and the reason for the other three hours. What is an object, what is an attribute, what is a relationship, which semantic definition wins where two sources disagreed this morning, and which actions a person or an agent may take against each object — because an ontology that cannot be acted on is a diagram.
- 15:00
Landing it, and asking it questions
Your extract into a lakehouse, the ontology on top of it, in a sandbox we bring — or, in the closed case, the same model against synthetic data. You watch it being built rather than receiving it. Then we put this morning's decision to the model: what it can already answer, what it cannot, and exactly which missing field or missing source is the reason.
- 16:30
What production would take
Read-back to the people who will have to fund it: what is needed to run this in your own perimeter, what a pilot would prove, and what would make one premature.
What you leave with
A first ontology, written down
Objects, links, actions, and the decision it was built to serve.
An integration sketch per source
How the data would actually move, how often, and what the volume or the freshness would break first.
A data-readiness verdict per source
Usable as-is, usable after work, or not usable and why. In writing, within five business days.
The definitions that conflict
The terms two of your systems disagree about, and which meaning we would make canonical.
The gap list
The fields, the sources and the ownership questions between this ontology and one that could run in production.
The build, if you can keep it
Where an extract was shared, you get the sandbox contents and the model definitions; nothing about the day is locked to us.
A recommendation on scale
Whether a six-week pilot is the right next step, and which use case would prove the most with the least.
What it costs
- Credited in full against a pilot or your first platform year
- A price, not a rate — it fits a purchase order, not a procurement cycle
- Travel outside Belgium at cost, agreed up front
The full amount is credited against a pilot or your first year of platform. All prices exclude VAT.
Where this sits
- 30 minutesOrganisational AI in 30 MinutesFree
- One day eachThe assessment days€1,500
- Six weeksThe Six-Week Pilot€25,000 – €90,000
Above this day is a briefing that costs nothing; below it, a pilot that proves one use case on your data inside your own perimeter. This day is the step in between: the one that says whether the data underneath any of it will hold.
The same length, the same price, aimed at the organisational half of the problem — which of your processes are ready to run as Human+AI, drawn end to end, and what sovereignty, the EU AI Act, NIS2 and DORA require before they can. This day answers can our data carry AI; that day answers which work should we run this way, and what must be true first. Run back to back, both days are €2,700. The questionnaire below decides which fits, on your answers rather than ours.
What we ask before the day
The questionnaire is the preparation. It takes about four minutes, it is the same set of questions we would otherwise spend the first hour on, and the answers decide which of the two days you actually need — including the answer "neither yet".
Who runs it
Scrydon's founders and architects have built ontologies and sovereign data platforms for governments, defence organisations and regulated enterprises. Meet the team.