The Lakehouse
Stop feeding two copies of the truth. One sovereign store where tables, documents and embeddings live together — fast enough for BI, open enough for AI, yours end to end.
Open Iceberg Tables
Native support for Apache Iceberg and other open table formats — your data stays yours, queried in place with no lock-in.
Lightning OLAP
StarRocks' vectorised engine and materialised views power real-time SQL — from dashboards to agent reasoning — without data duplication.
Integrated Vector Search
Store and query embeddings alongside traditional data, making the Lakehouse instantly ready for AI.
Organisations have historically kept one store for reporting and another for everything messier, then spent years copying between them. A lakehouse is a single store that does both jobs, so your analysts and your AI are looking at the same data rather than two ageing copies of it.
Read this if you're paying for two data platforms and a pipeline between them.
The Lakehouse is the high-performance data foundation of the platform. It merges the flexibility of a data lake with the speed of a data warehouse, unifying structured tables, unstructured AI knowledge, and vector embeddings in one sovereign store — so analytics and AI run on the same data without duplication.
Built on StarRocks over open table formats such as Apache Iceberg, this sovereign data platform keeps your data yours: no proprietary lock-in, no copying data between a lake and a warehouse. StarRocks' vectorised MPP engine delivers sub-second, high-concurrency SQL — powering everything from dashboards to agent reasoning — and integrated vector search makes the same store instantly ready for AI workloads, all underpinning the Cognitive Enterprise.
Lakehouse in the Scrydon platform
One integrated, sovereign architecture. Here is where Lakehouse sits — highlighted against the full stack it works with.
The AI OS for Humans & AI Agents
Ontology & Semantic Layer, one connected model for your data, knowledge & processes
Combining the best of data lakes, data warehouses and search
AI agents, workflows & automations that execute across your systems
Integrate across A2A, MCP, legacy systems and data sources
Secure domain federation, trusted data sharing, and cross-boundary intelligence
Sovereign Foundations
Lakehouse in depth
The Lakehouse is the high-performance data foundation underpinning the Cognitive Enterprise. It is built on StarRocks — a blazing-fast, vectorised MPP query engine delivering sub-second analytics, real-time updates, and high concurrency — and queries open Apache Iceberg tables directly, merging the flexibility of a data lake with the speed of a warehouse under a single, sovereign roof.
- Open Iceberg tables: Query Apache Iceberg and other open table formats directly — your data stays yours, with no proprietary lock-in and no data movement.
- Lightning OLAP: StarRocks' vectorised engine, cost-based optimiser, and materialised views power real-time SQL — from dashboards to agent reasoning — without data duplication.
- Integrated Vector Search: Store and query embeddings alongside traditional data, making the Lakehouse instantly ready for AI workloads.
One store for tables, knowledge, and vectors
Open Iceberg tables — Apache Iceberg and other open formats keep your data portable, queried in place with no lock-in.
Lightning OLAP — StarRocks' vectorised MPP engine delivers sub-second, high-concurrency SQL — no data movement required.
Integrated vector search — Embeddings stored and queried alongside your data, ready for AI workloads.
Foundation for the Cognitive Enterprise — Underpins the Cognitive Enterprise, feeding fresh data into the model.
No copies, no lock-in, no compromise
Splitting data across a lake and a warehouse means duplication, drift, and cost — and proprietary formats trap your data. The Lakehouse unifies storage and compute on open formats inside your perimeter, so analytics and AI work on one current, sovereign copy of the truth.
Open formats are your exit, not a feature
Every data platform promises no lock-in on the way in. The test is the way out: whether another engine can read your tables tomorrow, without an export job and without asking permission. That is what open table formats settle. The Lakehouse stores your data in open table formats such as Apache Iceberg, on storage you control, so the files, their schema and their history are readable by any compatible engine — not only ours.
| Question | Proprietary warehouse | Open lakehouse |
|---|---|---|
| Who can read the files on disk? | Only the vendor's engine | Any engine that speaks the open format |
| What does leaving look like? | A migration project and an export bill | Point a different engine at the same tables |
| Can two engines share one copy? | No — copy and reconcile | Yes — the table is the contract, not the engine |
| Where does schema history live? | Inside the vendor's catalogue | In the table metadata, with the data |
| Who decides retention and residency? | Constrained by the service | You, on storage inside your perimeter |
For a sovereign deployment this is not a nicety. Sovereignty that depends on one supplier's continued goodwill is a contract, not a property. Because the data sits in open formats on your own storage — on-premises, air-gapped or in a sovereign cloud — the exit is always available, and it is the same exit for the ontology, the embeddings agents retrieve over, and the tables your dashboards read. The platform does not hold your data hostage to keep you; it has to keep earning the query.
Frequently asked questions
What is a lakehouse?+
What is it built on?+
Does it support AI and vector search?+
How does the Lakehouse relate to the Cognitive Enterprise?+
Explore the platform
Prefer to write? Email hello [at] scrydon.com and we will get back to you.