Invinite Data Stack

A modern data platform.
In your environment.
Under your control.

Invinite Data Stack is a lakehouse platform built on open technologies for data production, analytics and AI. The same architecture can run in your own data centre, in a private cloud, in a European cloud or as a hybrid, without tying critical data production to a single platform vendor.

Most data platforms solve a different problem

Most data platforms are designed for one cloud and one vendor's environment. They work well until the data is sensitive, the regulation changes or the price goes up. At that point there are no alternatives, because data production is tied to the platform.

Invinite Data Stack solves a different problem: how to get the capabilities of a modern cloud platform while data, infrastructure and operations remain yours to decide, and the environment can be changed later.

Control stays with you

You decide where the data is located, who operates the platform and who develops it. The supplier can be changed without rebuilding the data or the data production.

The same architecture in every environment

Own data centre, private cloud, European cloud or public cloud: no functionality is lost, and changing the environment means moving a workload, not the whole platform.

All of data production, not only storage

From sources to reporting, sharing and AI on the same data foundation, with security and continuity built into the structure.

Open standards, replaceable components

Data in an open format in object storage and a catalogue behind an open interface, so the compute engine can be changed while the data stays where it is.

Sovereignty is more than where the data is located.

A sovereign data platform is one where the organisation itself decides where the data is located, on which infrastructure and technologies it is processed, who operates the platform and who may develop it. A data centre in the EU does not by itself mean that control is yours. In Invinite Data Stack, control covers five things: data, infrastructure, technology, operations and the supplier. Each of them is your decision.

Data

You decide where the data is located. Raw data can be kept entirely in your own environment, different use cases can be given different levels of protection, and the data stays in open formats that other tools can read.

Infrastructure

The platform is installed in your own data centre, in a private cloud, in a European cloud or in a public cloud, and where needed in an environment isolated from the network (air-gapped). The environment can be changed later without rebuilding data production.

Technology

The platform is based on open standards and replaceable components. The compute engine is chosen by use case, and a single technology can be replaced without renewing the whole platform.

Operations

Infrastructure is described as code, changes go through version control and automated testing and deployment (CI/CD), and identity, logs, monitoring and secrets management are part of the structure. The platform is operated by Invinite, by your own team or by both together.

Supplier

The infrastructure definitions and the data production implementation are yours. Development can be continued by Invinite, by your own team or by another partner, and the use of your data does not depend on one vendor's closed environment.

The same architecture in your own data centre, in a private cloud and in a public cloud

Invinite Data Stack works the same way wherever it is installed: in your own data centre, in a private cloud, in a European cloud, in a public cloud or in a combination of these. No functionality is lost in any environment.

This is why the environment can change without the foundation of data production having to be designed again. When regulation, prices or vendors change, you move the workload, not the whole platform.

A complete data production platform

Invinite Data Stack covers data production from sources to use. It is not a storage solution.

  1. Data ingestion. From interfaces, databases, files and event streams with repeatable, configurable loads.
  2. Open storage. The lakehouse model: data in an open format in object storage, with compute on top.
  3. Processing and modelling. Distributed or single-machine compute by use case, for example with Apache Spark.
  4. Governance. Catalogue, metadata, identity and access rights in one place.
  5. Security. Sized for high-risk systems: identity, access rights, logs, network isolation and monitoring are built into the structure along the whole chain, not added on.
  6. Queries and analysis. In SQL and notebooks, directly on the open data.
  7. Reporting. Dashboards and self-service analytics.
  8. Sharing. Interfaces and open sharing models for partners without moving raw data.
  9. AI. Data science, model lifecycle and open interfaces for AI tools on the same data foundation.

Architecture: open standards, replaceable components

The platform is built in layers, and every component in them can be replaced. What stays is the open standard between the layers.

LayerWhat it doesOpen standard
StorageTables and files in object storageApache Iceberg, Parquet
Catalogue and governanceTable metadata, access rights and versionsIceberg REST catalogue interface
ProcessingDistributed and single-machine computePython, SQL
OrchestrationScheduling and dependenciesreplaceable component
Queries and analyticsSQL queries and notebooksSQL
ReportingDashboards and self-service analyticsreplaceable component
IdentitySign-in and access rightsOIDC, SAML
InfrastructureEnvironment definition and deploymentinfrastructure as code (IaC)
MonitoringLogs, metrics and alertsreplaceable component
AI toolsInterfaces for applications and agentsopen interfaces

The choice of components and the complete reference architecture are covered in the architecture review.

What openness means in practice

An open-source licence does not by itself give room to move. Room to move comes from four things: the data is in open formats that other engines can read; the interfaces are open, so a tool can be replaced; the data production logic is Python and SQL, which travel with you; and the platform can be operated without one vendor's control plane.

When these four are in place, replacing a single component is a bounded project. The whole platform does not need to be renewed.

The platform decides where data lives. Data as Software decides how it is produced.

Data as Software is Invinite's method in which data production is done the way software is developed: as code, in version control and tested. In Invinite Data Stack this means that data pipelines and data products are in Git version control, changes are reviewed before production, testing and deployment are automated, and data contracts state what each data product promises. Tests run on four levels: unit, integration, acceptance and data quality. Every change is traceable, and releases are controlled.

This makes data production as portable as the platform: what is built for you travels with you as code.

Data as Software: the method →

Security and continuity are built into the structure

These are part of every installation. They are described as code and checked before the platform is taken into critical use.

Identity and access management

Secrets management

Traceability and logs

Network isolation

An option isolated from the network

Separate environments for development, testing and production

Monitoring and alerting

Backup and recovery

Automated enforcement of usage rules

Controlled release and rollback to the previous version

The platform supports meeting regulatory requirements. It does not by itself make an organisation meet them: data protection and security work, risk management and the processes regulation requires remain the organisation's own. The platform makes them verifiable.

Raw data in your own environment, compute where it is useful

All of data production, compute included, can run in Invinite Data Stack in the location you choose: your own data centre, a private cloud or a public cloud. The hybrid model is an option when part of the analytics and AI is wanted on your current cloud platform or in public cloud services. It then proceeds in three steps.

Controlled environment.

Raw data is ingested, pseudonymised and modelled in your own environment. The result is a data product without direct identifiers.

Selected cloud services.

The pseudonymised data product is taken to the public cloud or to your current cloud platform for scalable analytics and AI.

Results back.

Models, results and aggregates are returned to your own environment in a controlled way and shared from there.

Pseudonymised data is still personal data for whoever holds the key, so the structure alone does not decide what may be taken to the cloud. It makes data protection work verifiable.

Who the platform is built for

Invinite Data Stack suits an organisation where one or more of these holds:

  • the data is sensitive or critical to operations;
  • you want to keep your own ability to operate;
  • the cost of the current cloud platform or dependence on one vendor is a concern;
  • AI needs a reliable data foundation that machines can use;
  • preparedness requires alternative ways to run;
  • the same data has to be used in several environments.

Sectors: health and social care, public administration, security, critical industry and research.

Replaces selected layers or works alongside your current cloud platform

Invinite Data Stack does not have to be chosen in place of everything else. It can replace selected data production layers of your current cloud platform, for example raw data processing, or work alongside it in the hybrid model.

A comparison at category level, as of October 2026.

Managed cloud platformEuropean sovereign cloud platformInvinite Data Stack
Own data centreTypically noTypically noYes
Private cloudLimitedService by serviceYes
Public cloudYesYesYes
Environment isolated from the networkNoRarelyPossible
Open formatsVariesOftenThe starting point
Customer-controlled infrastructure as codeVariesVariesYes
Replaceability of componentsLimitedLimitedA design principle
Operation by your own teamLimitedRarelyPossible

Databricks and Snowflake are examples of managed cloud platforms alongside which Invinite Data Stack can be used.

Evidence

What the platform is based on

Invinite Data Stack has been productised from implementations in which open data platform technologies and Data as Software were taken into the customer's own environment in a regulated healthcare data environment. We will describe the customer-specific details once the customer's permission is in place.

In addition: a lakehouse and Data as Software at Western Uusimaa Wellbeing Services County in the customer's own public cloud environment; three hybrid cloud and sovereignty studies; the [production-readiness checklist](#security).

How the platform is taken into use

Architecture review.

We go through your current data platform, data, security and environments. You get a description of the current state, the critical dependencies and one realistic path.

Proof of capability or pilot.

One bounded domain or use case is taken through from source to use. You get a working data product and an estimate for taking it into production.

Production platform.

The platform is set up as infrastructure as code: integrations, identity, monitoring, and automated testing and deployment. You get a platform whose definitions are yours.

Operation and handover.

Invinite operates, you operate or we do it together. Skills and the implementation transfer to you during the work.

The platform is bought as a project or as a continuous allocation. The review is a bounded project.

Frequently asked questions

What does a sovereign data platform mean?

A data platform where the organisation holds control in five respects: where the data is located, on which infrastructure the platform runs, which technologies it is based on, who operates it and who may develop it. A data centre in the EU does not meet this if the platform's control plane is in one vendor's hands.

Is Invinite Data Stack a cloud service?

No. It is a platform installed in the environment you choose: your own data centre, a private cloud, a public cloud or a combination of these. It is not bought as a SaaS service from a single cloud vendor.

Can it be installed in our own data centre?

Yes. An own data centre is one of the platform's usual environments, and the functionality is the same as in the cloud. The platform is set up as infrastructure as code, so the installation is repeatable.

Show all questions (12)

Can it operate in an environment isolated from the network?

Yes. The part in your own environment can be isolated from the network entirely (air-gapped). This requires that package distribution, updates and monitoring are arranged on the isolated environment's terms. This is planned in the review.

Does it replace Databricks or Snowflake?

It can replace selected data production layers of theirs, for example raw data processing and storage, but it does not force this. Many organisations keep their current cloud platform for analytics and AI and move raw data processing to their own environment.

Can it be used alongside Databricks or Snowflake?

Yes. In the hybrid model, raw data is processed and pseudonymised in Invinite Data Stack in your own environment, and the pseudonymised data product is taken to your current cloud platform. The results are returned in a controlled way.

Which open standards is the platform based on?

Tables are stored in the Apache Iceberg format and files in the Parquet format, the catalogue works through the open Iceberg REST interface, data production is written in Python and SQL, and the infrastructure is described as code. The components are established open-source projects, and we go through them in the architecture review.

In what format is data stored?

In Apache Iceberg tables and Parquet files in object storage. Both are open formats that several engines and tools can read, so the data is not tied to the platform.

Who operates the platform?

You choose: Invinite operates, your own team operates or we do it together. Because the infrastructure is code and the documentation is yours, operation can be transferred step by step to your own team or to another partner.

How are updates carried out?

Changes are made to the infrastructure code, reviewed and taken into production with a controlled release. If something goes wrong, the previous version is restored with the same mechanism. In an environment isolated from the network, updates are brought in separately in an agreed way.

How does the platform support the requirements of the GDPR, NIS2, the EHDS and the Finnish Secondary Use Act?

The structure supports accountability: raw data can be kept in your own environment, use is traceable and access is controlled. Data released under the Secondary Use Act is processed in an audited secure operating environment, and the EHDS, the European Health Data Space, brings its own requirements for the secondary use of data. The platform does not by itself make an organisation compliant, but it makes data protection and security work verifiable.

How does Invinite Data Stack relate to Data as Software?

The platform decides where the data is located and with which technologies it is processed. Data as Software decides how it is produced: as code, in version control, tested and described with data contracts. The platform is delivered with the method, and together they make data production portable.

What would a sovereign data platform look like in your environment?

The architecture review is a bounded review of the kind we have carried out for several organisations. We go through your current data platform, the critical dependencies and one realistic path to a hybrid or sovereign environment. In the review we also show the reference architecture with its components.

Related pages: Data as Software · Definition