home / knowledge base / what is a data layer
knowledge base

What Is a Data Layer? An Honest Answer_

Data layer is one of those words that now appears in every software proposal and means something different in each one. Below is what it actually is, what it is not the same as, and when a company of twenty to two hundred people really needs one. Including the two situations where you are better off not building one.

the_definition

What a data layer is

A data layer is where your business data lives, separate from the programs you use to look at it. Not inside the CRM, not inside the accounting package and not in a folder full of spreadsheets, but in a place of its own that all those systems write to and read from. The software around it is then replaceable without your data moving along, simply because your data was never inside it.

That sounds like a database, and the storage part is indeed just a database. The difference is in the two layers that have to sit on top of it.

each layer is useless without the one below it
layer_03

Access

People, systems and AI applications can reach it, with permissions per role, without someone having to run an export first.

layer_02

Meaning

A record of what customer, order and revenue mean in your business, so every system and every model gives the same answer to the same question.

layer_01

Storage

All records in one place, in an open format, on infrastructure you own yourself.

all three layers, or it is not a data layer but a database with good intentions

The middle layer is the one that gets skipped in practice, and it is the layer where things go wrong. As long as nobody has written down that a customer in your business is someone with at least one paid invoice, the sales report gives a different number than the accounts, and an AI assistant adds a third. Not because the figures are wrong, but because three definitions sit side by side that nobody ever wrote down.

why_now

Why the word is suddenly everywhere

On the Dutch government's budget day, 15 September, digitalisation and AI were once again prominent in the King's Speech, with plans for a national investment institution and a separate organisation for breakthrough technology. That is national infrastructure. For an individual SME it changes little in the short term, because the brake sits one level lower: according to the European DESI report, around 18 percent of Dutch SMEs use big data, and roughly the same share use AI. Not because the tools are missing, but because the data those tools would have to run on is spread across four or five systems. That matches what SME owners say themselves: they use AI mostly for text, and far less for their own numbers.

At the same time there is a second movement, and it is more confusing. Software vendors have built a whole vocabulary around this problem in 2026: semantic layer, context layer, agentic data layer, each with its own product page and its own 2026 guide. The underlying observation is correct. An AI agent that does not know your pricing logic does not give a bad answer, it makes a bad decision, and Gartner expects more than 40 percent of agentic AI projects to be scrapped before the end of 2027. But in a company of twenty to two hundred people those are not three products. They are three things your one data layer has to do.

not_the_same

A data layer is not a data warehouse

This is the confusion we run into most often, and it costs money. A data warehouse is a copy to report on. A data layer is where the work happens.

data warehouse

A copy to report on

Records are loaded into it periodically so you can read figures out of it, usually about yesterday.

  • Read only, nobody works in it
  • The source systems stay in charge of their own data
  • Does not fix your fragmentation, it makes it visible
data layer

Where the work happens

Systems write into it and read from it live, so what you see is the current state.

  • Read and write, it is the source itself
  • The meaning of your terms is recorded here
  • Software around it is replaceable without data loss
they do not exclude each other, but a reporting copy does not make your data yours

A well built data warehouse is good work. It just does not answer the question this article is about: your figures become easier to read, your dependency stays. On top of that, a large part of what you see in such a copy is not stored but calculated, and that goes wrong quietly when you switch. The difference between stored and derived data is worked out in ERP migration: which data do you keep.

do_you_need_one

When you need one, and when you do not

This is where we part ways with almost every other piece on this subject, because the honest answer is that most companies do not need one. A data layer is a foundation, and a foundation without a building on top of it is just maintenance.

one system, a few edge integrations

No data layer needed. A good package does its job. Building a foundation here means buying maintenance for a problem you do not have.

two or three systems that need to know about each other

One integration is usually enough. Do not build a foundation for something an integration solves, but do write down which system is the source for what.

the same truth lives in several places

Here a data layer pays for itself. Not because it is tidier, but because nobody can say which version is right and every new plan runs aground on that.

the honest verdict: in two out of three situations a data layer is the wrong investment

There is a five minute test to work out which situation you are in. Ask three colleagues separately how many customers you have, and ask where they get that number from. If you get three numbers from three sources, you do not have a counting problem but a definition problem. The same goes for open quotes, active projects and revenue per customer. If you recognise that, the question is no longer whether you need a data layer, but when.

That pattern runs parallel to the signs that a package has been outgrown, which are set out in five signs your business has outgrown its CRM.

starting_small

How small you can start

A data layer is not a year long programme. The approach that works in practice starts with a process that hurts right now and with the data that process needs. Move those records into an open database in your own hands, leave the existing package standing next to it for the time being, and build the first solution on top. If it works, the next process follows. If it does not, you have lost a couple of weeks instead of a year.

What you should not do is start by moving everything. Moving is rarely the hard part; counting and checking is. That order is set out in the data migration plan, and how to pick the first process in automating repetitive tasks.

what_now

The foundation is the point, not the layer

A data layer is not a goal in itself. It is the reason you can switch vendors in two years without losing your history, that an AI application can reach a complete customer picture, and that a new dashboard takes days instead of months. That is the whole promise, and it is no more than that.

How we fill that in concretely, with an open data layer on PostgreSQL in your own hands that dashboards, AI chat and agents connect to, is on the page about the AI Native Data Layer. If you first want to know what automation on top of it looks like, read automating business processes with AI.

frequently_asked

Frequently asked questions about a data layer

What is a data layer?

A data layer is where your business data lives, separate from the programs you use to look at it. It does three things at once: it keeps all your records in one place in an open format, it records what terms like customer, order and revenue actually mean in your business, and it gives people, systems and AI applications access with permissions per role. If one of those three is missing, you have a database or a reporting copy, not a data layer.

What is the difference between a data layer and a data warehouse?

A data warehouse is a copy you report on: records are loaded into it periodically and you read figures out of it, usually about yesterday. A data layer is where the work itself happens: systems write into it and read from it live, so what you see is the current state. The two do not exclude each other, but a data warehouse does not fix your fragmentation. The separate systems stay in charge of their own data.

Does every SME need a data layer?

No. If your business runs on one system with a few edge integrations, a good package is fine and a data layer just buys you maintenance for a problem you do not have. With two or three systems that need to know about each other, one integration is usually enough. A data layer only pays for itself once the same truth lives in several places and nobody can say which version is right.

How do I know whether my business data is fragmented?

Ask three colleagues separately how many customers you have, and ask where they get that number from. If you get three numbers from three sources, that is not a counting problem but a definition problem: nobody ever wrote down what a customer is in your business. That is exactly the work a data layer does.

Is a data layer the same as a semantic layer or a context layer?

Those terms describe parts of the same job. A semantic layer is about what your terms mean, a context layer about the current state of a customer or order across systems. In an organisation with thousands of employees it makes sense to buy separate products for that. In a company of twenty to two hundred people they are simply two things your data layer has to do, not two invoices.

Further reading: the whitepaper

In the whitepaper on the AI Native Data Layer you can read what an open, in-house data layer looks like, what happens to your existing systems, and how to start small without turning everything upside down at once.

Download the whitepaper