“What was our revenue?”
It sounds like a simple question. The database contains orders, invoices, shipments, customers, and amounts. An AI system can read the schema, identify the tables, write a SQL query, and return a number in seconds.
But which number?
A real case emerged during a DataTalk demo: for that customer, “revenue” did not correspond only to issued invoices. For the analysis they needed at that moment, what mattered was the goods that had been shipped. Even an apparently obvious expression such as “Central Italy” required a specific business rule.
The problem, then, was not writing SQL. It was understanding the meaning the company assigned to the question.
This changes how we should evaluate AI applied to business data. Connecting a model to a database gives it access to a structure. It does not automatically give it access to the context the organisation uses to interpret that structure.
A query can be correct and the answer can still be wrong
When discussing AI for databases, quality is often measured with a technical question: is the generated query valid?
That check is necessary, but it is not enough.
A query can be syntactically correct, use existing columns, and run without errors. Yet it may select the wrong table from two similar sources, use a technically possible but unapproved join, apply a revenue definition that differs from the one used by the management accounting team, or interpret “active customer” according to a plausible logic that does not match the company’s rules.
In all these cases, the database provides an answer. The problem is that it answers a different question from the one the user thought they had asked.
This is why, with AI, it is useful to distinguish technical correctness from semantic correctness. The first checks whether the query can work. The second checks whether it truly represents the business concept being requested.
The database describes the structure. The business adds the meaning
A SQL schema can reveal a great deal: table and column names, data types, keys, relationships, and views. It is an essential foundation for navigating the data.
But much of the information that really matters is not explicitly contained in the schema.
A column called `revenue` does not necessarily reveal whether the value includes returns, credit notes, or orders that have not yet been fulfilled. A `customers` table is not guaranteed to be the official source used by finance. A geographical field does not explain how the company divides Northern, Central, and Southern Italy. And a technically possible relationship does not necessarily mean that the join is the right one for the KPI being calculated.
Meaning emerges where technical structure meets the organisation’s rules.
And that is precisely the problem a semantic layer is designed to solve.
What is a semantic layer in practice?
A semantic layer is an abstraction layer that connects technical data to the concepts the business uses. It translates tables, columns, and relationships into shared entities, metrics, definitions, and rules, allowing people, AI-powered Business Intelligence tools, and AI systems to work from a more consistent foundation.
The term is not new. Semantic layers have existed in Business Intelligence for years. What has changed is that the arrival of AI agents and natural-language interfaces has made their value more apparent: the easier it becomes to ask questions, the more important it becomes to establish what those questions mean.
Google Cloud describes Looker’s semantic layer as a shared foundation of metrics, dimensions, definitions, and relationships that can provide AI with the context needed to interpret business logic, rather than raw data alone. IBM defines it as the layer that translates complex technical structures into more meaningful business terms.
In practical terms, a semantic layer designed to support AI can include:
- The approved definition of metrics such as revenue, margin, active customer, or churn
- Which tables and views are considered canonical and which are legacy or supporting sources
- The relationships to use between entities and the correct level of granularity for analysis
- Synonyms, company terminology, exceptions, and rules that cannot be inferred from field names
- Access restrictions and scopes that affect what a user can ask and see
This does not necessarily mean starting a major new data project. It means making explicit and usable the knowledge that would otherwise remain scattered across dashboards, documentation, code, and people.
Why AI makes context more important, not less
With a traditional report, many decisions are made in advance: an analyst decides which sources to use, how to calculate an indicator, and how to present it. The dashboard incorporates those decisions.
With natural-language data analysis, the sequence changes. The user asks a new question, often without knowing the schema and without specifying all the implicit rules.
This is the major advantage of conversational interfaces: they lower the technical barrier. But it is also why relying solely on the model’s ability to “understand” the question is not enough.
A language model is very good at recognising intent and generating plausible code. It cannot, however, know a company’s private definition if no one has provided it. If “booked order” for sales and “recognised revenue” for finance follow different rules, the model cannot select the correct definition simply by looking at column names.
The issue, then, is not choosing between AI and a semantic layer. AI makes the semantic layer even more important.
Schema, business glossary, knowledge base, and semantic layer are not the same thing
These concepts are often treated as interchangeable, but they solve different problems. The schema describes how the database is organised: tables, columns, types, keys, and technical relationships.
A business glossary clarifies the meaning of terms: what the organisation means by active customer, net revenue, closed order, or qualified lead. A knowledge base can collect documentation, exceptions, procedures, and information that helps explain how a business domain works.
The semantic layer makes part of this knowledge operational by connecting it to the data that must be queried and the metrics that must be calculated.
In practice, architectures may differ and the boundaries are not always clear-cut. The point is not to impose a single terminology. It is to ensure that the knowledge required to answer a question is not left to the model’s ability to guess.
The simplest test: ask where the definition of “revenue” lives today
Before introducing any form of AI for data analysis, you can run a very simple exercise. Ask finance, sales, and operations how an important KPI is calculated. Then ask where that definition is documented and how it is linked to the tables that produce it.
If the answer is “that person knows,” “it is inside that query,” “it depends on the report,” or “we have always done it this way,” the problem is not AI yet. The problem is that an important part of the data’s meaning is not available in a usable, governable form.
It is precisely in these situations that connecting a general-purpose model directly to the database can create a false sense of autonomy: asking the question becomes extremely easy, while the logic needed to answer it remains implicit.
Where DataTalk fits in
DataTalk is not designed as a simple interface that takes a question and turns it into SQL.
During database onboarding, DataTalk explores and describes the structure, builds knowledge the agent can use, and employs a semantic layer to connect the technical schema, company terminology, definitions, and rules needed to interpret questions. The platform also provides a knowledge base and mechanisms for adapting to the organisation’s language.
This does not mean that AI can infer every rule on its own. If a definition is specific to the organisation, it must be stated, validated, and maintained. The system’s value lies in its ability to use that context when interpreting the question, rather than forcing the user to rewrite it in every prompt.
For complex databases, the same approach is also available through the DataTalk API: the point is not only to generate SQL, but to retrieve the relevant context before producing the query.
The right question is not, “Can AI query the database?”
Today, in many cases, the answer is yes.
The more useful question is this: what does the AI know about how your company interprets that database?
Does it know which source is authoritative? Does it know what your organisation means by revenue? Does it understand the exceptions? Can it distinguish an operational table from a legacy table? Can it use the same definitions people rely on when making decisions?
If the answers still depend on the model’s ability to infer the context, the quality of the AI remains fragile even when the generated code is flawless.
The decisive step is therefore to shift the focus from querying to understanding.
Because AI can write SQL. But business meaning does not live in SQL.
If you want to determine how much context is needed to query your company’s data reliably, start with a real question.
Frequently asked questions
No. A Data Warehouse centralises, stores, and organises data over time; a semantic layer adds a business-oriented representation of metrics, entities, relationships, and rules. They can work together, and the semantic layer often uses data that is already organised in the DWH.
No. A semantic layer does not necessarily replace SQL or the query engine. Its main purpose is to establish which meaning, metric, relationship, or rule should be applied when data is queried. A system can continue to generate SQL while using more reliable context.
It can help describe a schema, suggest relationships, document tables, and identify patterns. However, it cannot know with certainty an internal rule that is not present in the data or documentation. Business-specific definitions must be validated by people who understand the domain.