SELECT and wonders whether AI will take away their craft.Where this thesis comes from
In autumn 2025 I prepared a talk called "O SQL, AI i Przyszłość Baz Danych" ("On SQL, AI and the Future of Databases"). The plan was simple: three acts, SQL, AI and the databases of the future. The longer I collected material, the more one sentence kept coming back to me, and in the end it made it onto a slide:
"SQL isn't dying, it's mutating."
I talked about this in Kraków and Bielsko-Biała at meetups with students, where the topic was simple: is it still worth writing and learning SQL? It turns out it is, because SQL is alive and well.
This is not a sentimental defense of SQL. It is an observation from the history of the language and from everyday work on Databricks, where we call AI more and more often in exactly the places where we used to write a plain SELECT.
The foundation: we say "what", not "how"
SQL is a declarative language. We write what result we want, and the engine decides how to compute it. A classic example from the slide:
✓ Works on Free Edition (on any table with name and salary columns)
SELECT name, salary
FROM employees
WHERE salary > 5000
ORDER BY salary DESC;
There are no loops, no indexes and no file read order here. Thanks to that, the optimizer can change the way the query runs while the query itself stays the same. In my opinion this is the main reason SQL has survived more than half a century of technology changes: the "what" layer is separated from the "how" layer, so almost everything underneath can be swapped out.
In the lab (29.09.2026) I ran this query on a synthetic employees table with five rows. Three came back: 7300, 6250 and 5100, in that order. The employee earning 4999 was dropped by the salary > 5000 condition, and I did not have to write a single word about how the engine should read the data.

Lab result (Databricks, 29.09.2026): out of five rows, the three with a salary above 5000 remained, sorted descending, without a single line about "how" to compute it.
A bit of history, because it helps to see the scale:
- 1970 – Edgar F. Codd publishes "A Relational Model of Data for Large Shared Data Banks".
- 1970s – IBM creates SEQUEL, later renamed SQL.
- 1986–1987 – ANSI and ISO standardize the language (SQL-86). From that point on, SQL does not belong to one company.
On top of that, it is not a niche language. In the Stack Overflow Developer Survey 2023, 51.52% of professional developers used SQL (48.66% of all respondents).
Four mutations
The most interesting part is what happened to SQL each time a technology appeared that was supposed to "kill" it.
| Era | What appeared | How SQL responded |
|---|---|---|
| Relational | tables, ACID, foreign keys | SQL as the query language for relations |
| NoSQL | documents, JSON | JSONB in PostgreSQL, JSON functions in SQL Server |
| Big Data | distributed processing, files in a data lake | Spark SQL, Presto, BigQuery, Snowflake |
| AI | embeddings, semantic search | vector types, AI functions and vector search called from SQL |
In the talk I summed it up with the sentence I like most in the whole deck: "Each era did not replace the previous one – it extended it." Relational databases still handle transactions, warehouses handle analytics, the lakehouse handles data science, and vectors handle AI. Evolution goes through specialization, and SQL stays the common language on top.
Vectors show this well. On the slide "The giants didn't wait" I showed that PostgreSQL got the pgvector extension and SQL Server 2025 got a native VECTOR type. Instead of moving data to a separate vector database, we can add an embeddings column next to the existing tables.
A correction to my own slide: the function for generating embeddings in SQL Server 2025 is called AI_GENERATE_EMBEDDINGS, not GENERATE_EMBEDDING. The demo scripts had the correct name, the slide did not.
What this mutation looks like on Databricks
In Databricks SQL, semantic search is simply another function in the FROM clause. We do not have to compute the question vector ourselves, because the index with managed embeddings does it for us.
✓ Works on Free Edition (requires an AI Search endpoint and index; Free Edition is limited to one endpoint)
SELECT *
FROM vector_search(
index => '<your_catalog>.<schema>.<index>',
query_text => 'how to brew coffee properly',
num_results => 3
);
It is worth adding that I did not run this snippet in the lab on 29.09.2026. The notebook keeps it behind the RUN_VECTOR_SEARCH switch, and the workspace did not have an AI Search endpoint or index yet, so the cell was skipped. Treat the syntax as a sketch to check against the documentation.
It is still a SELECT. What changed is what happens underneath: instead of comparing letters, we compare meanings. How exactly LIKE, full-text and vectors differ is covered in "A magnifying glass, a book and a buddy".
AI without data is an empty box
The second thesis from the talk matters more to me than the first: "AI without data is an empty box."
The best language model does not know what products we have, who our customers are, or what happened yesterday in the sales system. It can only guess. On the other hand, we have petabytes of data, neatly arranged in tables, but a classic query requires knowing the structure, and full-text search looks for words, not meanings. We cannot ask "what do our customers think about product X?" in a WHERE clause alone.
That is why I do not see competition here. RAG, agents and Genie need data that someone has modeled, cleaned and described beforehand. When I was preparing a workshop on AI agents a year later, many of the problems were not in the model but in the data: duplicated customer rows, an unclear measure definition, permissions that disappeared after a function was recreated. This is exactly the area where SQL people feel most confident.
Since the 1990s someone has predicted the death of SQL every few years. Today we have Spark SQL, and we even call AI models from SQL.
What this means in practice
A few things I do myself and recommend to people on data teams:
- Let's not drop SQL for "AI alone". If an agent or Genie generates a query, someone has to be able to read it and judge whether it computes the right thing.
- Let's learn new functions in the old language. On Databricks, we call
ai_query, AI functions and vector search from SQL. The barrier to entry is low, because the syntax is familiar. - Let's take care of data semantics. Comments on tables and columns, measure definitions, clear names. This is what Genie and agents use.
- Let's pick the method for the problem.
LIKEand regex are still the best choice when we know the pattern (a product code, an email). Vectors make sense when we are looking for intent.
What this thesis does not say
"SQL isn't dying" does not mean everything should be done in SQL. We train an ML model in Python, and we build an agent in code, not in a stored procedure. It also does not mean that SQL understands meaning by itself. The embedding model understands it, and SQL gives us a convenient interface to it. And finally: the fact that SQL has survived so far is no guarantee that it will last forever. It depends :) So far, though, every wave of novelty has ended the same way: with a new data type and new functions in a familiar language.
See it run
A short recording from the lab shows the notebook going step by step through preparing the table, the SELECT query and the skipped vector_search() step.
The full notebook is in the code/ folder (sql_mutuje.py).
Summary
To close the talk I used a sentence that sums up the whole series well:
"Structure without semantics is dead data. Semantics without structure is chaos."
- SQL survived because it separates "what" from "how", so the engine underneath can be swapped.
- Each era extended the previous one: JSON, Big Data, and now vectors and AI.
- AI needs data, and data needs people who understand it.
Next: "A magnifying glass, a book and a buddy: LIKE, full-text and vectors in plain words".
As of:
Want more posts? Follow along via RSS or on LinkedIn.



Comments
Quiet on the trail so far. Be the first to comment.