By Dr. Guendalina CALDARINI, Program Director, MSc in Data Management | Professor at aivancity |
AI education today faces a paradox. Schools and boot camps are training graduates to refine language models, deploy generation pipelines augmented by retrieval, and build computer vision systems. Yet when asked who is responsible for the quality, consistency, and governance of the data on which these models depend, the answers quickly become hesitant. Data management—which ensures that AI systems are reliable in production—remains the least represented area in AI curricula.
How the Cycle of Enthusiasm Shaped Education
The boom in AI training in the 2020s closely followed the hype cycle. As transformer models became accessible, training programs quickly began producing prompt engineering specialists. The arrival of large language models in enterprise software then caused a surge in demand for ML engineers capable of fine-tuning and deploying them. Training organizations therefore adopted a core curriculum that has since become standard: Python, PyTorch, LangChain, vector databases, and model evaluation. This choice made sense, as these skills are genuinely useful.
This shift has also created a structural blind spot. A model is easy to showcase and highlight in program communications. The work involved in processing the data that feeds the model attracts less attention and remains difficult to incorporate into a 12-week curriculum. Yet, in production environments, data issues derail far more AI projects than model issues do.
What Data Management Entails
Data management is often reduced to simply “storing data.” In reality, it encompasses several interdependent areas: data quality management (ensuring completeness, consistency, and accuracy throughout the data lifecycle); metadata management (describing data so that it can be found, understood, and governed); ontology and taxonomy management (building structured representations of context); entity resolution and data alignment (identifying when different records refer to the same real-world entity); and data governance (defining rules for ownership, accountability, and lifecycle management).
Together, these fields form the foundation of reliable AI systems. They also constitute a coherent discipline—one that is still rarely taught—at the intersection of computer science, information science, and management programs.
Why AI Makes Data Management More Critical
The shift to large-scale AI has exacerbated data management problems rather than solving them. Language models trained on poorly prepared data replicate those errors at scale. Retrieval-augmented generation systems are only as reliable as the knowledge bases they query. Agent-based AI systems, which act autonomously, introduce new categories of risk when the underlying data is inconsistent, undocumented, or ungoverned.
“In production environments, data issues derail far more AI projects than model issues. Yet most AI curricula devote only a tiny fraction of their instruction time to the discipline that would help prevent them.”
The European AI Act has translated this reality into legal obligations. Article 10 of the regulation establishes binding data governance requirements for high-risk AI systems. These requirements pertain to the quality, relevance, representativeness, and documentation of data. Compliance is therefore no longer merely a matter of best practices. It is a legal requirement that will apply to a significant portion of the AI systems deployed in Europe. However, most AI graduates today have little or no training on what these requirements actually entail.
The Components of a Data Management Program
A comprehensive data management training program is structured around four interrelated areas of expertise. Technical skills include data modeling, SQL and graph query languages, ETL design, metadata standards (Dublin Core, DCAT, OWL), and data quality frameworks. Governance skills focus on GDPR compliance, stewardship models, data catalogs, audit trails, and the requirements of the European AI Act. Semantic skills include ontology design, controlled vocabularies, knowledge graph construction, and entity resolution techniques. Finally, strategic skills focus on product thinking applied to data and aligning data architecture with business objectives.
Data management is a cross-disciplinary field at the intersection of computer science, information science, law, and organizational design. This diversity explains why it is difficult to integrate into existing curricula and why dedicated programs are necessary.
The Consequences of This Blind Spot
Every year, companies invest heavily in AI systems whose performance is limited by inconsistent, undocumented, or siloed data—rather than by the models themselves. Industry analyses consistently identify data quality as one of the main causes of AI project failure. Demand for professionals in data governance and management is rising sharply due to regulatory pressure, the enterprise-wide deployment of AI, and the proliferation of agent-based systems. However, the supply remains limited, as training has not yet caught up with market needs.
Training machine learning engineers without also preparing data management specialists can compromise the reliability of the systems they build. As AI moves from pilot projects to systems that organizations rely on every day, data management will take center stage in training. Programs that incorporate it now will produce graduates capable of building reliable, large-scale AI systems that comply with the regulatory framework governing their deployment.
Learn more
Learn about the MSc in Data Management program at aivancity and our work on data governance in the era of the European AI Act.
