When an AI project quietly dies, the post-mortem tends to blame the model. It hallucinated, it wasn’t accurate enough, the vendor overpromised. Sometimes that’s true. More often, if you trace the failure back far enough, you end up somewhere far less exciting: a nightly job that silently dropped a third of the rows, a customer table with four spellings of the same company, or a “single source of truth” that turned out to be three sources and a spreadsheet on someone’s laptop.
Models have become a commodity remarkably quickly. Any competent team can call a capable large language model through an API in an afternoon. What they can’t do in an afternoon is give that model trustworthy, current, well-described data to work with. That part still takes proper engineering, and rather a lot of it.
Where it goes wrong
Data engineering failures rarely announce themselves. A pipeline that crashes is almost a gift, because somebody notices. The dangerous ones keep running and produce slightly wrong answers for months, and catching those is most of what the teams that build the pipelines AI depends on spend their days doing.
Take a retrieval-based assistant built to answer questions about product documentation. It works beautifully in testing. In production, it starts citing a pricing policy that was withdrawn last spring, because the ingestion job indexed the archive folder along with the live one and nobody set up a way to expire old content. The model did exactly what it was asked. The data told it something untrue.
Or take a forecasting model trained on sales history. The source system changed how it recorded returns partway through the year, and nobody downstream was told. The model learned that returns fell off a cliff in June, which will come as news to the returns team.
These aren’t exotic problems. They’re the ordinary consequences of treating data movement as plumbing that someone will sort out later. For traditional reporting, a bit of mess was tolerable, because a human analyst looked at the dashboard and could spot that a number looked off. AI systems consume data at a scale and speed where nobody is looking at each record, so errors propagate quietly and get presented with great confidence.
The basics that most failed projects skipped are not complicated:
- Data contracts or at least agreed schemas between the teams producing data and the teams consuming it, so a change upstream doesn’t break things downstream without warning.
- Automated quality checks on freshness, volume and null rates, run as part of the pipeline rather than as an afterthought.
- Lineage, so that when an AI answer looks wrong you can trace which source and which transformation produced the data behind it.
- Clear ownership, meaning a named person who’s responsible for each important dataset.
None of this is glamorous. All of it is cheaper than explaining to the board why the customer-facing assistant told a client something that stopped being true a year ago.
Lakehouses, Fabric and Databricks
The platform conversation has settled, broadly, on the lakehouse: data stored in open formats in cheap object storage, with warehouse-style tables, governance and query engines layered on top. It avoids the old split between a data lake that nobody could find anything in and a warehouse that was expensive and rigid. Both of the main options Microsoft-centric organisations weigh up follow this pattern.
Databricks, available on Azure as a first-party service, has been at this longest. It’s strong on large-scale engineering and machine learning workloads, and its Unity Catalog gives a mature way to govern tables, permissions and lineage. It tends to suit organisations with substantial engineering teams who are happy working in code.
Microsoft Fabric, generally available since late 2023, takes a different angle. It bundles data engineering, warehousing, real-time analytics and Power BI into a single SaaS product built on OneLake, with capacity-based pricing. For businesses already living in Microsoft 365 and Power BI, the appeal is obvious: fewer moving parts, one bill, familiar tooling. The trade-off is less fine-grained control, and some teams find the capacity model harder to predict than they’d hoped.
Both store data in Delta format, and Fabric can read Databricks-managed tables through shortcuts without copying them, so the choice is less binary than vendor marketing sometimes suggests. Plenty of organisations end up running Databricks for heavy engineering and Fabric for business-facing analytics, which is fine as long as someone has decided where the authoritative copy of each dataset lives.
What neither platform does is fix bad data for you. A lakehouse full of untested pipelines is just a more modern place to keep the same problems. The organisations getting real value from AI tend to be the ones that invested in boring engineering first: testing, monitoring and documentation, applied consistently.
There’s one gap the tooling hasn’t closed yet. Unstructured data, such as contracts, emails and support transcripts, now matters as much to AI projects as tidy tables do, and most organisations have far weaker quality controls for it. Nobody has quite worked out what a data quality check for a PDF should look like.

