Most businesses assume their software works together. They’ve invested in APIs, maybe a middleware layer, possibly a cloud data data warehouse. On paper, the stack looks connected. In practice, the finance team is working off a different revenue number than the sales team, and nobody knows which one is right. That’s the integration problem – and it has almost nothing to do with technology.
Best-of-breed software created fragmented data by design
Moving towards best-of-breed software was a good decision back then. Instead of having one overly popular alternative trying to carry out every task, businesses chose the most suitable tool for each aspect. The outcome is that business became more efficient – but also more scattered. Marketing uses one alternative. Sales use another. Operations work on their own set, sometimes without IT even being considered during the process. Shadow IT exclusively introduced dozens of unapproved software to the majority of business systems.
On average, businesses now use 976 distinct applications, while only on average 28% of those applications are connected (MuleSoft Connectivity Benchmark Report). This disproportion is not a technical error. It’s the gathered consequences of sections solving their problems using business software that was right for them then, without considering how data integrates among these solutions.
Each new solution produces several additional data silos. This means that there will be someone somewhere who is extracting data out of a given system manually and also someone who is injecting into another system. Manual actions lead to mistakes, as well as to latency times worsening, and also to the truth being a unique affair at regular operation meetings.
Centralized platforms only work with visibility into data flow
The push toward centralized data cloud platforms addresses the fragmentation problem at the storage layer. Rather than reconciling fifteen sources at query time, everything lands in one place. But centralization isn’t the same as clarity.
A data warehouse can hold information from fifty source systems and still be opaque. If data teams don’t know how each pipeline was built, what assumptions were made during transformation, or which downstream reports depend on which upstream tables, the platform becomes a large, expensive version of the same problem. This is why governance and discovery layers matter as much as the storage infrastructure itself. Tools like snowflake Horizon address exactly this layer – providing the cataloging and access management capabilities that make a centralized environment something you can actually trust and act on.
That all sounds good if you were starting from scratch right here, right now. The reality is that most large enterprises have generational technical debt in operational systems that were designed before the world wide web, let alone APIs. These were never architected to work in an interconnected way (hence they don’t). But these core platforms are where the most critical data lives, so just “ripping them out and starting again” isn’t an option. Integration strategy for these environments isn’t about connecting everything cleanly – it’s about controlling the risk of the connections that do exist.
The hidden cost isn’t the software – it’s the manual work around it
When integrations are not successful, individuals find a way. They share CSVs. They create spreadsheets and link two systems together manually. They cut and paste data from one tool to another and cross their fingers. This is known as manual data input exhaustion, and it’s expensive in a way that is rarely documented in a software budget.
The human error factor of manual data work is difficult to measure, but it’s easy to identify – one wrong number in a pricing spreadsheet, one obsolete customer entry in a campaign list, one misclassified transaction in a report. It all adds up. Even if each mistake seems small, after months and years and many employees managing this on a daily basis, the overall impact on decision-making becomes substantial.
ETL processes and API connections provide a solution for the transfer issue. However, transferring data more quickly doesn’t resolve the problem of governance. You can automate the synchronization between two systems and still have two contradictory “active customer” definitions – one in your CRM and one in your data warehouse.
Integration is now a governance problem, not a plumbing problem
This is where most integration conversations go wrong. They focus on connectivity: can system A talk to system B? The harder question is whether anyone knows what’s happening to data once it moves. Who transformed it? What rules were applied? Which business definition is being used here versus there?
Data lineage – the ability to trace where data came from and how it changed as it moved – is what separates useful integration from dangerous integration. When data lineage is invisible, integrated systems create a false sense of reliability. Numbers match on the surface. The logic behind them doesn’t.
Metadata management is what makes lineage possible. When every data asset carries consistent metadata – ownership, classification, transformation history – you can interrogate your integrated environment instead of just trusting it. Metadata is the context that makes data meaningful regardless of which system it currently lives in.
Integration without visibility is just risk at scale
Businesses that regard integration as a governance issue instead of an IT chore are the ones who derive actionable business intelligence from their software stack. Businesses that treat integration as plumbing, eventually uncover the fact that connected systems are equally unreliable as disconnected ones. The data is transferred, but there is no insight into it on arrival.

