The data platform vendors are industrializing how fast data reaches AI models — which makes the unresolved question of what that data means the binding constraint.
This week Cloudera and VAST Data announced a jointly delivered AI factory: a production environment where enterprise data is continuously ingested, refined, governed, and delivered to AI models for training and inference, deployable across on-premises and cloud environments. It is a credible piece of engineering aimed at a real problem — most enterprises struggle to keep AI infrastructure fed from governed sources, and this architecture takes that seriously rather than assuming it away.
Vector Infrastructure Inside the Governance Perimeter
One design choice deserves particular attention. The VAST AI Operating System integrates vector database services with NVIDIA cuVS for GPU-accelerated vector indexing and search — the storage and retrieval substrate for embeddings shipping inside the data platform rather than as a separate product alongside it. That is the right instinct: embeddings belong inside the governance perimeter, not in a shadow store outside it. The announcement describes the result as turning latent enterprise data into AI-ready data. That phrase is doing a lot of work, and chief data officers are the ones who will find out how much.
Two Definitions of AI-Ready Data
AI-ready, as the platform vendors tend to define it, means accessible, performant, and governed for access. AI-ready, as an agent actually requires it, means semantically resolved — retrieval lands on the right entities, and "customer," "account," and "exposure" mean one thing across the dozen systems the factory ingests, not twelve. Those are different standards, and the gap between them is where agentic deployments tend to stumble in practice. A pipeline that moves unresolved data faster does not produce better answers; it produces confidently wrong ones at higher utilization, with the infrastructure invoice to match. It is a pattern that comes up not infrequently in the field: an organization stands up retrieval over a well-governed lakehouse, and the first hard question exposes the same counterparty living under four names across three systems. No amount of throughput fixes that. Ontology, entity resolution, and a maintained semantic layer do — and that work is measured in quarters, not sprints, which is why it belongs on the agenda before the factory goes live rather than after.
The Layer Above the Factory
None of this is a gap in the factory; it is the layer above it, and it was never the platform's job to build. The signal in the announcement is that the platform layer is converging toward the foundations the semantic layer sits on — vector infrastructure inside the data platform at least puts embeddings within reach of the platform's governance, though whether they inherit its access policies and lineage in practice is an integration question worth verifying in any deployment. The substrate is arriving, and arriving well built. What sits on top of it — the meaning — remains the enterprise's job.
The CDOs who treat semantic readiness as a prerequisite for the factory, not a fast follow, will be the ones whose agents can be trusted with real decisions.