Lakehouse for AI: Avoid a Second Vector Store, Be ADMP Audit Ready

A lakehouse is the right foundation for AI when we need one authoritative copy of data that serves analytics and machine learning with ACID guarantees, time travel and unified governance, built on open table formats over object storage. It suits retrieval‑augmented generation, recommendations and document AI, where embeddings and tabular data must stay consistent. A separate vector service still wins when latency requirements are extreme and the dataset is small enough that synchronisation overhead is trivial.
TL;DR:
Choosing an open table format like Iceberg, Delta Lake, or Hudi enables ACID guarantees, schema enforcement, and time travel for AI workloads on object storage.
Recipes for AI, like retrieval-augmented generation or document AI, benefit from a structured medallion pattern where raw, cleaned, and curated data are stored in a unified lakehouse.
Integrating vectors directly into Parquet files using the AI-Lake approach maintains transactional integrity and simplifies governance compared to separate vector databases.
Ensuring delete consistency, role-based access, and audit trails across the entire data stack is crucial for compliance with regulations like Malaysia’s ADMP.
Successful production deployment requires small, controlled pilots focusing on data governance, correct workflows, and load testing, with specialist help recommended for platform design and automation.
Table of Contents
Architecture overview: mapping the medallion pattern to AI artefacts
Storing and serving embeddings: unified formats vs separate vector services
How training, feature stores and MLOps work with a lakehouse
Governance and data protection: operational controls and ADMP implications
Author perspective: common mistakes and when to call a specialist
Architecture overview: mapping the medallion pattern to AI artefacts
The medallion pattern gives AI platforms a predictable shape. Bronze holds raw ingested data exactly as it arrived, silver cleans and conforms it, and gold exposes curated, business‑ready tables. AI artefacts fit naturally into this structure rather than sitting beside it in a separate system.
Bronze stores raw documents, logs and streaming events before any transformation.
Silver holds deduplicated, schema‑enforced records, including extracted text and intermediate feature calculations.
Gold carries curated tables used for retrieval, training and serving, including embeddings and model snapshots versioned alongside the data that produced them.
A metadata and catalogue layer sits across all three, recording schema changes, access grants and lineage. This control plane is what gives compliance teams a single audit trail instead of three disconnected ones. According to Analyst Engineering’s comparison, the practical benefit of a lakehouse is that many engines, Spark for batch, a streaming engine for near‑real‑time features, and a SQL engine for BI, can read the same open tables without a copy step. That removes the duplication that typically causes drift between an analytics warehouse and a separate ML feature store.
Core technologies and open table formats to evaluate
Choosing a table format is the architectural decision that shapes everything downstream, more than any single compute engine choice. Analyst Engineering notes that Iceberg, Delta Lake and Hudi each add ACID transactions, schema enforcement and time travel to files sitting on object storage, but they differ in operational surface and ecosystem maturity.
Apache Iceberg suits teams wanting strong multi‑engine compatibility and granular snapshot control across Spark, Trino and DuckDB.
Delta Lake fits teams already committed to a Spark‑centric pipeline and wanting tightly integrated streaming support.
Apache Hudi favours workloads with frequent upserts and near‑real‑time ingestion, such as change data capture feeds.
Object storage decisions follow a similar logic: lifecycle policies that tier cold data, encryption at rest, and access control enforced through the catalogue rather than bucket‑level policies scattered across teams. Metadata and catalogue patterns, whether a Nessie‑style versioned catalogue or a unified catalogue concept that tracks tables, permissions and lineage centrally, matter because they turn a pile of files into something an auditor can actually interrogate. We find that teams comparing engines against this backdrop often benefit from a structured platform decision guide before committing to a stack.
Storing and serving embeddings: unified formats vs separate vector services
Vector search has traditionally meant standing up a dedicated database alongside the lakehouse, which introduces a second system that must stay synchronised with the source of truth. The AI‑Lake project proposes an alternative: appending vectors and an HNSW index directly to Parquet files, with centroid and radius metadata kept in Iceberg manifests so files can be pruned without reading them in full. This keeps tabular and vector data under the same transactional boundary, which means a delete is atomic and a time‑travel query returns vectors exactly as they existed at that version.
Unified formats avoid the sync failure modes common when two systems hold the same entity under different consistency models.
Separate vector databases still offer simpler operational tooling and managed control planes for teams without lakehouse experience.
Latency on a dedicated vector service can be lower for very high query‑per‑second workloads, at the cost of a second governance surface.
Pro Tip: If you already have significant Iceberg investment and need atomic deletes for compliance reasons, extend that investment rather than bolting on a second vector store.
For a broader comparison of where vector infrastructure fits across search, RAG and fraud detection, our vector database use cases guide walks through the trade‑offs in more detail.

How training, feature stores and MLOps work with a lakehouse
Training on a lakehouse means training against a table snapshot, not a loosely defined extract. That snapshot becomes the dataset’s version of record, so a model’s lineage can point to the exact rows, schema and timestamp used.
Feature engineering reads and writes through the same gold tables that serve BI, keeping training and serving features consistent.
Model CI/CD pipelines tie dataset snapshot identifiers to model artefacts, so retraining triggers (scheduled or streaming‑driven) produce a reproducible provenance chain.
Serving splits into batch scoring against gold tables and near‑real‑time joins for use cases needing fresher features.
Drift monitoring compares live feature distributions against the training snapshot, with rollback achieved by reverting to a prior table version rather than re‑engineering a pipeline under pressure.
This is the part of the stack where the discipline, not the tooling, decides whether a platform stays reliable in production.
Governance and data protection: operational controls and ADMP implications
Malaysia’s Automated Decision‑Making and Profiling guideline requires informing data subjects when automated decisions affect them, and it explicitly warns against relying solely on AI for a decision. The guideline recommends designating reviewers for AI used in automated decision‑making, trained in risk assessment and compliance oversight, which have direct architectural consequences.
Per‑tenant isolation at the table and catalogue level, so a reviewer can audit exactly which data fed a given decision.
Atomic delete behaviour, inherited from the table format, so a data subject’s deletion request removes both tabular and vector representations in one transaction.
Access audit trails tied to the catalogue, giving reviewers and auditors a single log rather than scattered application logs.
Designated human reviewers are a named recommendation, not an afterthought, under the ADMP guideline, which means reviewer training and documented purpose statements belong in the runbook, not the appendix. Our AI security controls guide covers tenant isolation and permission design in more depth, and the AI governance framework from Mindpod offers a complementary enterprise‑risk perspective.
Representative production use cases and deployment patterns
Three patterns recur across most lakehouse‑for‑AI deployments we encounter.
Retrieval‑augmented generation: a curated gold layer feeds the embedding refresh pipeline on a defined cadence, with the serving layer querying the same table version used for retrieval to avoid stale context.
Recommendations: features join on a near‑real‑time schedule against gold tables, with evaluation metrics tracked against the same snapshot used for the most recent retraining cycle.
Document AI: ingestion and OCR land in bronze, extraction populates governed silver tables, and a vectorised QA pipeline, similar to the pattern behind our intelligent document processing work, serves answers from gold.
Stepwise checklist for pilots and rolling to production
Inventory existing data assets, assign owners and bring in governance sponsors, including a data protection officer where ADMP‑style decisions are in scope.
Scope a pilot around one medallion‑layered dataset, one embedding pipeline and one reviewer workflow, kept deliberately small.
Verify delete consistency, role‑based access and audit completeness before expanding the pilot’s data volume.
Load‑test the catalogue and compute engines against realistic concurrency before declaring production readiness.
Pro Tip: Treat the pilot’s audit trail as the actual deliverable, not the model, because a reviewer who cannot reconstruct a decision will stall the rollout regardless of accuracy.
Author perspective: common mistakes and when to call a specialist
The recurring failure we see is not a bad table format choice. It is two teams, one owning the warehouse and one owning the vector store, drifting apart until nobody can explain which copy of a customer record is authoritative. Governance vacuums follow close behind. Specialist help earns its cost in platform design and MLOps automation, not in picking a brand. A realistic pilot‑to‑production timeline runs over a moderate period once reviewer workflows are included.
— Thomas Samuel
How we help you operationalise a lakehouse for AI
We help design and build lakehouse platforms that carry AI workloads from strategy through to daily operation, aiming to maintain project continuity. Our Data & Platform Engineering service covers table format selection and catalogue design, our Deployment & MLOps service handles reproducible training and serving pipelines, and Managed AI Operations keeps the platform running once it is live.

If you are assessing whether your current data estate is ready for this architecture, our Readiness & Data Diligence engagement is the practical starting point, or browse the full services catalogue for strategy, build and run options.
FAQ
What is a lakehouse?
A lakehouse combines object storage with an open table format, such as Iceberg, Delta Lake or Hudi, to add ACID transactions, schema enforcement and time travel to data that previously lived only as loose files. It lets analytics, streaming and machine learning read the same tables rather than maintaining separate copies, as Analyst Engineering explains.
Is Lakehouse the same as Databricks?
No, a lakehouse is an architectural pattern, not a single product; Databricks is one vendor that implements it, alongside other engines that can read the same open table formats. The distinction matters when choosing tools, since the underlying table format commitment, not the engine brand, determines portability.
Why choose a lakehouse over a data warehouse?
A warehouse typically owns its storage and offers a polished SQL experience, while a lakehouse keeps data as open files on object storage and adds table formats for the same guarantees, which suits workloads that span analytics and AI rather than SQL alone. Analyst Engineering frames this as a trade‑off between workload plurality and vendor lock‑in tolerance.
Is Snowflake a lakehouse?
Snowflake began as a managed warehouse and has added support for open table formats, which brings it closer to lakehouse characteristics, but the comparison depends on which features a given workload actually uses. Readers evaluating this trade‑off in practice may find our platform decision guide useful for weighing engine‑specific differences.
Sources
Recommended