What The Medallion Architecture Actually Is
The Medallion is a data lakehouse design pattern. It organizes raw and processed data into three concentric layers: bronze, silver, and gold. You drop raw ingestion into bronze, clean and transform it in silver, then build curated business entities in gold. That is the textbook version. The reality on the ground is a lot messier. I set this up for a logistics company a couple of years ago and spent about three weeks arguing with people about whether their source data could go straight into silver. It can't. The bronze layer exists for a reason. I have seen teams skip it and then lose the ability to reprocess downstream when they realized the transformation logic was wrong. Once the raw stream is gone, you cannot undo bad assumptions. The bronze layer stores raw events exactly as they arrive. JSON blobs, CSV dumps, Kafka snapshots, whatever. No schema enforcement, no transformations, no deduplication. You record the ingestion timestamp and the source identifier so you can trace every record back later. Silver applies schema conformance, handles deduplication, resolves slowly changing dimensions, and enriches records with lookup tables. Gold folds silver into fact and dimension tables that analytics tools can consume without questioning anything.
How to Set It Up Properly
Start by picking your storage layer. Delta Lake, Apache Iceberg, or Apache Hudi all work. I prefer Delta Lake for this because the time travel feature lets you query what bronze looked like yesterday without running a separate snapshot job. That saved me during an incident when a vendor changed their API response format without documentation. We restored bronze, reran the silver pipeline, and caught the breakage in under an hour. Next, define the schema at the silver boundary. This is where most people make mistakes. They let silver schemas drift, which means gold tables quietly degrade over time. Enforce the schema at the bronze-to-silver transition with a fail-fast validation step. If a record does not match the expected structure, route it to a quarantine partition instead of silently dropping it or forcing a cast. I found this approach necessary after a payment processing service started sending null values in fields that had been non-nullable for two years. A silent cast would have corrupted the revenue metrics. The quarantine partition caught it immediately. For the gold layer, build out star or snowflake schemas depending on your query engine and the complexity of the relationships. Keep the grain consistent. A common failure mode is mixing order-level and line-item-level data in the same gold table. It works until someone writes a query that double-counts revenue across joins.
Common Pitfalls
The biggest one is treating bronze as a permanent archive. It is not. Bronze should be retained only as long as you need it for reprocessing. Storage is cheap enough that people keep everything forever, which then creates confusion about which dataset is the source of truth. I established a policy of keeping bronze at monthly granularity for six months, then archiving to cold storage with a manifest file listing what was included. That kept the active workspace clean without losing the ability to replay. Another issue is the overuse of silver. Some teams create multiple silver zones for different domains, which fragments the pipeline and makes downstream orchestration complicated. A single silver layer with clearly partitioned schemas is easier to maintain. If you find yourself creating silver_v2 or silver_enriched, stop and redesign the schema instead. The medallion also does not solve data quality problems. It isolates them. A broken enrichment job in silver will still produce valid but meaningless gold tables. You need monitoring at each layer, not just at the end.
Get the Full Details
![The Medallion [Blu-ray]: Amazon.ca: John Rhys-Davies, Anthony Carpio, Jackie Chan, Lee Evans ...](https://m.media-amazon.com/images/I/819lbmfYraL._AC_SL1500_.jpg)
When It Fails
The Medallion architecture breaks down when your data is primarily write-heavy with very low query volume. Building bronze, silver, and gold layers for a system that processes a few hundred rows per day and gets queried twice a month is overengineering. In those cases, a simple two-tier setup with raw and processed tables is sufficient. The overhead of managing three layers with schema enforcement, quarantine routing, and time travel tracking becomes a burden rather than a benefit. It also does not fit well with streaming-only workloads that require sub-second latency. The medallion pattern assumes batch-oriented processing between layers. If your use case demands real-time dashboards fed directly from a Kafka stream, a lambda or kappa architecture serves you better. You can still borrow the concept of separating raw from enriched, but the three-layer model is not the right fit.
A Practical Workaround I Used
At one point, a downstream team wanted to query bronze directly because they needed fields we had not yet decided to include in silver. The correct answer was to add the fields to the silver schema, but they were under time pressure. Instead of refusing or allowing direct bronze queries that would bypass governance, I created a shadow view on bronze that applied only the columns they needed, registered it in the same catalog as silver, and gave them read access. This kept the governance boundary intact while giving them what they needed. The shadow view lasted about a month before we promoted the columns into silver properly. Databricks has built-in medallion support through delta live tables. Spark Structured Streaming works well if you are already in the Spark ecosystem. For smaller teams, Snowflake with its native time travel and zero-copy cloning features handles the pattern without requiring a separate compute layer. dbt can manage the silver-to-gold transformations reliably. Airflow or Prefect handles orchestration. The specific tools matter less than getting the layer boundaries right. If you want something you can clone and adapt, the GitHub repository for the medallion architecture reference implementation from Databricks gives you a reasonable starting point. It includes sample pipelines for each layer and schema validation templates. I used it as a baseline and modified it heavily for our use case, but the core structure was sound.
Bottom Line
The Medallion gives you a clear separation between raw data and business-ready data. It makes reprocessing possible. It keeps analytics teams from querying things they should not be querying. It does not make your data better, and it does not replace monitoring or schema governance. Build it if your data volume and complexity justify it. Skip it if they do not.
