Manichandu2210/retail-data-engineering-pipeline- ? reverse-engineered prompt

Reverse engineered prompt

Build me an end to end retail data pipeline in Databricks that takes raw CSV transaction files from Amazon S3, cleans them up, and turns them into something I can trust for reporting.

I want a simple Medallion flow with Bronze, Silver, and Gold layers, where the raw data is ingested first, then standardized and cleaned, then modeled into a star schema with customer, product, location, date, and sales tables. Please handle messy dates, duplicate orders, null values, and negative quantity edge cases in a sensible way, and add some basic data checks so I can catch bad records before they reach the dashboard.

Also set it up so new data can be loaded incrementally instead of rebuilding everything every time, and make the pipeline run in the right order as an automated Databricks job. Finish it with a small SQL dashboard that shows useful retail KPIs. Use current Databricks docs online if you need to.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab