Divy-dev/retail-sales-etl ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python PySpark ETL pipeline for retail sales data.
It should read a raw CSV of sales records, check the data for missing values and duplicates, clean it up, and save a processed version as Parquet. I want it to cast the important fields to the right types, fill missing quantity values, calculate total amount for each order, and add useful date fields like month and year. Please also classify orders by size and add a simple product label based on revenue.
After that, generate a few business reports like sales by category, product, and city, plus basic totals like revenue, orders, quantity sold, average order value, top product, and top category. I’d also like examples of window functions, Spark SQL, and a custom UDF so the project shows a few common data engineering techniques. Save all outputs in clear Parquet folders under a processed and analytics area, and make sure I can run the whole thing from main.py. If anything is unclear, look up current docs online if you need to.
Are you gonna build this?
make sure you review the code using coderabbit