# Data cleaning and preparation — sample project

A deliberately messy order export turned into a clean, analysis-ready table, with a log of every change.
**All data is synthetic.** No real company or person is involved.

## The problem
The raw export has duplicate rows, three date formats, inconsistent names (`Line A`, `line a`, `LINE-A`, `A`),
weights in kg or g with decimal commas, missing values and impossible quantities.

## What the script does
1. Removes exact duplicates.
2. Trims spaces and standardises customer, material and line names.
3. Converts all dates to `YYYY-MM-DD`.
4. Converts weights to kg (values without a unit are assumed to be kg).
5. Rejects rows with a missing, negative or implausible quantity — they go to `output/rejected_rows.csv` instead of disappearing.
6. Writes `output/clean_orders.csv` and a `output/cleaning_log.csv` that counts what every step changed.

## Run it
```bash
pip install pandas
python src/generate_data.py
python src/clean.py
```
