Data preparation takes more than half the time of many data science projects. The data quality part of preparation is often seen as junior work, but reality is almost the opposite. The earlier you are in a project the more influence your decisions have. Understanding data quality is not easy – it requires judgement, context and curiosity.
Unfortunately, most data scientists and analysts have very little training in data quality. Even in specialist degree courses, minimal time is devoted to data quality topics (10% of a single module, if you are lucky). The lecturers and professors who teach those courses tend to be experts in modelling or statistics, so the vicious circle continues.
We created the 6-step method when data scientists and analysts told us they:
- Regret not doing more,
- Want to know what constitutes good practice for investigating data quality, and
- Need a rulebook or process to follow.
I have now turned that into free-to-use teaching materials by adding seven YouTube videos that bring the method alive with real-world examples (see playlist). You can use any software to perform the 6-step method (Python, R, commercial tools, etc.). The materials include six Jupyter Notebooks (all part of the vizdataquality Python package) that you can run and customise.
I recommend you use these materials as follows:
- Read the 6-step Data Quality Method (it’s only 8 pages) to learn about the 69 tasks and 127 questions to consider asking about your data.
- Look at the downloadable spreadsheet that maps the tasks into a carefully thought-out 6 step order, to help you detect major issues early and clean data on the fly.
- Watch the Overview (8 minutes) to understand the what, when and how of the method, and why it can save you time, reduce cost and improve your results.
- Work through each step, watching the video and downloading and running the Jupyter Notebook yourself:
- Step 1: Is anything obviously wrong (10 minutes), featuring a car park fines example in the Step 1 Notebook.
- Step 2: Watch out for special values (15 minutes), featuring data about traffic collisions, bus stops, water leaks, health records and water meter readings in the Step 2 Notebook.
- Step 3: Is any data missing? (9 minutes) This features data about traffic collisions, potholes, water meter readings and traffic collisions in the Step 3 Notebook.
- Step 4: Check each variable (20 minutes), featuring the car park fines in the Step 4 Notebook.
- Step 5: Check combinations of variables (13 minutes), featuring the car park fines in the Step 5 Notebook.
- Step 6: Profile the cleaned data (9 minutes), featuring the car park fines in the Step 6 Notebook.
If you use these materials then please link directly to them rather than downloading a local copy. Please let me know if you use the materials. Just a quick email to “info AT surprisinganalytics.co.uk” telling me the organisation, course and approximate number of students will do. Any additional comments will be very welcome too, of course.
Finally, if you would like individual guidance, or me to give a talk or run a data quality course for you then do get in touch via the above email or LinkedIn.