Some time ago I needed a lightweight tool that could quickly scan a folder, list all files, and profile any data files for analytics. Nothing quite fit my needs, so I built one.
It’s still a prototype, fully functional but not yet engineered for production. Think of it as a sketch that shows what data profiling and scoring is and how it can be applied to geology data for data analytics.
It was initially designed for projects with a lot of data files.
The workflow uses R and Quarto to automatically profile data files, extract metadata (column names, data types, encodings, basic statistics), and feed these results into a set of Scoring Functions that quantify data profiling consistently and deterministically. All profiling and scoring outputs are written to both SQLite and Excel for easy inspection and downstream use.
Data profiling is about extracting metadata but it also involves clustering and classification. CSV and Excel table header information is clustered into groups of similar data using the CLOPE transaction clustering algorithm. The header information is also used to classify the data into geological classes such as laboratory assays major elements or drill hole collar or drill hole survey. The classification uses simple matching of known column names.
The data profile scoring is similar to data quality scoring but the intent differs. A low profile score does not mean the data is poor or unusable. In this workflow, the score simply indicates where more effort may be required during data wrangling and data analysis. Lower scores highlight uplift opportunities, not data failures and indicate greater effort maybe required to process that data.
Data profile scoring is done across these dimensions:
- Accessibility: How easily data can be accessed and understood.
- Completeness: Missing values.
- Complexity: Structural complexity of the data (number of columns).
- Consistency: Uniformity of data encoding.
- Metadata: Availability of metadata documentation.
- Skewness: Distribution characteristics and data balance.
- Uniqueness: Duplication patterns.
- Usability: How well suited data is for analysis.
To make exploration of the results easier, I added a Shiny dashboard that visualizes the scoring results and highlights patterns across datasets.
The workflow lists all file types but only processes the CSV data files for now.
A complete working prototype—fully implemented in R and Quarto—is now available in the mining-ds-toolkit under tools-gui/.
The working example shown in the screenshots used the public data from the East Queensland Mineral Deposit Atlas found on the Queensland Government GSQ Open Data Portal.
It’s simple, practical, and ready to adapt to your own datasets.