Site LogoContact
+ Explore the Problem Space. Discover the Value.

Mining DS Blog

A centralized blog for the mining-ds ecosystem which consists of the GitHub repositories mining-ds-vault and mining-ds-toolkit. Research, case studies, R notebooks, workflows, examples, code fragments, and GUI/CLI tools.

Kriging with Fuzzy Domain Boundaries

In the early 2000 when I was first doing Mineral Resource estimation I followed the rule book. I processed the geological data in the standard way, dealing with outliers using the typical clip at a threshold method (98th percentile being common). I happily built complex 3D wireframe models of the geology. And I did spatial estimation using Ordinary Kriging.

In 2006 it struck me that doing Ordinary Kriging or any form of kriging was compute expensive. The problem was matrix inversion. This was being done for every sample estimate. I surmised that their must be a way to minimize the cost of the matrix inversion. I started playing with Kalman Filters.

Mining-Ds-Vault Mining-Ds-Toolkit
Research Kriging Fuzzy Boundaries Go
Open Report ➔
DAGs aren’t daggy

They’re the most powerful tool we have to stop fooling ourselves with historical data. In geometallurgy and mining data science, the biggest risks aren’t bad assays or noisy test‑work. It’s the hidden structural bias we introduce without realising it. And once it’s baked into a model, the damage is done.

The challenge:

Across geology, metallurgy, and mining analytics, selection bias creeps in quietly.

I’ve seen it first-hand. Two labs, two rounds of metallurgy test‑work, and a big discrepancy. At first glance, Lab B looked wrong, Until the DAG revealed the real culprit. Different rock properties sent to each lab. The “lab bias” wasn’t a lab bias at all. It was a hidden selection variable.

Mining-Ds-Vault
Discover Geometallurgy DAG R
Open Report ➔
DPS

Some time ago I needed a lightweight tool that could quickly scan a folder, list all files, and profile any data files for analytics. Nothing quite fit my needs, so I built one.

It’s still a prototype, fully functional but not yet engineered for production. Think of it as a sketch that shows what data profiling and scoring is and how it can be applied to geology data for data analytics.

Mining-Ds-Toolkit
Tools GUI Shiny Dashboard Scoring R
Scoring Functions

Geological data validation can include the use of twin drill holes to verify historic drill hole data, verify high-grade assay intersections, and check for sampling bias. Traditionally, assessing these paired data involves subjective visual inspections of scatter plots, quantile-quantile plots, and down-hole logs, or the use of global statistics—a process that does not scale across large databases.

To overcome the problem of scaling the analysis, Scoring Functions are introduced to illustrate how the observed geological variability between the paired data can be translated into a normalized, deterministic metric, which can then be used as a standardized quality KPI.

Mining-Ds-Vault
Discover Data-Profiling Scoring R
Open Report ➔
CLOPE in R

Most clustering algorithms fall apart when you give them text‑heavy, high‑dimensional, or irregular data. CLOPE doesn’t. It was designed specifically for transaction data–fast, scalable, and with no need to pre‑specify the number of clusters.

K‑Means is the classic choice for numeric data, but it breaks in two predictable ways:

  • You must guess the number of clusters up front.
  • It relies on a distance matrix, which becomes painful once you pass ~10,000 rows.

CLOPE avoids both problems.

Mining-Ds-Vault Mining-Ds-Toolkit
Discover Clustering R
Exploratory Modelling with Bayesian Networks

Most people who work with the Titanic dataset follow the same path: load the data, engineer a few features, and train a random forest to predict survival. It’s a good exercise. But it skips the most important part of the analysis.

Understanding the data.

For this example, I use the Titanic dataset as a statistical storytelling problem, not a machine-learning (ML) competition. The goal isn’t to build the most accurate model. It is to uncover the relationships that explain why certain groups survived and others didn’t.

Mining-Ds-Vault
Discover Bayesian Networks Data Analysis Titanic Dataset R
Open Report ➔
DMCONV

Sometimes a simple tool is all you need. In the age of AI it is AI this and AI that, but here is a tool that does what it needs to do, no fuss. The tool simply reads in a ‘.dm’ Datamine Studio binary file and outputs to either a CSV or Apache Arrow Parquet file. It is a simple CLI. Reminiscent of the old but useful GSLIB geostatistics Fortran library. It can be used as a one off or in a workflow or called as an agentic AI tool.

Mining-Ds-Toolkit
Tools CLI File Conversion Go