Site LogoContact
+ Case Studies // in Action

Examples of domain-hard problems.

Series of case studies demonstrating how Data Science can help with geological, geometallurgical, and operational problems. Examples of known problems with engineered solutions and repeatable workflows.

+ Case File #01

Assessing geology drillhole logging consistency.

In resource estimation we assess quality assurance and quality control (QAQC) data to ensure the data is fit for use and assign confidence. We might indirectly validate drillhole geology logging during 3D geological interpretation. What about validating the consistency of geological logging with respect to assays and density and hardness data (and other numeric geological data)? This is important when we use all the geological data for predictive modeling. The consistency of the data becomes very important.

Errors in geological logging create noise for an ML model which can lead to poor outcomes. The assays don't match the logging. The challenge is to identify inconsistencies in the geological logging and measurements across all the data. Including exploration drilling, production drilling, and ore control data.

A subtle inconsistency that is often overlooked is the definition of rock types, material types, and oretype. The exploration drilling database might log lithology one way, the ore control data might label the data using grouped litholgies, while the metallurgy data might use oretype. This inconsistency across data sets becomes a headache for data scientists to resolve before data can be combined. What if the ore control and metallurgy data label definitions then change?

A variety of statistical and machine learning techniques can be used to identify inconsistencies in the data prior to data joining and modeling. The data is flagged as either being consistent or as being likely to contain errors. This gives geologists a means to apply domain experience and expertise to correct the data before it is used in ML or AI.

The next challenge is then how to ensure that the data stays consistent. Label definitions can very easily drift over time. New labels get added. Consistent data can very quickly become inconsistent. Once the inconsistencies are created it is difficult to correct as they become embedded in the data and down-stream processes.

Why can't conventional methods solve this?

Conventional techniques struggle with large amounts of data scattered across multiple data sets which are not easily joined. The ML approach handles the different data and captures subtle associations that would not normally be identified using manual methods such as visualizing the data in 3D.

What assumptions must be true?

For multiple data sets the assumption is that there is some overlap in the data to allow data fusion techniques to be used or data stacking or traditional data joining. It is also assumed that there is more than one numeric field and that there is geological logging. The more relevant variables that are available the better the outcome.

What is the expected outcome and benefits?

The outputs would be a data set with each row flagged with either being consistent or not. The output will also include an additional field or fields that suggest the most likely geological logging label (for instance, the most likely lithology).

The benefits include uplifting data quality, monitoring geological logging quality, better ML performance, and a better understanding of the data.

+ Case File #02

Predicting rock density from geological logs and assay data.

Rock density is one of the three fundamental components of estimating a mineral resource along with grade and volume. Without rock density it is not possible to estimate tonnage or contained metal. A lot of effort is put into quantifying grade and even building the 3D geological interpretation. Density tends to be undersampled and under-analyzed. A consequence of this is that there typically is very limited density measurements with sparse spatial distribution. At worst all that is available is rock type and from that an average global density is inferred.

One solution to this is to build a predictive model trained using machine learning (ML) where density is estimated from the more dense data that includes assays and geological logging (lithology, alteration, oxidation, mineralisation, domains). However, this hides a fundamental problem with resource classification. If the rock density data is sparse and the uncertainty in the predicitons from a ML model is not known then how can a Measured classification be applied? Especially if we train the model with domain features and grade bins which amplifies problems with limited and sparse data. How can we trust the ML model? How do we know if the model suffers from data traps or statistical traps such as the Simpsons Paradox and Berksons Paradox?

The Nimmo Analytics 4D's Data Science workflow (NA-4DS) aims to address these uncertainties before any code or analysis is done and well before they can become a bigger problem.

After building the predictive ML model the next challenge is keeping the model up to date and prevent model drift from creeping in and impacting the estimates. The better question though is how can we use the ML model to help us identify where to add additional sampling to improve the model with the least cost.

Why can't conventional methods solve this?

Conventional methods can do this problem. The conventional methods are using statistical regression to predict rock density. However, the resulting regression model is likely to not be the best model. Also, if not done correctly, the regression model will likely include variables that should not be included. Or variables that should be included but aren't. A bigger problem lurks in how the modelling is done. The modeling needs to carefully consider hidden (unobserved) features that affect the target variable (confounders) and the statistical traps (Simpsons Paradox and Berksons Paradox).

Performing statistical regression is more involved than simply training a regression model. Data Science experts and Statisticians know how to build an unbiased regression model. Data Science allows selecting regression techniques that can, and often do, outperform a standard statistical (ordinary least squares linear) regression. A statistical model aims to inform causal relationships. A data science model aims to predict.

What assumptions must be true?

The data must be clean and complete. The measured rock geochemistry and optional properties (such as moisture, hardness, and porosity, and geological logging) must be sufficient to explain variability in rock density. A single assay (such as gold grade) is usually not sufficient.

What is the expected outcome and benefits?

The geological drillhole data with missing density estimated using the regression model along with a flag to identify between predicted and measured rock density. Depending on the regression model selected there may also be a measure of uncertainty for each estimate. A CLI tool can also be provided that predicts rock density from available rock properties (excluding density) where the inputs are in CSV format and the output is in the same format. The specifications of the tool should be defined before any modeling begins. An ML model may be too large to provide the model formula or model weights.

As part of the analysis the impact of data gaps and coverage can be added to the model uncertainty to get an indication of the impact on resource classification.

+ Case File #03

Fusing geology and metallurgy when the crushed ore stockpile prevents the join.

Sometimes we have lots of data across multiple tables that can easily be joined using standard relational database joins or simply stacking the data. Sometimes we have data that has limited overlap of fields between tables and a relational join is not possible. Other times we have two data sets (our inputs and our ouputs) in two separate files with no way of joining the data.

For example consider the situation where we are trying to build a geometallurgy predictive model using geological inputs to predict metallurgy response. Our inputs are the geology ore control data that is split into batches by oretype. Our targets are in a separate data set that is mineral processing plant daily and monthly data of blended ore. The inputs don't match the target outputs. The cause of the blend is a crushed ore stockpile (COS) which mixes the rock over an indeterminate time frame. We can guess the merging or matching of the data by date and time. We could try a physics simulation to approximate the COS. Or we could try exploiting ML learning and fuse the data without joining the data. The later avoids unnecessary assumptions.

The Nimmo Analytics 4D workflow exposes this problem early so a solution can be designed before any analysis is done. If there is no solution then there is no analysis. However, in the case of a COS resulting in mixed ore the solution is relatively simple and robust. The solution uses concepts from mathematics and data science and requires both daily and monthly data. It is quick to build and easy to test and maintain compared with a complex digital twin.

Building a geometallurgy predictive model when there is a COS that prevents direct data joins between geology and metallurgy data is the easy part. The next challenge is ensuring that the model is general and future looking. The mineral processing data from the process plant is likely to suffer from sever data bias. Not all oretypes or rock and assay combinations will be present in the data. A bigger problem would arise if the process plant handled only oxide but the plant is expanded to handle transitional or sulphide ore. The process plant data is likely to be incomplete which means some form of data augmentation would likely be needed to ensure the model is suitable for the entire deposit and not just for the areas that had alread been mined.

Why can't conventional methods solve this?

The only other way to build an ML model using two data sets where the inputs are in one data table and the target variables are in another table is to build a complex digital twin. A digital twin can be built to simulate the flow of data from insitu to process plant to product. This approach requires complex physics models to simulate the crushed ore stockpile and its effect on blending. The ML approach is simple, requires very few assumptions compared with the digital twin, and can be done quickly.

What assumptions must be true?

All that is required is the input geology ore control data and the process plant daily and monthly data. It is assumed that the geology ore control data contains sufficient geology detail that contains features that impact the metallurgy response. Consideration needs to be given to whether the process plant and its conditions have and or will change over time and whether the ore control data has also changed within the time period being assessed.

What is the expected outcome and benefits?

A trained geometallurgy predictive ML model that accounts for actual process plant performance. The outcome is also the benefit. Being able to use data that would not ordinarily be able to be used for training a geometallurgy model other than for an expensive digital twin.

+ Case File #04

Prescriptive analytics from a predictive geometallurgy model.

Most geometallurgy ML modeling and analysis ends when the predictive model is built. What if the real challenge is in answering the question of how to optimally blend material types to ensure optimal process plant performance. The ML model can answer the question: What is the predicted metal recovery given the rock properties? But it does not answer the question: What should I do given that the predicted recovery is not optimal?

The solution is using prescriptive statistics and conter-factual ML analysis combined with simulation and discrete deterministic rules and constraints. The predictive ML model is useful but it can be even more useful. It now becomes a matter of deployment. How will the ML model be used? To predict or to help decide? But to do full counter-factural ML analysis we will need a type of model that is suited for this task. Not all models are. This is why it is very important to ask the questions around how the model will be used before the model is built.

The next challenge is to test the validity of the model. Manipulating the inputs for an ML model is easy. Manipulating the physical inputs to achieve the target outcome is much harder. As blending two different rocks with different properties with different metallurgy performance is not necessarily going to result in the metallurgy performance being a linear weighted average of the metallurgy performance of the two rocks. Nor is it going to match the performance of a rock that has the same properties as the blend. The trained ML model has learnt to map the rock properties with the target metallurgy performance. This is not the same as physically mixing rock to match the properties of a rock sample in the data and hoping that that blend gives the same results as the physical rock sample. The training set needs to include blended rock samples and their metallurgy performance and the ML model needs to account for the potential non-linear mixing effects of each rock in the blend.

Why can't conventional methods solve this?

Machine learning is suited for this task. The approach uses standard ML counter-factual Analysis and simulation to identify what needs to change in the inputs to achieve the desired target. The ML counter-factual analysis can the be used along with simulation to identify the optimal blending rule under operational constraints.

What assumptions must be true?

The ML model contains the necessary features to accurately predict geometallurgy material type or oretype along with metallurgy performance. The features being adjusted to optimise the blend must be valid operational rules. It is assumed that the blending of different material types or oretypes can adjust the metallurgy response in a favourable manner. Using batched geometallurgy data used to train the geometallurgy predictive ML model may not be sufficient. Taking a weighted average of the mixed rock properties and using it in the geometallurgy model to predict metallurgy performance may not be the same as the effect of physically blending that rock. To prove that linear weighting of rock properties is sufficient to approximate the effect of blending a set of expensive testwork on the various blends would be needed. Or, alternatively, use the process plant data to build an ML model that learns the effects of blending of oretypes (Case File #03).

What is the expected outcome and benefits?

The outcome depends on the project requirements and specifications. It might include a CLI tool that determines the optimal blend given batches of rock. Or it might be recommendations or set of rules to apply given the available rock on the run-of-mine (ROM) and available stockpiles.

+ Case File #05

Predicting the next equipment component to fail.

Consider the challenge of trying to understand equipment failure modes and causes. We might immediately jump straight into applying our favourite ML method or might expend a lot of time and energy exploring the data. What we failed to do is ask the right questions at the right time. Understanding our goal for the project is the first challenge. In the case of equipment failure our real goal might be to understand the sequence of failures. If component X fails then what is the most likely component to fail next. Or a better question would be to ask is there a better way to reduce equipment downtime by optimising warehousing by anticipating component usage to ensure that replacement parts are available when they are needed.

Statistical models can be used to answer questions like this. Rather than building a complex AI model. If the answer from the model is what we need such as giving us a prediction of the next component to fail on an equipment given its maintenance history (a Markov chain), then that is the best model for the solution. If not then we continue to explore options to design the right solution.

The next challenge is how to deploy a predictive model to maximize its value. How do we integrate the model into the warehousing stock optimisation and procurement process? If we ask the right questions at the beginning before we build the solution the answer will be part of the solution and it will be easier to deploy.

Why can't conventional methods solve this?

Machine learning is widely used for predicting component failure.

What assumptions must be true?

The assumptions are dependent on the project. They are captured as part of the process of defining the problem and designing the solution.

What is the expected outcome and benefits?

The outcome depends on the project and its goals and is specified as part of the requirements and specifications.

+ Case File #06

How to do DS without transferring data

A typical mining consulting engagement has the consultant requesting as much site data as possible. This puts the data at risk of being caught by third parties. An NDA will not necessarily prevent this. Often data is transferred via a digital folder. The risk in the transfer and storage of data on the third-party vendor. The challenge is to do data science and ML without risking data security.

The solution is to not transfer data. ML can train a model on the raw data, processed data, or a statistical facsimile—synthetic data. The technology does not need to be invented, it exists now. Concepts such as data masking, hashing, adding noise is well known for highly sensitive environments such as medical.

Nimmo Analytics uses an encoder-decoder model with an internal representation (IR) to allow building ML models without risking data security. The IR is a combination of techniques to obfuscate the data of sensitive information while retaining the salient statistical relationships and patterns. Data stays under your control.

The IR can also be used to fine tune a small language model (SLM) for use as a model to explain reasoning and confidence.

Why can't conventional methods solve this?

Data masking and techniques such as Federated Learning are used in highly sensitive environments such as medical. It is rarely used in mining industry consulting. Training an ML model directly on raw data is usually the prefered approach. But in consulting this puts the data at risk when transfering between client and consultant.

Nimmo Analytics still uses the conventional approach where the raw data is transferred and used in data analysis and modeling. For some projects this may be what is required for best results. But it is not a necessary condition.

What assumptions must be true?

That the data can be trusted to be representative of the population.

What is the expected outcome and benefits?

Total control over your data and no leakage of sensitive information. The internal representation provides additional capabilities that the raw data alone does not. It captures the data quality, the data gaps, and the generative models. This is a step that is needed whether an IR is generated or not.

+ Work With Us

Bring us your hardest domain problem.

We'll determine whether the challenge can be engineered into a repeatable, defensible workflow.

Discuss a Challenge