Errors in geological logging create noise for an ML model which can lead to poor outcomes. The assays don't match the logging. The challenge is to identify inconsistencies in the geological logging and measurements across all the data. Including exploration drilling, production drilling, and ore control data.
A subtle inconsistency that is often overlooked is the definition of rock types, material types, and oretype. The exploration drilling database might log lithology one way, the ore control data might label the data using grouped litholgies, while the metallurgy data might use oretype. This inconsistency across data sets becomes a headache for data scientists to resolve before data can be combined. What if the ore control and metallurgy data label definitions then change?
A variety of statistical and machine learning techniques can be used to identify inconsistencies in the data prior to data joining and modeling. The data is flagged as either being consistent or as being likely to contain errors. This gives geologists a means to apply domain experience and expertise to correct the data before it is used in ML or AI.
The next challenge is then how to ensure that the data stays consistent. Label definitions can very easily drift over time. New labels get added. Consistent data can very quickly become inconsistent. Once the inconsistencies are created it is difficult to correct as they become embedded in the data and down-stream processes.