Most clustering algorithms fall apart when you give them text‑heavy, high‑dimensional, or irregular data. CLOPE doesn’t. It was designed specifically for transaction data–fast, scalable, and with no need to pre‑specify the number of clusters.
K‑Means is the classic choice for numeric data, but it breaks in two predictable ways:
- You must guess the number of clusters up front.
- It relies on a distance matrix, which becomes painful once you pass ~10,000 rows.
CLOPE avoids both problems.
Where I first used it
About a decade ago I was analysing SAP work‑order data that was predominately text. I needed a way to cluster work orders after tokenising and cleaning the text in R. CLOPE was the right algorithm, but there was no R implementation in R at that time.
Switching tools mid‑workflow wasn’t an option. Exporting, running another tool, re‑importing… that breaks the entire fluent R analysis chain that I use. So I implemented CLOPE myself.
The first version was a literal translation of the paper which was slow and loop‑heavy but correct. Then I refactored it into vectorised R. It is now faster.
Where I use it now
One of the most useful applications is clustering file headers.
When I’m profiling large collections of CSVs, I extract each header, treat it as a transaction (a set of tokens), and run CLOPE. The clusters reveal structure immediately, the type of data:
- assay files
- drillhole collar files
- drillhole survey files
- geometallurgy testwork files
- anything else hiding vast list of files
It’s quick, reliable, and surprisingly effective.
Where you can use it
Geology summary logging Most operations ignore the free‑text summary column. Tokenise it, treat each record as a transaction, run CLOPE, and you suddenly have a new categorical feature for ML or exploratory analysis. Unused data becomes used.
Whole‑dataframe clustering If you want to cluster mixed data, assays, geology, geotech, metallurgy, simply discretise the continuous variables into categorical bins. Then treat each row as a transaction and run CLOPE. Same pattern as the header clustering workflow.
The R implementation can be found in my mining-ds-toolkit. It is in library/clope. The source R code is in the file clope.R and there is an example Quarto markdown (example.qmd) that shows basic examples of how to use the algorithm.