Method & references

Identification matrices and Willcox probabilities

PaleoID (BULKMAT 2.0)

 

- Overview

- User guide

- Method & references

- Feedback

 

- Launch the tool

 

 

Method & references

The identification matrix

 

The knowledge base is a matrix of n taxa × m characters (species). Each cell is the percentage of samples of that taxon in which the character is present (the “percent positive” value). Taxa may be depositional environments, foram bands or pollen zones, depending on the matrix used.

 

The matrix supplied with the program (MATBASIC-1-NORMAL) describes the North West Borneo environmental scheme developed by Sarawak Shell Berhad in the 1970s: thirteen environmental units from the lower coastal plain (LCP) through the holomarine and fluviomarine neritic realms (HIN/FIN, HMN/FMN, HON/FON, and their shallower “S” subdivisions) to outer–bathyal (O‑BAT). See the diagram on the overview page.

 

Willcox probability

 

For a sample, every character is scored present (1) or absent (0). The likelihood of the sample for a given taxon is the product over all characters of the taxon’s percent-positive value for present characters and its complement (100 − value) for absent characters, after division by 100. The likelihoods are normalised over all taxa so that they sum to one; these normalised values are the Willcox probabilities, and the three highest are reported for each sample.

 

Calculations are performed in logarithmic space, which removes the risk of numerical underflow on long species lists; the results are identical to the original algorithm to machine precision. An optional “positive entries only” mode computes the likelihood from present characters alone.

 

Diagnostics and diversity

 

- Species against: characters whose expected frequency in a taxon differs from the sample by more than 90% — the species that argue against (or conspicuously for) each of the three best identifications.

- Diversity indices (quantitative samples): total specimens, planktonic/benthonic ratio, Yule–Simpson index and Fisher’s alpha.

 

History

 

The method derives from Sneath (1979). The original program, BULKMAT, was written by P. Lesslar (XGS/1, October 1984) for the identification of well samples using presence–absence data. The program and the North West Borneo environmental scheme it applies are described in Lesslar (1987). This web edition reproduces its calculations and report format exactly, and adds a searchable species list, on-page options and structured logging of submissions for the further development of the identification knowledge base.

 

References

 

- Lesslar, P. (1987). Computer-assisted interpretation of depositional palaeoenvironments based on foraminifera. Geological Society of Malaysia, Bulletin 21, 103–119. Download the paper (PDF) — describes the BULKMAT program, the identification-matrix approach and the N.W. Borneo environmental scheme.

- Sneath, P.H.A. (1979). BASIC program for identification of an unknown with presence–absence data against an identification matrix of percent positive characters. Computers & Geosciences 5, 195–213.

- Willcox, W.B., Lapage, S.P., Bascomb, S. & Curtis, M.A. (1973). Identification of bacteria by computer: theory and programming. Journal of General Microbiology 77, 317–330.