Agglomerative Hierarchical Clustering: An Introduction to Essentials. (3) Standardization, Normalization and Dimensionality Reduction of a Data Matrix
Keywords:
corpus, vector, matrix, standardization, coefficient of variation, normalization, dimensionality reduction
Abstract
In a previous tutorial article I looked at a proximity coefficient and, in the light of that proximity created a vector-distance matrix and used it to construct a hierarchical tree using different hierarchical clustering methods which will be the basis for exploratory multivariate analysis. The present article deals with three topics: (i) standardization for variable scales variation, (ii) normalization for sample length variation, and (iii) dimensionality reduction or minimization of data space. These techniques reflect the author's academic background and particular area of interest and are, by necessity, not a particular purpose and are straightforwardly applicable to other kinds of data, and thus to a wide range of analysis in Linguistics. My treatment of these techniques is, necessarily, introductory and brief. I hope that this article will provide practitioners with an introductory overview of these techniques used for cluster analysis of electronic corpora of linguistic data.
Downloads
- Article PDF
- TEI XML Kaleidoscope (download in zip)* (Beta by AI)
- Lens* NISO JATS XML (Beta by AI)
- HTML Kaleidoscope* (Beta by AI)
- DBK XML Kaleidoscope (download in zip)* (Beta by AI)
- LaTeX pdf Kaleidoscope* (Beta by AI)
- EPUB Kaleidoscope* (Beta by AI)
- MD Kaleidoscope* (Beta by AI)
- FO Kaleidoscope* (Beta by AI)
- BIB Kaleidoscope* (Beta by AI)
- LaTeX Kaleidoscope* (Beta by AI)
How to Cite
References
R Belew (2000) Finding Out About: A Cognitive Perspective on Search Engine Technology and the WWW. 58(1), 125-128.
I Borg, P Groenen (2005) Modern Multidimensional Scaling.
C Chu, J Holliday, P Willett (2009) Effect of data standardization on chemical clustering and similarity searching. 49, 155-161.
J Dy (2008) Unsupervised feature selection. 19-39.
J Dy, C Bodley (2004) Feature selection for unsupervised learning. 5, 845-889.
R Gnanadesikan, J Kettenring, S Tsao (1995) Weighting and selection of variables for cluster analysis. 12(1), 113-136.
A Gordon, A Chapman And Halljain, M Murty, P Flynn (1999) Data clustering: a review. 31, 264-323.
Hermann Moisl (2015) Cluster Analysis for Corpus Linguistics.
T Kohonen (2001) Self-Organizing Maps.
G Milligan, Cooper (1985) An examination of procedures for determining the number of clusters in a data set. 50, 159-179.
Kevin Priddy, Paul Keller (2005) Artificial Neural Networks: An Introduction.
A Singhal, C Buckley, M Mitra (1996) Pivoted document length normalization. 21-29.
Amit Singhal, Gerard Salton, Mandar Mitra, Chris Buckley (1995) Document length normalization. 32(5), 619-633.
J Tenenbaum, V, J Langford (2000) A global geometric framework for nonlinear dimensionality reduction. 290, 2319-2323.
Published
2016-04-29
Issue
Section
License
Copyright (c) 2016 Authors and Global Journals Private Limited

This work is licensed under a Creative Commons Attribution 4.0 International License.