Microarrays Research - Experiments, Designs, Statistics, Analysis, Software

Microarrays Research Today is a free monthly online journal that collates and summarizes the latest research about Microarrays, including details on experiments, designs, statistics, analysis, software.


Microarrays Research Today

Home

View Latest Issue

Information About Microarrays

Books on Microarrays

Advertising in Research Today

View Other Research Today Publications



Randomized maps for assessing the reliability of patients clusters in DNA microarray data analyses.

Bertoni A, Valentini G

DSI, Dipartimento di Scienze dell’ Informazione, Università degli Studi di Milano, Via Comelico 39, Milano, Italy.

OBJECTIVE:: Clustering algorithms may be applied to the analysis of DNA microarray data to identify novel subgroups that may lead to new taxonomies of diseases defined at bio-molecular level. A major problem related to the identification of biologically meaningful clusters is the assessment of their reliability, since clustering algorithms may find clusters even if no structure is present. METHODOLOGY:: Recently, methods based on random "perturbations" of the data, such as bootstrapping, noise injections techniques and random subspace methods have been applied to the problem of cluster validity estimation. In this framework, we propose stability measures that exploits the high dimensionality of DNA microarray data and the redundancy of information stored in microarray chips. To this end we randomly project the original gene expression data into lower dimensional subspaces, approximately preserving the distance between the examples according to the Johnson-Lindenstrauss (JL) theory. The stability of the clusters discovered in the original high dimensional space is estimated by comparing them with the clusters discovered in randomly projected lower dimensional subspaces. The proposed cluster-stability measures may be applied to validate and to quantitatively assess the reliability of the clusters obtained by a large class of clustering algorithms. RESULTS AND CONCLUSION:: We tested the effectiveness of our approach with high dimensional synthetic data, whose distribution is a priori known, showing that the stability measures based on randomized maps correctly predict the number of clusters and the reliability of each individual cluster. Then we showed how to apply the proposed measures to the analysis of DNA microarray data, whose underlying distribution is unknown. We evaluated the validity of clusters discovered by hierarchical clustering algorithms in diffuse large B-cell lymphoma (DLBCL) and malignant melanoma patients, showing that the proposed reliability measures can support bio-medical researchers in the identification of stable clusters of patients and in the discovery of new subtypes of diseases characterized at bio-molecular level.

Published 7 June 2006 in Artif Intell Med, 37(2): 85-109.
Full-text of this article is available online (may require subscription).

Place a permanent text-link or advertisement here for just US$15.

© 2004-2008 Microarrays Research Today. All Rights Reserved.



Microarrays Research Today Archive:

Volume 1 (2004)
  Issue 1 (June)
  Issue 2 (July)
  Issue 3 (August)
  Issue 4 (September)
  Issue 5 (October)
  Issue 6 (November)
  Issue 7 (December)

Volume 2 (2005)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)
  Issue 11 (November)
  Issue 12 (December)

Volume 3 (2006)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)
  Issue 11 (November)
  Issue 12 (December)

Volume 4 (2007)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)
  Issue 11 (November)
  Issue 12 (December)

Volume 5 (2008)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)



Microarrays Books

The Analysis of Gene Expression Data

The Analysis of Gene Expression Data