7 Cluster Analysis

Data Warehousing and Data Mining · Unit 7 · 8 hrs

Cluster Analysis

Exam-focused notes for Cluster Analysis (Data Warehousing and Data Mining, CSC420): what the TU syllabus asks and how it has actually been tested, with 9 solved past questions from this unit.

What this unit covers

  • Types of data in cluster analysis
  • Similarity and dissimilarity between objects
  • Clustering techniques: Partitioning (k-means, k-means++, Mini-Batch k-means, k-medoids)
  • Hierarchical (Agglomerative and Divisive)
  • Density based (DBSCAN)
  • Outlier analysis

Clustering techniques

20815 marks

Using k-means++ algorithm and Euclidean distance, find the initial 3 cluster centroids from A1 = (3, 11), A2 = (3, 6), A3 = (9, 5), A4 = (6, 9), A6 = (7, 5), A7 = (2, 3), A8 = (5, 10). Choose (3, 11) as one of the initial centroids. [5]

Data points: Point Coordinates -------------------- A1 (3, 11) A2 (3, 6) A3 (9, 5) A4 (6, 9) A6 (7, 5) A7 (2, 3) A8 (5, 10) Number of clusters: $k = 3$ First centroid (given): $C1 = (3, 11)$ 1. First centroid is fixed as $C1 = (3,11)$. 2. For each point com...

Full solved answer →
20815 marks

Differentiate between k-means and k-medoids clustering algorithm. [5]

Both K-Means and K-Medoids are partitioning-based clustering algorithms that divide a dataset of n objects into k clusters. However, they differ significantly in how they represent cluster centers and handle data. --- Feature K-Means K-Medoids --------- Cen...

Full solved answer →
20805 marks

How K-medoids clustering differs from K-means clustering? Divide the following data points into two clusters using kmedoids algorithm. Show computation up to 3 iterations. {(70,85), (65,80), (72,88), (75,90), (60,50), (64,55), (62,52), (63,58)}. [5]

Data points (8 points), to be split into $k = 2$ clusters: Label Point ------ P1 (70, 85) P2 (65, 80) P3 (72, 88) P4 (75, 90) P5 (60, 50) P6 (64, 55) P7 (62, 52) P8 (63, 58) Requirements: show computation up to 3 iterations. Distance metric not specified, s...

Full solved answer →
20795 marks

Discuss the concept of K-means++ and Mini-batch K-means algorithm. [5]

--- In the standard K-means algorithm, the initial cluster centers (centroids) are chosen randomly. This random initialization leads to a problem called initialization sensitivity, where the final clusters formed depend heavily on which points were chosen i...

Full solved answer →
20785 marks

Apply K(=2)-Means algorithm over the data (185, 72), (170, 56), (168, 60), (179, 68), (182, 72), (188, 77) up to two iterations and show the clusters. Initially choose first two objects as initial centroids. [5]

Point Data ------------- P1 (185, 72) P2 (170, 56) P3 (168, 60) P4 (179, 68) P5 (182, 72) P6 (188, 77) Initial Centroids: $C1 = (185, 72)$, $C2 = (170, 56)$ Distance metric: Euclidean, $d = \sqrt{(x2-x1)^2 + (y2-y1)^2}$ --- Distances to $C1=(185,72)$: - P1:...

Full solved answer →

Density based

20805 marks

Discuss working of DBSCAN algorithm. [5]

DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. It is a density-based clustering method where clusters are defined as dense regions in the data space, separated by regions of lower density points. The key idea is: for each poi...

Full solved answer →
207810 marks

How do you compare two classifiers? Given the points A(3,7), B(4,6), C(5,5), D(6,4), E(7,3), F(6,2), G(7,2) and H(8,4), find the core points, border points and outliers using DBSCAN. Take Eps 2.5 and MinPts = 3.[10]

Two classifiers can be compared using the following techniques: 1. Confusion Matrix based metrics Build a confusion matrix (TP, TN, FP, FN) for each classifier and compute: - Accuracy $= \dfrac{TP+TN}{TP+TN+FP+FN}$ - Precision $= \dfrac{TP}{TP+FP}$ - Recall...

Full solved answer →

Hierarchical

20795 marks

What are two categories of hierarchical clustering? Divide the following data points into two clusters using agglomerative clustering. { (2,10), (2,5), (8,4), (5,8), (7,5), (6,4) } [5]

Data points to cluster: - P1 = (2, 10) - P2 = (2, 5) - P3 = (8, 4) - P4 = (5, 8) - P5 = (7, 5) - P6 = (6, 4) Target: 2 clusters. Linkage method not specified; I will use single linkage (minimum distance) as is standard for this textbook problem. --- 1. Aggl...

Full solved answer →

Types of data in cluster analysis

20755 marks

What is the purpose of cluster analysis in data mining? Explain. [5]

Cluster analysis is a data mining technique in which data objects are grouped into clusters based on their similarity. Objects within the same cluster have high similarity to one another, while objects belonging to different clusters are dissimilar to each ...

Full solved answer →