Data Warehousing and Data Mining · Unit 7 · 8 hrs
Cluster Analysis
Exam-focused notes for Cluster Analysis (Data Warehousing and Data Mining, CSC420): what the TU syllabus asks and how it has actually been tested, with 9 solved past questions from this unit.
What this unit covers
- Types of data in cluster analysis
- Similarity and dissimilarity between objects
- Clustering techniques: Partitioning (k-means, k-means++, Mini-Batch k-means, k-medoids)
- Hierarchical (Agglomerative and Divisive)
- Density based (DBSCAN)
- Outlier analysis
Clustering techniques
Using k-means++ algorithm and Euclidean distance, find the initial 3 cluster centroids from A1 = (3, 11), A2 = (3, 6), A3 = (9, 5), A4 = (6, 9), A6 = (7, 5), A7 = (2, 3), A8 = (5, 10). Choose (3, 11) as one of the initial centroids. [5]
Data points: Point Coordinates -------------------- A1 (3, 11) A2 (3, 6) A3 (9, 5) A4 (6, 9) A6 (7, 5) A7 (2, 3) A8 (5, 10) Number of clusters: $k = 3$ First centroid (given): $C1 = (3, 11)$ 1. First centroid is fixed as $C1 = (3,11)$. 2. For each point com...
Full solved answer →Differentiate between k-means and k-medoids clustering algorithm. [5]
Both K-Means and K-Medoids are partitioning-based clustering algorithms that divide a dataset of n objects into k clusters. However, they differ significantly in how they represent cluster centers and handle data. --- Feature K-Means K-Medoids --------- Cen...
Full solved answer →How K-medoids clustering differs from K-means clustering? Divide the following data points into two clusters using kmedoids algorithm. Show computation up to 3 iterations. {(70,85), (65,80), (72,88), (75,90), (60,50), (64,55), (62,52), (63,58)}. [5]
Data points (8 points), to be split into $k = 2$ clusters: Label Point ------ P1 (70, 85) P2 (65, 80) P3 (72, 88) P4 (75, 90) P5 (60, 50) P6 (64, 55) P7 (62, 52) P8 (63, 58) Requirements: show computation up to 3 iterations. Distance metric not specified, s...
Full solved answer →Discuss the concept of K-means++ and Mini-batch K-means algorithm. [5]
--- In the standard K-means algorithm, the initial cluster centers (centroids) are chosen randomly. This random initialization leads to a problem called initialization sensitivity, where the final clusters formed depend heavily on which points were chosen i...
Full solved answer →Apply K(=2)-Means algorithm over the data (185, 72), (170, 56), (168, 60), (179, 68), (182, 72), (188, 77) up to two iterations and show the clusters. Initially choose first two objects as initial centroids. [5]
Point Data ------------- P1 (185, 72) P2 (170, 56) P3 (168, 60) P4 (179, 68) P5 (182, 72) P6 (188, 77) Initial Centroids: $C1 = (185, 72)$, $C2 = (170, 56)$ Distance metric: Euclidean, $d = \sqrt{(x2-x1)^2 + (y2-y1)^2}$ --- Distances to $C1=(185,72)$: - P1:...
Full solved answer →Density based
Discuss working of DBSCAN algorithm. [5]
DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. It is a density-based clustering method where clusters are defined as dense regions in the data space, separated by regions of lower density points. The key idea is: for each poi...
Full solved answer →How do you compare two classifiers? Given the points A(3,7), B(4,6), C(5,5), D(6,4), E(7,3), F(6,2), G(7,2) and H(8,4), find the core points, border points and outliers using DBSCAN. Take Eps 2.5 and MinPts = 3.[10]
Two classifiers can be compared using the following techniques: 1. Confusion Matrix based metrics Build a confusion matrix (TP, TN, FP, FN) for each classifier and compute: - Accuracy $= \dfrac{TP+TN}{TP+TN+FP+FN}$ - Precision $= \dfrac{TP}{TP+FP}$ - Recall...
Full solved answer →Hierarchical
What are two categories of hierarchical clustering? Divide the following data points into two clusters using agglomerative clustering. { (2,10), (2,5), (8,4), (5,8), (7,5), (6,4) } [5]
Data points to cluster: - P1 = (2, 10) - P2 = (2, 5) - P3 = (8, 4) - P4 = (5, 8) - P5 = (7, 5) - P6 = (6, 4) Target: 2 clusters. Linkage method not specified; I will use single linkage (minimum distance) as is standard for this textbook problem. --- 1. Aggl...
Full solved answer →Types of data in cluster analysis
What is the purpose of cluster analysis in data mining? Explain. [5]
Cluster analysis is a data mining technique in which data objects are grouped into clusters based on their similarity. Objects within the same cluster have high similarity to one another, while objects belonging to different clusters are dissimilar to each ...
Full solved answer →Make Unit 7 stick
Practice CSC420 with flashcards & quizzes