Data Warehousing and Data Mining · Unit 5
Clustering Methods
Exam-focused notes for Clustering Methods (Data Warehousing and Data Mining, BIT454): what the TU syllabus asks and how it has actually been tested, with 5 solved past questions from this unit.
What this unit covers
- Clustering approaches and types
- K-means algorithm and limitations
- Agglomerative clustering approach
- Divisive clustering approach
- Types of data in clustering
- High dimensional data clustering
- Similarity measures for ordinal data
K-means algorithm and limitations
List different clustering approaches. Write the algorithm of K-means algorithm. What are its limitations?[10]
Clustering approaches can be classified into the following main categories: 1. Partitioning Methods - Divide data into k non-overlapping clusters - Examples: K-means, K-medoids, PAM (Partitioning Around Medoids) 2. Hierarchical Methods - Build a hierarchy o...
Full solved answer →Similarity measures for ordinal data
How do you find similarities between ordinal data attributes? Explain. [5]
Ordinal data are categorical attributes with a natural ordering or ranking among the categories. Examples include: education level (high school < bachelor < master < PhD), satisfaction ratings (poor < fair < good < excellent), or grades (A B C D F). Convert...
Full solved answer →Types of data in clustering
Explain the types of data in clustering. [5]
Since reference notes are not provided, this answer draws from standard clustering theory in computer science: - Consists of continuous or discrete numeric values - Examples: age, salary, temperature, distance measurements - Can be directly used in distance...
Full solved answer →High dimensional data clustering
What is outlier? How do you cluster high dimensional data? [2+3]
(a) What is an Outlier? An outlier is a data point that deviates significantly from the normal pattern or distribution of the rest of the dataset. Key characteristics: - It is substantially different from other observations - It lies far away from the centr...
Full solved answer →Agglomerative clustering approach
Differentiate between agglomerative and divisive clustering approach. [5]
Agglomerative clustering is a bottom-up hierarchical clustering approach that works as follows: - Starts with each data point as an individual cluster - Iteratively merges the two closest/most similar clusters into a single cluster - Continues merging until...
Full solved answer →Make Unit 5 stick
Practice BIT454 with flashcards & quizzes