5 Clustering Methods

Data Warehousing and Data Mining · Unit 5

Clustering Methods

Exam-focused notes for Clustering Methods (Data Warehousing and Data Mining, BIT454): what the TU syllabus asks and how it has actually been tested, with 5 solved past questions from this unit.

What this unit covers

  • Clustering approaches and types
  • K-means algorithm and limitations
  • Agglomerative clustering approach
  • Divisive clustering approach
  • Types of data in clustering
  • High dimensional data clustering
  • Similarity measures for ordinal data

K-means algorithm and limitations

208110 marks

List different clustering approaches. Write the algorithm of K-means algorithm. What are its limitations?[10]

Clustering approaches can be classified into the following main categories: 1. Partitioning Methods - Divide data into k non-overlapping clusters - Examples: K-means, K-medoids, PAM (Partitioning Around Medoids) 2. Hierarchical Methods - Build a hierarchy o...

Full solved answer →

Similarity measures for ordinal data

2082.15 marks

How do you find similarities between ordinal data attributes? Explain. [5]

Ordinal data are categorical attributes with a natural ordering or ranking among the categories. Examples include: education level (high school < bachelor < master < PhD), satisfaction ratings (poor < fair < good < excellent), or grades (A B C D F). Convert...

Full solved answer →

Types of data in clustering

2082.15 marks

Explain the types of data in clustering. [5]

Since reference notes are not provided, this answer draws from standard clustering theory in computer science: - Consists of continuous or discrete numeric values - Examples: age, salary, temperature, distance measurements - Can be directly used in distance...

Full solved answer →

High dimensional data clustering

2082.15 marks

What is outlier? How do you cluster high dimensional data? [2+3]

(a) What is an Outlier? An outlier is a data point that deviates significantly from the normal pattern or distribution of the rest of the dataset. Key characteristics: - It is substantially different from other observations - It lies far away from the centr...

Full solved answer →

Agglomerative clustering approach

2082.15 marks

Differentiate between agglomerative and divisive clustering approach. [5]

Agglomerative clustering is a bottom-up hierarchical clustering approach that works as follows: - Starts with each data point as an individual cluster - Iteratively merges the two closest/most similar clusters into a single cluster - Continues merging until...

Full solved answer →