CSC420 · TU past paper
Data Warehousing and Data Mining 2078 question paper
The complete TU 2078 exam paper for Data Warehousing and Data Mining (CSC420), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksSigned networkHideAnswer
Write down any one advantage and disadvantage of MOLAP over ROLAP. Define signed network and how do you check whether it is balanced or not? How beam search reduces the space complexity? Illustrate with an example.[10]
--- Advantage of MOLAP over ROLAP: MOLAP stores data in a multidimensional cube (pre-aggregated), which allows faster query performance. Since data is pre-computed and stored in optimized multidimensional arrays, analytical queries are a...
- 210 marksNumericalFinding frequent itemsetHideAnswer
How Concept Hierarchy is Used in Extracting Information
Concept hierarchy is used in extracting information by:
- Organizing data at different levels of abstraction
- Enabling generalization and specialization of patterns
- Allowing drill-down and roll-up operations for multi-level analysis
- Facilitating hierarchical association rule mining
- Supporting semantic understanding of discovered patterns
--- A concept hierarchy is a sequence of mappings from low-level (specific) concepts to higher-level (general) concepts. It organizes data values into levels of abstraction. How it is used in extracting information: 1. Multilevel mining:...
- 310 marksNumericalDensity basedHideAnswer
How do you compare two classifiers? Given the points A(3,7), B(4,6), C(5,5), D(6,4), E(7,3), F(6,2), G(7,2) and H(8,4), find the core points, border points and outliers using DBSCAN. Take Eps 2.5 and MinPts = 3.[10]
Two classifiers can be compared using the following techniques: 1. Confusion Matrix based metrics Build a confusion matrix (TP, TN, FP, FN) for each classifier and compute: - Accuracy $= \dfrac{TP+TN}{TP+TN+FP+FN}$ - Precision
- 45 marksIssues and ApplicationsHideAnswer
When a pattern is said to be interesting? List the issues of data mining. [5]
A pattern is said to be interesting if it is: 1. Easily understood by humans 2. Valid on new or test data with some degree of certainty 3. Novel (previously unknown or validates a hypothesis) 4. Useful (can be acted upon to gain some adv...
- 55 marksSpatial data miningHideAnswer
Define spatial data mining. What are the challenges of multimedia mining? Describe with an example. [5]
--- Spatial data mining refers to the extraction of knowledge, spatial relationships, or other interesting patterns that are not explicitly stored in spatial databases. Such mining demands an integration of data mining with spatial datab...
- 65 marksNumericalBayesian classificationHideAnswer
Consider the following data set. Find out whether the object with attribute Confident = Yes, Sick = No will Fail or Pass using Bayesian classification.
Confident Studied Sick Result Yes No No Fail Yes No No Pass No Yes Yes Fail No Yes Yes Pass Yes Yes Yes Pass [5]
Row Confident Studied Sick Result --------------------------------------- 1 Yes No No Fail 2 Yes No No Pass 3 No Yes Yes Fail 4 No Yes Yes Pass 5 Yes Yes Yes Pass Query object: Confident = Yes, Sick = No (Studied not specified). Total in...
- 75 marksCube materializationHideAnswer
What are the choices for data cube materialization? Explain the strategies for cube computation. [5]
Data Cube Materialization: Choices and Strategies for Cube Computation
Introduction
Data cube computation is an essential task in data warehouse implementation. The precomputation (materialization) of all or part of a data cube can greatly reduce response time and enhance OLAP performance. However, it is challenging because it may require huge computational time and storage space.
Choices for Data Cube Materialization
There are three main choices for data cube materialization:
1. Full Materialization
- All cuboids in the data cube lattice are precomputed and stored.
- Provides the fastest query response time since all aggregations are already computed.
- Disadvantage: Requires enormous storage space and high computation cost, especially when the number of dimensions is large.
2. No Materialization
- No cuboids are precomputed; all queries are computed on-the-fly from the base cuboid (raw data).
- Disadvantage: Very slow query response time, especially for complex OLAP queries on large datasets.
3. Partial Materialization
- Only a selected subset of cuboids is precomputed and stored.
- Provides a balance between storage cost and query response time.
- This is the most practical and widely used approach.
Strategies for Cube Computation
i) Full Cube Computation
- Computes every possible cuboid for all combinations of dimensions.
- For an n-dimensional data cube, the total number of cuboids is 2^n.
- Suitable only when n is small and storage is not a constraint.
ii) Iceberg Cube
- Instead of computing the complete cube, only those cuboids (cells) that satisfy a minimum support threshold (called the iceberg condition) are computed.
- Cells with aggregate values below the threshold are discarded.
- Example condition:
HAVING count(*) >= min_support - This avoids computing and storing many low-value, sparse cells, saving both time and space.
- The name comes from the idea that only the "tip of the iceberg" (the significant aggregates) is materialized.
iii) Shell Cube (Partial Materialization)
- A strategy for partial materialization where only cuboids involving a small number of dimensions (such as 3 to 5) are precomputed.
- These precomputed cuboids form a shell cube.
- For example, in an n-dimensional data cube, we could compute all cuboids with 3 dimensions or less, resulting in a Shell Cube of size 3.
- Remaining higher-dimensional cuboids are computed on-the-fly using the precomputed shell cuboids.
- Advantage: Significantly reduces storage while still speeding up many common queries.
iv) Closed Cube
- Only closed cells are materialized. A cell is closed if there is no supercell with the same measure value (e.g., same count).
- Reduces redundancy in stored data by avoiding cells whose values can be directly derived from a more specific cell.
Summary Table
Strategy What is Stored Storage Cost Query Speed Full Materialization All cuboids Very High Very Fast No Materialization None None Very Slow Iceberg Cube Cells above threshold Low Fast Shell Cube Cuboids with few dimensions Medium Moderate Closed Cube Non-redundant closed cells Low Fast
Conclusion
Among all strategies, partial materialization (especially Iceberg Cube and Shell Cube) provides the best trade-off between storage cost and query performance, making them the most practical choices for real-world data warehouse implementations.
- 85 marksSigned networkHideAnswer
Show the conflict between theory of balance and status. How do you improve Apriori? [5]
--- In signed social networks, links between users can be either positive (+) or negative (-). - Balance Theory predicts the sign of a link based on the principle: "the friend of my friend is my friend." - Status Theory predicts the sign...
- 95 marksA multidimensional data modelHideAnswer
Differentiate between star schema and snow flake schema. List any two methods for data normalization. [5]
Basis Star Schema Snowflake Schema --------- Structure A central fact table is connected directly to dimension tables, forming a star shape A variant of star schema where some dimension tables are further normalized and split into additi...
- 105 marksK-fold cross validationHideAnswer
How do you evaluate the accuracy of a classifier? Discuss the advantages of using K-fold cross validation. [5]
The accuracy of a classifier measures how well it correctly predicts the class label of unseen tuples. The common methods are: - The dataset D is randomly partitioned into two independent sets: - Training set (typically 2/3 of data): use...
- 115 marksNumericalClustering techniquesHideAnswer
Apply K(=2)-Means algorithm over the data (185, 72), (170, 56), (168, 60), (179, 68), (182, 72), (188, 77) up to two iterations and show the clusters. Initially choose first two objects as initial centroids. [5]
Point Data ------------- P1 (185, 72) P2 (170, 56) P3 (168, 60) P4 (179, 68) P5 (182, 72) P6 (188, 77) Initial Centroids: $C1 = (185, 72)$, $C2 = (170, 56)$ Distance metric: Euclidean, $d = \sqrt{(x2-x1)^2 + (y2-y1)^2}$ --- Distances to ...
- 125 marksData discretization and Concept Hierarchy HideAnswer
Define data discretization. Describe the tasks for data preprocessing. [5]
Data Discretization and Data Preprocessing Tasks
Part 1: Data Discretization (Definition)
Data discretization is a technique of data preprocessing that transforms continuous (numeric) data into discrete intervals or categories. Instead of working with raw continuous values, the data is divided into a finite number of intervals or bins, and each interval is assigned a label or concept.
For example, a continuous attribute like age (values: 1 to 100) can be discretized into categories such as:
- Young: 1-30
- Middle-aged: 31-60
- Senior: 61-100
Discretization reduces the number of distinct values for a continuous attribute, making data easier to handle in mining tasks like classification and association analysis.
Part 2: Tasks for Data Preprocessing
Raw data collected from the environment can be dirty - meaning it is incomplete, noisy, or inconsistent. Poor quality data leads to poor quality mining results. Therefore, data preprocessing is a critical step in data mining.
The main tasks (categories) of data preprocessing are:
1. Data Cleaning
- The process of removing noise and inconsistent data from the dataset.
- Handles missing values, smooths noisy data, identifies or removes outliers, and resolves inconsistencies.
- Example: Filling in missing age values with the mean age of the dataset.
2. Data Integration
- The process of combining data from multiple sources (databases, files, data cubes) into a coherent data store.
- Handles issues like redundancy, duplicate records, and conflicting attribute names.
- Example: Merging customer data from two different databases into one unified dataset.
3. Data Transformation
- Data is transformed into forms appropriate for mining.
- Includes operations such as:
- Normalization: Scaling values to a small range (e.g., 0.0 to 1.0)
- Aggregation: Summarizing data
- Generalization: Replacing low-level data with higher-level concepts
- Example: Normalizing salary values so they fall between 0 and 1.
4. Data Reduction
- The process of reducing the volume of data while maintaining analytical integrity.
- Produces a reduced representation of the dataset that is much smaller in size yet produces the same (or nearly the same) analytical results.
- Techniques include: dimensionality reduction, numerosity reduction, and data compression.
- Example: Removing irrelevant attributes that do not contribute to the mining task.
5. Data Discretization
- Transforms continuous attributes into discrete intervals or categories as defined above.
- Reduces the number of distinct values for continuous attributes.
- Particularly useful for algorithms that require categorical data.
- Example: Converting income (continuous) into Low, Medium, High categories.
Summary Table
Task Purpose Data Cleaning Remove noise and inconsistencies Data Integration Combine multiple data sources Data Transformation Convert data into suitable forms Data Reduction Reduce data volume while preserving quality Data Discretization Convert continuous data into discrete intervals
Note: Data preprocessing is one of the most critical steps in data mining because the quality of data directly affects the quality of the mining results.