CSC420 · TU past paper
Data Warehousing and Data Mining 2076 question paper
The complete TU 2076 exam paper for Data Warehousing and Data Mining (CSC420), all 13 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksDifferences between operational database aHideAnswer
Do pattern and information refer to same aspect? Justify. Differentiate between data warehouse and operational database.[10]
--- No, pattern and information do not refer to the same aspect. They are related but distinct concepts in the context of Knowledge Discovery from Data (KDD) and Data Mining. A pattern is a structural regularity, relationship, or trend d...
- 210 marksNumericalLimitation and improving AprioriHideAnswer
List the problems of Apriori algorithm with its possible solutions. Consider the following transaction dataset. What association rules can be found in this set, if the minimum support is 3 and the minimum confidence is 80%?
Transaction_ID Item_List T1 {K, A, D, B} T2 {D, A, C, E, B} T3 {C, A, B, E} T4 {B, A, D} [10]
Apriori Algorithm: Problems, Solutions, and Association Rule Mining
Part 1: Problems of Apriori and Solutions
Problems
- Huge number of candidate itemsets. At each level Apriori generates many candidates. For example, $10^4$ frequent 1-itemsets produce roughly $\binom{10^4}{2}\approx 5\times10^7$ candidate 2-itemsets.
- Multiple database scans. The full database must be scanned once for every level $k$, causing heavy I/O cost.
- Costly pattern matching / support counting. Each candidate must be checked against every transaction.
- High memory usage. Storing all candidates at each iteration is infeasible for large data.
Possible Solutions
Problem Solution Too many candidates Hash-based technique to bucket and prune candidate k-itemsets Repeated scans FP-Growth (only 2 scans, no candidate generation) Large candidate sets Transaction reduction (drop transactions with no frequent itemsets) Memory bottleneck Partitioning the database to fit in memory Full-database cost Sampling a representative subset, then verify Support counting cost Dynamic Itemset Counting (DIC) to reduce passes
Part 2: Association Rule Mining
Given data
- Transactions: T1{K,A,D,B}, T2{D,A,C,E,B}, T3{C,A,B,E}, T4{B,A,D}
- Minimum support count = 3
- Minimum confidence = 80%
- Total transactions = 4
Step 1: L1 (1-itemsets)
Item Count Frequent (>=3)? K 1 No A 4 Yes D 3 Yes B 4 Yes C 2 No E 2 No $L_1 = { {A}, {D}, {B} }$
Step 2: L2
Candidates: {A,D}, {A,B}, {D,B}
Itemset Transactions containing Count Frequent? {A,D} T1,T2,T4 3 Yes {A,B} T1,T2,T3,T4 4 Yes {D,B} T1,T2,T4 3 Yes $L_2 = { {A,D}, {A,B}, {D,B} }$
Step 3: L3
Candidate {A,D,B}. All subsets ({A,D},{A,B},{D,B}) are frequent, so retain.
Itemset Transactions Count Frequent? {A,D,B} T1,T2,T4 3 Yes $L_3 = { {A,D,B} }$
No 4-itemsets possible. Terminate.
Step 4: Association Rules
$$\text{Conf}(X\Rightarrow Y)=\frac{\text{Sup}(X\cup Y)}{\text{Sup}(X)}$$
Support values: A=4, D=3, B=4, {A,D}=3, {A,B}=4, {D,B}=3, {A,D,B}=3.
From {A,D} (sup 3):
Rule Confidence >=80%? A ⇒ D 3/4 = 75% No D ⇒ A 3/3 = 100% Yes From {A,B} (sup 4):
Rule Confidence >=80%? A ⇒ B 4/4 = 100% Yes B ⇒ A 4/4 = 100% Yes From {D,B} (sup 3):
Rule Confidence >=80%? D ⇒ B 3/3 = 100% Yes B ⇒ D 3/4 = 75% No From {A,D,B} (sup 3):
Rule Confidence >=80%? A ⇒ D,B 3/4 = 75% No D ⇒ A,B 3/3 = 100% Yes B ⇒ A,D 3/4 = 75% No A,D ⇒ B 3/3 = 100% Yes A,B ⇒ D 3/4 = 75% No D,B ⇒ A 3/3 = 100% Yes Final Strong Association Rules (support >= 3, confidence >= 80%)
- $D \Rightarrow A$ (100%)
- $A \Rightarrow B$ (100%)
- $B \Rightarrow A$ (100%)
- $D \Rightarrow B$ (100%)
- $D \Rightarrow A,B$ (100%)
- $A,D \Rightarrow B$ (100%)
- $D,B \Rightarrow A$ (100%)
- 310 marksWeb miningHideAnswer
Discuss the types of web mining. Explain why K-means is sensitive to outlier and how does K-Medoid minimize this issue.[10]
--- Web mining is the application of data mining techniques to extract useful knowledge and patterns from the World Wide Web. It involves discovering patterns from web data including web content, web structure, and web usage. Web mining ...
- 45 marksDefinitionHideAnswer
How classification plays significance role in data mining? Explain. [5]
Classification is a supervised data mining technique used to predict the class label (category) of unknown data objects based on a trained model. It maps input data into predefined classes or categories by learning patterns from labeled ...
- 55 marksIssues and ApplicationsHideAnswer
Are the information given by data mining is always useful? What are the issues in data warehousing and data mining? [5]
No, the information given by data mining is not always useful. Data mining may produce patterns or results that are: - Trivial - already known to domain experts and provide no new insight - Irrelevant - not applicable to the business pro...
- 65 marksData warehouse and data warehousingHideAnswer
Explain the four characteristics of data warehouse. [5]
A data warehouse is a repository of information collected from multiple sources that stores historical data and provides support for decision-makers for data analysis and reporting. The four key characteristics of a data warehouse are: -...
- 75 marksGeneral strategies for cube computationHideAnswer
Explain the optimization techniques in data cube computation. [5]
Data cube computation is an essential task in data warehouse implementation. Precomputing all or part of a data cube can greatly reduce response time and enhance OLAP performance. However, it is challenging because it may require huge co...
- 85 marksA multidimensional data modelHideAnswer
How multidimensional data model helps in retrieving information? Explain with suitable example. [5]
A multidimensional data model organizes data around a central theme (such as sales) represented by a fact table. Facts are numerical measures (e.g., sales amount, units sold). Each dimension (e.g., time, location, product) may have an as...
- 95 marksArchitecture of data warehouseHideAnswer
Compare the OLAP servers, ROLAP, MOLAP and HOLAP. [5]
As mentioned in the notes, the middle tier of a data warehouse architecture contains the OLAP Server, which can be implemented using ROLAP, MOLAP, or HOLAP. These are three different approaches to storing and processing multidimensional ...
- 105 marksMotivation for data miningHideAnswer
Give a syntax and example of data mining query language. [5]
A data mining query is a structured specification that allows users to interactively communicate with the data mining system to direct the mining process. Data mining tasks can be specified in the form of a data mining query, which is de...
- 115 marksKDDHideAnswer
Differentiate between KDD and data mining. [5]
Difference Between KDD and Data Mining
Definition
KDD (Knowledge Discovery from Databases) is the overall process of discovering useful, valid, and understandable knowledge from large amounts of data. It is a complete, iterative, and interactive process.
Data Mining is a single step within the KDD process where intelligent methods (algorithms) are applied to extract patterns from data.
As stated in the notes: "Data mining can be viewed as a step within a larger process called KDD."
Key Differences
Basis KDD Data Mining Scope Broad, end-to-end process of knowledge discovery A single core step within the KDD process Nature Iterative sequence of multiple steps Application of intelligent algorithms to find patterns Goal To convert raw data into useful knowledge To extract interesting patterns from preprocessed data Steps Involved Includes cleaning, integration, selection, transformation, mining, evaluation, and presentation Focuses only on applying mining algorithms (association, classification, clustering, etc.) User Involvement Involves user at multiple stages Primarily algorithmic and automated Output Final knowledge presented to users Patterns or models (not yet validated knowledge) Dependency KDD cannot exist without data mining Data mining can be studied independently
KDD Process Steps (from notes)
The KDD process consists of the following iterative steps:
- Data Cleaning: Removing noise and inconsistent data
- Data Integration: Combining multiple data sources
- Data Selection: Retrieving data relevant to the analysis task
- Data Transformation: Transforming data into forms appropriate for mining
- Data Mining: Applying intelligent methods to extract data patterns
- Pattern Evaluation: Identifying truly interesting patterns based on interestingness measures
- Knowledge Presentation: Using visualization and knowledge representation techniques to present mined knowledge to users
Summary
KDD is the complete pipeline from raw data to final knowledge, while data mining is the core engine inside that pipeline responsible for pattern extraction. Without the surrounding KDD steps (cleaning, transformation, evaluation), data mining alone cannot produce reliable and meaningful knowledge.
- 125 marksData warehouse implementationHideAnswer
What does data warehouse tuning mean? Describe the parameters. [5]
Data warehouse tuning refers to the process of optimizing the performance of a data warehouse system so that queries execute faster, data loading (ETL) is more efficient, and overall system resources are utilized effectively. Since a dat...
- 135 marksClassification by decision tree inductionHideAnswer
Write short notes on (Any Two): a. Evolution analysis b. Decision trees c. Text mining d.Classification using Regression [5]
--- Evolution analysis refers to the description and modeling of regularities or trends for objects whose behavior changes over time. - It is one of the core functional modules found in the Data Mining Engine, which is the central compon...