2076

CSC420 · TU past paper

Data Warehousing and Data Mining 2076 question paper

The complete TU 2076 exam paper for Data Warehousing and Data Mining (CSC420), all 13 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksDifferences between operational database aAnswer

    Do pattern and information refer to same aspect? Justify. Differentiate between data warehouse and operational database.[10]

    --- No, pattern and information do not refer to the same aspect. They are related but distinct concepts in the context of Knowledge Discovery from Data (KDD) and Data Mining. A pattern is a structural regularity, relationship, or trend d...

  2. 210 marksNumericalLimitation and improving AprioriAnswer

    List the problems of Apriori algorithm with its possible solutions. Consider the following transaction dataset. What association rules can be found in this set, if the minimum support is 3 and the minimum confidence is 80%?

    Transaction_IDItem_List
    T1{K, A, D, B}
    T2{D, A, C, E, B}
    T3{C, A, B, E}
    T4{B, A, D}

    [10]

    Apriori Algorithm: Problems, Solutions, and Association Rule Mining

    Part 1: Problems of Apriori and Solutions

    Problems

    1. Huge number of candidate itemsets. At each level Apriori generates many candidates. For example, $10^4$ frequent 1-itemsets produce roughly $\binom{10^4}{2}\approx 5\times10^7$ candidate 2-itemsets.
    2. Multiple database scans. The full database must be scanned once for every level $k$, causing heavy I/O cost.
    3. Costly pattern matching / support counting. Each candidate must be checked against every transaction.
    4. High memory usage. Storing all candidates at each iteration is infeasible for large data.

    Possible Solutions

    ProblemSolution
    Too many candidatesHash-based technique to bucket and prune candidate k-itemsets
    Repeated scansFP-Growth (only 2 scans, no candidate generation)
    Large candidate setsTransaction reduction (drop transactions with no frequent itemsets)
    Memory bottleneckPartitioning the database to fit in memory
    Full-database costSampling a representative subset, then verify
    Support counting costDynamic Itemset Counting (DIC) to reduce passes

    Part 2: Association Rule Mining

    Given data

    • Transactions: T1{K,A,D,B}, T2{D,A,C,E,B}, T3{C,A,B,E}, T4{B,A,D}
    • Minimum support count = 3
    • Minimum confidence = 80%
    • Total transactions = 4

    Step 1: L1 (1-itemsets)

    ItemCountFrequent (>=3)?
    K1No
    A4Yes
    D3Yes
    B4Yes
    C2No
    E2No

    $L_1 = { {A}, {D}, {B} }$

    Step 2: L2

    Candidates: {A,D}, {A,B}, {D,B}

    ItemsetTransactions containingCountFrequent?
    {A,D}T1,T2,T43Yes
    {A,B}T1,T2,T3,T44Yes
    {D,B}T1,T2,T43Yes

    $L_2 = { {A,D}, {A,B}, {D,B} }$

    Step 3: L3

    Candidate {A,D,B}. All subsets ({A,D},{A,B},{D,B}) are frequent, so retain.

    ItemsetTransactionsCountFrequent?
    {A,D,B}T1,T2,T43Yes

    $L_3 = { {A,D,B} }$

    No 4-itemsets possible. Terminate.

    Step 4: Association Rules

    $$\text{Conf}(X\Rightarrow Y)=\frac{\text{Sup}(X\cup Y)}{\text{Sup}(X)}$$

    Support values: A=4, D=3, B=4, {A,D}=3, {A,B}=4, {D,B}=3, {A,D,B}=3.

    From {A,D} (sup 3):

    RuleConfidence>=80%?
    A ⇒ D3/4 = 75%No
    D ⇒ A3/3 = 100%Yes

    From {A,B} (sup 4):

    RuleConfidence>=80%?
    A ⇒ B4/4 = 100%Yes
    B ⇒ A4/4 = 100%Yes

    From {D,B} (sup 3):

    RuleConfidence>=80%?
    D ⇒ B3/3 = 100%Yes
    B ⇒ D3/4 = 75%No

    From {A,D,B} (sup 3):

    RuleConfidence>=80%?
    A ⇒ D,B3/4 = 75%No
    D ⇒ A,B3/3 = 100%Yes
    B ⇒ A,D3/4 = 75%No
    A,D ⇒ B3/3 = 100%Yes
    A,B ⇒ D3/4 = 75%No
    D,B ⇒ A3/3 = 100%Yes

    Final Strong Association Rules (support >= 3, confidence >= 80%)

    1. $D \Rightarrow A$ (100%)
    2. $A \Rightarrow B$ (100%)
    3. $B \Rightarrow A$ (100%)
    4. $D \Rightarrow B$ (100%)
    5. $D \Rightarrow A,B$ (100%)
    6. $A,D \Rightarrow B$ (100%)
    7. $D,B \Rightarrow A$ (100%)
  3. 310 marksWeb miningAnswer

    Discuss the types of web mining. Explain why K-means is sensitive to outlier and how does K-Medoid minimize this issue.[10]

    --- Web mining is the application of data mining techniques to extract useful knowledge and patterns from the World Wide Web. It involves discovering patterns from web data including web content, web structure, and web usage. Web mining ...

  4. 45 marksDefinitionAnswer

    How classification plays significance role in data mining? Explain. [5]

    Classification is a supervised data mining technique used to predict the class label (category) of unknown data objects based on a trained model. It maps input data into predefined classes or categories by learning patterns from labeled ...

  5. 55 marksIssues and ApplicationsAnswer

    Are the information given by data mining is always useful? What are the issues in data warehousing and data mining? [5]

    No, the information given by data mining is not always useful. Data mining may produce patterns or results that are: - Trivial - already known to domain experts and provide no new insight - Irrelevant - not applicable to the business pro...

  6. 65 marksData warehouse and data warehousingAnswer

    Explain the four characteristics of data warehouse. [5]

    A data warehouse is a repository of information collected from multiple sources that stores historical data and provides support for decision-makers for data analysis and reporting. The four key characteristics of a data warehouse are: -...

  7. 75 marksGeneral strategies for cube computationAnswer

    Explain the optimization techniques in data cube computation. [5]

    Data cube computation is an essential task in data warehouse implementation. Precomputing all or part of a data cube can greatly reduce response time and enhance OLAP performance. However, it is challenging because it may require huge co...

  8. 85 marksA multidimensional data modelAnswer

    How multidimensional data model helps in retrieving information? Explain with suitable example. [5]

    A multidimensional data model organizes data around a central theme (such as sales) represented by a fact table. Facts are numerical measures (e.g., sales amount, units sold). Each dimension (e.g., time, location, product) may have an as...

  9. 95 marksArchitecture of data warehouseAnswer

    Compare the OLAP servers, ROLAP, MOLAP and HOLAP. [5]

    As mentioned in the notes, the middle tier of a data warehouse architecture contains the OLAP Server, which can be implemented using ROLAP, MOLAP, or HOLAP. These are three different approaches to storing and processing multidimensional ...

  10. 105 marksMotivation for data miningAnswer

    Give a syntax and example of data mining query language. [5]

    A data mining query is a structured specification that allows users to interactively communicate with the data mining system to direct the mining process. Data mining tasks can be specified in the form of a data mining query, which is de...

  11. 115 marksKDDAnswer

    Differentiate between KDD and data mining. [5]

    Difference Between KDD and Data Mining

    Definition

    KDD (Knowledge Discovery from Databases) is the overall process of discovering useful, valid, and understandable knowledge from large amounts of data. It is a complete, iterative, and interactive process.

    Data Mining is a single step within the KDD process where intelligent methods (algorithms) are applied to extract patterns from data.

    As stated in the notes: "Data mining can be viewed as a step within a larger process called KDD."


    Key Differences

    BasisKDDData Mining
    ScopeBroad, end-to-end process of knowledge discoveryA single core step within the KDD process
    NatureIterative sequence of multiple stepsApplication of intelligent algorithms to find patterns
    GoalTo convert raw data into useful knowledgeTo extract interesting patterns from preprocessed data
    Steps InvolvedIncludes cleaning, integration, selection, transformation, mining, evaluation, and presentationFocuses only on applying mining algorithms (association, classification, clustering, etc.)
    User InvolvementInvolves user at multiple stagesPrimarily algorithmic and automated
    OutputFinal knowledge presented to usersPatterns or models (not yet validated knowledge)
    DependencyKDD cannot exist without data miningData mining can be studied independently

    KDD Process Steps (from notes)

    The KDD process consists of the following iterative steps:

    1. Data Cleaning: Removing noise and inconsistent data
    2. Data Integration: Combining multiple data sources
    3. Data Selection: Retrieving data relevant to the analysis task
    4. Data Transformation: Transforming data into forms appropriate for mining
    5. Data Mining: Applying intelligent methods to extract data patterns
    6. Pattern Evaluation: Identifying truly interesting patterns based on interestingness measures
    7. Knowledge Presentation: Using visualization and knowledge representation techniques to present mined knowledge to users

    Summary

    KDD is the complete pipeline from raw data to final knowledge, while data mining is the core engine inside that pipeline responsible for pattern extraction. Without the surrounding KDD steps (cleaning, transformation, evaluation), data mining alone cannot produce reliable and meaningful knowledge.

  12. 125 marksData warehouse implementationAnswer

    What does data warehouse tuning mean? Describe the parameters. [5]

    Data warehouse tuning refers to the process of optimizing the performance of a data warehouse system so that queries execute faster, data loading (ETL) is more efficient, and overall system resources are utilized effectively. Since a dat...

  13. 135 marksClassification by decision tree inductionAnswer

    Write short notes on (Any Two): a. Evolution analysis b. Decision trees c. Text mining d.Classification using Regression [5]

    --- Evolution analysis refers to the description and modeling of regularities or trends for objects whose behavior changes over time. - It is one of the core functional modules found in the Data Mining Engine, which is the central compon...