2082.1

BIT454 · TU past paper

Data Warehousing and Data Mining 2082.1 question paper

The complete TU 2082.1 exam paper for Data Warehousing and Data Mining (BIT454), all 12 questions with solved model answers written to the mark scheme.

Past Papers2082.12081

Tap a question to open its answer.

  1. 110 marksData mining goals and objectivesAnswer

    List some of the data mining goals.Explain the components of data warehouse.[3+7]

    (a) Data Mining Goals The main goals of data mining are: 1. Prediction - Predicting future values or trends based on historical data patterns. For example, predicting customer churn or sales forecasts. 2. Classification - Assigning data ...

  2. 210 marksMarket basket analysis conceptAnswer

    What is the concept behind market basket analysis?How do you generate frequent item sets using Apriori algorithm? Explain.[4+6]

    Market Basket Analysis and Apriori Algorithm

    (a) Concept Behind Market Basket Analysis

    Market Basket Analysis is a data mining technique used to discover associations and relationships between items purchased together by customers.

    Key Concepts:

    1. Purpose:

    • Identifies which products are frequently bought together
    • Discovers customer purchasing patterns and behavior
    • Helps in product placement, cross-selling, and promotional strategies

    2. Core Idea:

    • Analyzes transaction data to find item combinations
    • Determines how often certain items co-occur in transactions
    • Establishes rules like "if customer buys X, they likely buy Y"

    3. Business Applications:

    • Supermarket shelf arrangement (placing related items near each other)
    • Recommendation systems (suggesting complementary products)
    • Bundle pricing and promotional campaigns
    • Inventory management and stock optimization

    4. Association Rules:

    • Expressed as: If {Item A} then {Item B}
    • Measured using metrics like support, confidence, and lift
    • Helps predict customer behavior and increase sales

    (b) Generating Frequent Itemsets Using Apriori Algorithm

    Algorithm Overview:

    The Apriori Algorithm generates frequent itemsets by iteratively finding itemsets that appear in at least a minimum support threshold of transactions.

    Key Principle:

    "If an itemset is frequent, then all of its subsets must also be frequent" (Apriori property)

    Steps:

    Step 1: Set Minimum Support Threshold

    • Define minimum support (e.g., 50% or 2 transactions out of 4)
    • Only itemsets meeting this threshold are considered frequent

    Step 2: Generate 1-Itemsets (L₁)

    • Count occurrences of each individual item
    • Keep items with support ≥ minimum support
    • Example: {A}, {B}, {C}, {D}

    Step 3: Generate Candidate k-Itemsets (Cₖ)

    • Join frequent (k-1)-itemsets to create candidate k-itemsets
    • Apply pruning: remove candidates containing infrequent subsets
    • Example: From {A}, {B}, {C} → candidates {A,B}, {A,C}, {B,C}

    Step 4: Count Support for Candidates

    • Scan database to count occurrences of each candidate itemset
    • Calculate support = (frequency of itemset) / (total transactions)

    Step 5: Filter by Minimum Support

    • Keep only candidates with support ≥ threshold
    • These become frequent k-itemsets (Lₖ)

    Step 6: Repeat Until No New Itemsets

    • Continue with k+1 until no new frequent itemsets are generated

    Example:

    TransactionItems
    T1{A, B, C}
    T2{A, B}
    T3{A, C}
    T4{B, C}

    With minimum support = 50% (2 transactions):

    • L₁: {A}(3), {B}(3), {C}(3) - all frequent
    • C₂: {A,B}, {A,C}, {B,C}
    • L₂: {A,B}(2), {A,C}(2), {B,C}(2) - all frequent
    • C₃: {A,B,C}
    • L₃: {A,B,C}(1) - NOT frequent (support < 50%)

    Result: Frequent itemsets are {A}, {B}, {C}, {A,B}, {A,C}, {B,C}

    Advantages:

    • Efficient pruning reduces search space
    • Guarantees finding all frequent itemsets
    • Widely used in practice
  3. 310 marksGini index for attribute selectionAnswer

    How can you use Gini index as attribute selection algorithm? Illustrate with an example.Describe the working mechanism of support vector machine.[5+5]

    (a) Gini Index as Attribute Selection Algorithm The Gini index (or Gini impurity) measures the probability of incorrectly classifying a randomly chosen element if it were randomly labeled according to the class distribution in a dataset....

  4. 45 marksSimilarity measures for ordinal dataAnswer

    How do you find similarities between ordinal data attributes? Explain. [5]

    Ordinal data are categorical attributes with a natural ordering or ranking among the categories. Examples include: education level (high school < bachelor < master < PhD), satisfaction ratings (poor < fair < good < excellent), or grades ...

  5. 55 marksFull cube and closed cube conceptsAnswer

    Differentiate between full cube and closed cube. [5]

    A full cube is a 3D representation of a cube where all six faces are visible in the isometric or orthographic projection. It shows the complete structure of the cube with: - Top face visible - Front face visible - Right side face visible...

  6. 65 marksLazy learners and instance-based learningAnswer

    Discuss about lazy learners and ensemble method. [5]

    Definition: Lazy learners are machine learning algorithms that defer the learning process until a query (prediction request) is made. They store the entire training dataset and perform minimal processing during the training phase. Charac...

  7. 75 marksTypes of data in clusteringAnswer

    Explain the types of data in clustering. [5]

    Since reference notes are not provided, this answer draws from standard clustering theory in computer science: - Consists of continuous or discrete numeric values - Examples: age, salary, temperature, distance measurements - Can be direc...

  8. 85 marksHigh dimensional data clusteringAnswer

    What is outlier? How do you cluster high dimensional data? [2+3]

    (a) What is an Outlier? An outlier is a data point that deviates significantly from the normal pattern or distribution of the rest of the dataset. Key characteristics: - It is substantially different from other observations - It lies far...

  9. 95 marksViral marketing conceptAnswer

    Define viral marketing. Describe densification power law. [2+3]

    Model Answer: Viral Marketing and Densification Power Law

    1. Viral Marketing (2 marks)

    Definition:

    Viral marketing is a marketing strategy that leverages social networks and word-of-mouth communication to promote products, services, or content. It relies on users voluntarily sharing information with their networks, causing exponential growth in message reach and brand awareness with minimal paid advertising.

    Key characteristics:

    • Self-replicating and spreads organically through networks
    • Users become active promoters by sharing content
    • Achieves rapid, widespread dissemination at low cost
    • Often uses emotionally engaging or novel content

    2. Densification Power Law (3 marks)

    Definition and Description:

    The densification power law is an empirical observation about the growth patterns of real-world networks (social networks, web graphs, etc.). It states that the number of edges in a network grows superlinearly with the number of nodes over time.

    Mathematical Expression:

    $$E(t) \propto N(t)^a$$

    where:

    • E(t) = number of edges at time t
    • N(t) = number of nodes at time t
    • a = densification exponent (typically 1 < a < 2)

    Key Implications:

    1. Superlinear Growth: As networks grow, they become denser (average degree increases)
    2. Exponent Range: For most real networks, a ≈ 1.1 to 1.6
    3. Network Evolution: Networks don't just add nodes; they add edges at a faster rate, meaning nodes form more connections over time
    4. Practical Significance: Explains why social networks become increasingly interconnected as they scale

    This law has been observed empirically in online social networks, citation networks, and the World Wide Web.

  10. 105 marksAgglomerative clustering approachAnswer

    Differentiate between agglomerative and divisive clustering approach. [5]

    Agglomerative clustering is a bottom-up hierarchical clustering approach that works as follows: - Starts with each data point as an individual cluster - Iteratively merges the two closest/most similar clusters into a single cluster - Con...

  11. 115 marksTypes of network analysisAnswer

    What type of analysis can be performed in social network? Explain. [5]

    Social network analysis (SNA) can perform several key types of analysis: - Examines the overall structure and patterns of connections within the network - Identifies clusters, communities, and subgroups of interconnected nodes - Analyzes...

  12. 125 marksWeb structure miningAnswer

    Discuss about web structure, web content and web usage mining. [5]

    Web mining is the application of data mining techniques to discover patterns and knowledge from web data. It comprises three main categories: --- Definition: Web structure mining discovers knowledge from the organization and links of the...