2081

CSC420 · TU past paper

Data Warehousing and Data Mining 2081 question paper

The complete TU 2081 exam paper for Data Warehousing and Data Mining (CSC420), all 12 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksA multidimensional data modelAnswer

    When do we prefer trim mean for statistical description of data? Justify with an example. Describe about multi-dimensional data model and conceptual modeling of data warehouse.[10]

    Trim Mean for Statistical Description & Multidimensional Data Model / Conceptual Modeling of Data Warehouse


    PART 1: Trim Mean - When to Prefer It and Justification

    Definition of Trim Mean

    The trimmed mean (trim mean) is a measure of central tendency calculated by removing a specified percentage of the smallest and largest values from a dataset before computing the arithmetic mean. If we trim p% from each end, it is called a p% trimmed mean.

    Formula:

    $$\bar{x}{trim} = \frac{1}{n - 2k} \sum{i=k+1}^{n-k} x_{(i)}$$

    where:

    • $x_{(i)}$ are the sorted (ordered) values
    • $k = \lfloor p \times n / 100 \rfloor$ is the number of values trimmed from each end
    • $n$ is the total number of observations

    When Do We Prefer Trim Mean?

    We prefer the trimmed mean over the ordinary arithmetic mean in the following situations:

    1. Presence of Outliers: When the dataset contains extreme values (outliers) that distort the arithmetic mean, the trim mean provides a more robust and representative central value.

    2. Skewed Distributions: When data is heavily skewed (positively or negatively), the ordinary mean is pulled toward the tail. The trim mean reduces this effect.

    3. Noisy or Erroneous Data: In real-world data collection, some values may be recorded incorrectly or represent measurement errors. Trimming removes these suspicious extreme values.

    4. When Median is Too Extreme: The median ignores all values except the middle one, while the arithmetic mean is too sensitive to extremes. The trim mean offers a balance between the two - it is more robust than the mean but uses more data than the median.

    5. Data Mining and Statistical Description: In data preprocessing for data mining, when describing large datasets with potential noise, the trim mean gives a cleaner summary statistic.


    Justification with Example

    Example:

    Consider the monthly salaries (in thousands) of 10 employees:

    $${12, 14, 15, 16, 17, 18, 19, 20, 21, 150}$$

    Step 1: Compute the Arithmetic Mean

    $$\bar{x} = \frac{12 + 14 + 15 + 16 + 17 + 18 + 19 + 20 + 21 + 150}{10} = \frac{302}{10} = 30.2$$

    The mean is 30.2, but 9 out of 10 employees earn between 12 and 21. The value 150 is clearly an outlier (perhaps the CEO's salary), and it has inflated the mean significantly.

    Step 2: Compute the 10% Trimmed Mean

    • Trim 10% from each end: $k = 0.10 \times 10 = 1$ value removed from each side
    • Remove the lowest value: 12
    • Remove the highest value: 150
    • Remaining values: ${14, 15, 16, 17, 18, 19, 20, 21}$

    $$\bar{x}_{10%} = \frac{14 + 15 + 16 + 17 + 18 + 19 + 20 + 21}{8} = \frac{140}{8} = 17.5$$

    Conclusion: The trimmed mean of 17.5 is far more representative of the typical employee salary than the arithmetic mean of 30.2. This justifies using the trim mean when outliers are present.


    PART 2: Multidimensional Data Model

    Definition

    A multidimensional data model is a data model that organizes data in multiple dimensions to facilitate complex analytical queries and reporting. It is the foundation of OLAP (Online Analytical Processing) systems and data warehouses.

    • Data is viewed as a data cube with multiple dimensions.
    • Each dimension represents a perspective or entity of interest (e.g., Time, Location, Product).
    • The model is typically organized around a central theme (e.g., Sales), represented by a fact table.
    • Facts are numerical measures such as sales amount, units sold, profit, etc.
    • Each dimension may have a dimension table associated with it.

    Key Components

    ComponentDescription
    Fact TableCentral table containing numerical measures (facts)
    Dimension TableTables describing the dimensions (e.g., Time, Product, Location)
    Measures/FactsQuantitative data: sales, revenue, units sold
    DimensionsQualitative context: who, what, when, where

    OLAP Operations in Multidimensional Data Model

    1. Roll-Up (Consolidation / Aggregation)

    • Generates summary from lower level to higher level.
    • Can be performed by:
      • Reducing dimensions
      • Climbing up a concept hierarchy
    • Concept Hierarchy: A system of grouping things based on their order level. For example, from daily summary, weekly summary is generated; from weekly, monthly summary is generated.
    • Example: Roll-up on location from cities to countries - individual city sales are aggregated into country-level totals.

    2. Drill-Down

    • The reverse of roll-up; navigates from summary data to more detailed data.
    • Example: From yearly sales, drill down to quarterly, then monthly sales.

    3. Slice

    • Selects one particular dimension from a data cube, resulting in a sub-cube.
    • Example: Slice on Time = "Q1 2024" to get all sales data for Q1 only.

    4. Dice

    • Selects two or more dimensions from a data cube to produce a sub-cube.
    • Example: Dice on Location = {Nepal, India} AND Time = {Q1, Q2}.

    5. Pivot (Rotate)

    • Rotates the data axes to provide an alternative presentation of data.
    • Example: Swapping rows and columns in a cross-tabulation.

    PART 3: Conceptual Modeling of Data Warehouse

    Definition

    A **conceptual

  2. 210 marksNumericalFinding frequent itemsetAnswer

    How do you generate strong association rules? From the following dataset find the frequent item set using FP growth algorithm using 3 as minimum support.

    Transaction IDItems
    T1{K, E, M, O, Y}
    T2{K, E, O, Y}
    T3{K, E, M}
    T4{K, M, Y}
    T5{K, E, O}

    [10]

    Strong Association Rules and FP-Growth

    Given Data

    Transactions:

    • T1: {K, E, M, O, Y}
    • T2: {K, E, O, Y}
    • T3: {K, E, M}
    • T4: {K, M, Y}
    • T5: {K, E, O}

    Minimum support = 3

    Part 1: Generating Strong Association Rules

    A strong association rule $X \Rightarrow Y$ satisfies both thresholds:

    • $\text{support}(X \cup Y) \geq \text{min_support}$
    • $\text{confidence}(X \Rightarrow Y) = \dfrac{\text{support}(X \cup Y)}{\text{support}(X)} \geq \text{min_confidence}$

    Procedure:

    1. Find all frequent itemsets satisfying min support (via Apriori or FP-Growth).
    2. For each frequent itemset $l$, generate every non-empty proper subset $s$.
    3. For each such $s$, form the rule $s \Rightarrow (l - s)$ and compute confidence $= \dfrac{\text{support}(l)}{\text{support}(s)}$.
    4. Output the rule only if confidence $\geq$ min_confidence. These are the strong rules.

    Part 2: FP-Growth (min_support = 3)

    Step 1: Item frequency (1st scan)

    ItemCount
    K5
    E4
    M3
    O3
    Y3

    All satisfy min support. Order (descending, ties alphabetical): K:5, E:4, M:3, O:3, Y:3

    Step 2: Reordered transactions

    TIDOrdered
    T1K, E, M, O, Y
    T2K, E, O, Y
    T3K, E, M
    T4K, M, Y
    T5K, E, O

    Step 3: FP-Tree

    null
     └── K:5
          ├── E:4
          │    ├── M:1
          │    │    └── O:1
          │    │         └── Y:1
          │    └── O:2
          │         └── Y:1
          └── M:2
               └── Y:1
    

    Note on M placement:

    • T1 (KEMOY) and T3 (KEM) go under E-branch: E→M appears twice (M:2 under E).
    • T4 (KMY) goes under K directly: K→M (M:1 under K, no E).

    So M under E has count 2, and M under K (no E) has count 1. Total M = 3. ✓

    Corrected FP-Tree:

    null
     └── K:5
          ├── E:4
          │    ├── M:2   (from T1, T3)
          │    │    └── O:1   (T1)
          │    │         └── Y:1
          │    └── O:2   (T2, T5)
          │         └── Y:1   (T2)
          └── M:1   (T4)
               └── Y:1
    

    Take care with the M counts: the correct placement is M:2 under E and M:1 under K, not the other way round.

    Header table:

    ItemSupportNode links
    K5K:5
    E4E:4
    M3M:2, M:1
    O3O:1, O:2
    Y3Y:1, Y:1, Y:1

    Step 4: Mining (bottom-up)

    Item Y (support 3): Conditional pattern base:

    • {K,E,M,O}:1 (T1)
    • {K,E,O}:1 (T2)
    • {K,M}:1 (T4)

    Counts: K=3, E=2, M=2, O=2. Only K:3 ≥ 3. Frequent: {K,Y}:3

    Item O (support 3): Conditional pattern base:

    • {K,E,M}:1 (T1)
    • {K,E}:2 (T2, T5)

    Counts: K=3, E=3, M=1. Survive: K:3, E:3. Conditional FP-tree: K→E (3). Frequent: {K,O}:3, {E,O}:3, {K,E,O}:3

    Item M (support 3): Conditional pattern base:

    • {K,E}:2 (T1, T3)
    • {K}:1 (T4)

    Counts: K=3, E=2. Only K:3 survives. Frequent: {K,M}:3

    Item E (support 4): Conditional pattern base:

    • {K}:4

    Frequent: {K,E}:4

    Item K: root, no conditional patterns.

    Final Frequent Itemsets (support ≥ 3)

    1-itemsets:

    • {K}:5, {E}:4, {M}:3, {O}:3, {Y}:3

    2-itemsets:

    • {K,E}:4
    • {K,M}:3
    • {K,O}:3
    • {E,O}:3
    • {K,Y}:3

    3-itemsets:

    • {K,E,O}:3

    The maximal/most useful frequent itemset is {K, E, O} with support 3.

  3. 310 marksNumericalID3 as attribute selection algorithmAnswer

    Decision Tree Classification with ID3 Algorithm

    Overfitting and Underfitting + ID3 Decision Tree

    Step 1 - EXTRACT: Given Data

    TIDAgeCar TypeClass
    1≤30FamilyHigh
    2≤30SportsHigh
    3>30SportsHigh
    4>30FamilyLow
    5>30TruckLow
    6≤30FamilyHigh
    • Total tuples = 6
    • Class High: TID 1, 2, 3, 6 → 4 tuples
    • Class Low: TID 4, 5 → 2 tuples
    • Attributes: Age (≤30, >30), Car Type (Family, Sports, Truck)

    Part 1: Definitions

    Overfitting

    Overfitting occurs when a model learns the training data too well, capturing noise and random fluctuations rather than the underlying pattern. It shows high accuracy on training data but poor accuracy on unseen/test data. In decision trees, this happens when the tree grows too deep.

    Underfitting

    Underfitting occurs when a model is too simple to capture the underlying structure of the data. It shows poor accuracy on both training and test data. In decision trees, this happens when the tree is too shallow.


    Part 2: ID3 Algorithm

    Step 1: Entropy of the whole dataset

    $$Entropy(S) = -\frac{4}{6}\log_2\frac{4}{6} - \frac{2}{6}\log_2\frac{2}{6}$$

    $$= -\tfrac{2}{3}(-0.585) - \tfrac{1}{3}(-1.585) = 0.390 + 0.528 = 0.918$$

    Step 2: Gain for Age

    Age = ≤30: TID 1, 2, 6 → High:3, Low:0 → $Entropy = 0$

    Age = >30: TID 3, 4, 5 → High:1, Low:2

    $$Entropy(>30) = -\tfrac{1}{3}\log_2\tfrac{1}{3} - \tfrac{2}{3}\log_2\tfrac{2}{3} = 0.528 + 0.390 = 0.918$$

    $$Info_{Age}(S) = \tfrac{3}{6}(0) + \tfrac{3}{6}(0.918) = 0.459$$

    $$Gain(Age) = 0.918 - 0.459 = 0.459$$

    Step 3: Gain for Car Type

    Family: TID 1, 4, 6 → High:2, Low:1 → $Entropy = 0.918$ Sports: TID 2, 3 → High:2, Low:0 → $Entropy = 0$ Truck: TID 5 → High:0, Low:1 → $Entropy = 0$

    $$Info_{CarType}(S) = \tfrac{3}{6}(0.918) + \tfrac{2}{6}(0) + \tfrac{1}{6}(0) = 0.459$$

    $$Gain(CarType) = 0.918 - 0.459 = 0.459$$

    Step 4: Select Root

    AttributeGain
    Age0.459
    Car Type0.459

    Both are tied at 0.459. Selecting Age as root (either is valid).

    Step 5: Split on Age

                    Age
                  /      \
               ≤30        >30
            [1,2,6]      [3,4,5]
           H:3,L:0      H:1,L:2
           PURE→High    impure
    
    • Age ≤30: all High → leaf Class = High
    • Age >30: TID 3 (High), 4 (Low), 5 (Low) → not pure, split further.

    Step 6: Split the (Age > 30) subset

    Remaining attribute: Car Type. Subset = {3, 4, 5}.

    • Sports: TID 3 → High → pure → High
    • Family: TID 4 → Low → pure → Low
    • Truck: TID 5 → Low → pure → Low

    Entropy of each subset = 0, so Car Type perfectly classifies this subset.

    Final Decision Tree

                         Age
                      /        \
                  ≤30            >30
                   |              |
                [High]         Car Type
                             /    |     \
                        Sports  Family  Truck
                           |       |       |
                        [High]  [Low]   [Low]
    

    Classification Rules

    1. If Age ≤ 30 → High
    2. If Age > 30 AND Car = Sports → High
    3. If Age > 30 AND Car = Family → Low
    4. If Age > 30 AND Car = Truck → Low

    This tree classifies all 6 training tuples correctly.


    The completed tree splits the (Age > 30) subset on Car Type, with entropy = 0.918, both gains = 0.459 and Age chosen as the root, the left branch being pure High.

  4. 45 marksNumericalClustering techniquesAnswer

    Using k-means++ algorithm and Euclidean distance, find the initial 3 cluster centroids from A1 = (3, 11), A2 = (3, 6), A3 = (9, 5), A4 = (6, 9), A6 = (7, 5), A7 = (2, 3), A8 = (5, 10). Choose (3, 11) as one of the initial centroids. [5]

    K-Means++ Initial Centroid Selection

    Step 1 - Given Data

    Data points:

    PointCoordinates
    A1(3, 11)
    A2(3, 6)
    A3(9, 5)
    A4(6, 9)
    A6(7, 5)
    A7(2, 3)
    A8(5, 10)

    Number of clusters: $k = 3$ First centroid (given): $C_1 = (3, 11)$

    Step 2 - Solve

    K-Means++ Rule

    1. First centroid is fixed as $C_1 = (3,11)$.
    2. For each point compute $D^2$ = squared Euclidean distance to the nearest chosen centroid.
    3. Choose the point with the largest $D^2$ (highest selection probability) as the next centroid.

    Selecting $C_2$: distances to $C_1 = (3,11)$

    $$D^2 = (x-3)^2 + (y-11)^2$$

    PointCalculation$D^2$
    A2 (3,6)$0 + 25$25
    A3 (9,5)$36 + 36$72
    A4 (6,9)$9 + 4$13
    A6 (7,5)$16 + 36$52
    A7 (2,3)$1 + 64$65
    A8 (5,10)$4 + 1$5

    Total $= 25+72+13+52+65+5 = 232$

    Probability $= D^2 / 232$:

    PointP
    A20.108
    A30.310
    A60.224
    A70.280
    A40.056
    A80.022

    Largest $D^2$ is A3 (72).

    $$\boxed{C_2 = (9, 5)}$$

    Selecting $C_3$: nearest distance to ${C_1, C_2}$

    $D^2(C_2)$ with $C_2 = (9,5)$:

    Point$D^2(C_1)$$D^2(C_2)$$\min$
    A2 (3,6)25$36+1=37$25
    A4 (6,9)13$9+16=25$13
    A6 (7,5)52$4+0=4$4
    A7 (2,3)65$49+4=53$53
    A8 (5,10)5$16+25=41$5

    Total $= 25+13+4+53+5 = 100$

    Probability $= \min D^2 / 100$:

    PointP
    A20.25
    A40.13
    A60.04
    A70.53
    A80.05

    Largest $\min D^2$ is A7 (53).

    $$\boxed{C_3 = (2, 3)}$$

    Final Result

    The 3 initial centroids chosen by K-Means++ are:

    CentroidCoordinates
    $C_1$(3, 11)
    $C_2$(9, 5)
    $C_3$(2, 3)
  5. 55 marksGeneral strategies for cube computationAnswer

    Explain the general strategies for cube computation. [5]

    Data cube computation is an essential task in data warehouse implementation. The precomputation of all or part of a data cube can greatly reduce response time and enhance the performance of OLAP. However, it is challenging because it may...

  6. 65 marksAttribute oriented induction for data charAnswer

    Distinguish between data characterization and data discrimination. What are the challenges of multimedia mining? [5]

    --- Both are descriptive data mining functions used to summarize and compare data. They are distinguished as follows: Aspect Data Characterization Data Discrimination --------- Definition Summarization of the general characteristics or f...

  7. 75 marksSigned networkAnswer

    Define graph mining. Discuss the conflict between theory of balance and theory of status. [5]

    Graph Mining and Conflict Between Theory of Balance and Theory of Status


    Part 1: Definition of Graph Mining (1 mark)

    Graph Mining is a subfield of data mining that focuses on discovering useful patterns, structures, and knowledge from graph-structured data. It applies data mining techniques to graphs, where data is represented as nodes (vertices) representing entities (such as users in a social network) and edges (links) representing relationships between those entities.

    Graph mining is closely related to Social Network Analysis (SNA), which is the process of investigating social structures through the use of networks and graph theory. It characterizes networked structures in terms of:

    • Nodes: Individual actors or entities
    • Edges/Links: Relationships connecting them (which can be positive or negative)

    Graph mining tasks include link prediction, community detection, classification, and clustering over graph data.


    Part 2: Conflict Between Theory of Balance and Theory of Status (4 marks)

    In social networks, relationships can be either positive (trust, friendship) or negative (distrust, opposition). Two major theories attempt to predict the sign of links in such signed networks: Balance Theory and Status Theory. These two theories sometimes come into direct conflict with each other.


    Theory of Balance

    Balance theory is based on the idea that relationships in a social network tend toward consistency. In simple terms:

    • A positive link from A to B means A considers B a friend.
    • A negative link from A to B means A considers B an enemy.

    The classic rule is: "A friend of my friend is my friend."


    Theory of Status

    In the Theory of Status, a signed directed link from A to B is interpreted in terms of relative status:

    • A positive directed link from A to B means A regards B as having higher status than A.
    • A negative directed link from A to B means A regards B as having lower status than A.

    These relative levels of status can then be propagated along multi-step paths of signed links, often leading to different predictions than balance theory.


    The Conflict: A Concrete Example

    Consider the following scenario:

    User A links positively to User B, and B links positively to User C. If C then forms a link to A, what sign should we expect this link to have?

    Prediction by Balance Theory:

    • A is a friend of B (positive link A -> B)
    • B is a friend of C (positive link B -> C)
    • Therefore, C is a friend of A's friend
    • Balance theory predicts: C should link POSITIVELY to A

    Prediction by Status Theory:

    • A links positively to B => A regards B as having higher status than A
    • B links positively to C => B regards C as having higher status than B
    • Therefore, C has higher status than B, who has higher status than A
    • So C should regard A as having low status
    • Status theory predicts: C should link NEGATIVELY to A

    Summary of Conflict

    AspectBalance TheoryStatus Theory
    Positive link A->B meansA and B are friendsB has higher status than A
    Prediction for C->APositive (friend of friend)Negative (A has low status)
    BasisFriendship/enmity consistencyRelative social status propagation

    This conflict shows that the same network configuration (A->B positive, B->C positive) leads to opposite predictions depending on which theory is applied. Machine learning algorithms and matrix factorization approaches have been proposed to handle both theories together for predicting positive and negative links in social networks.

  8. 85 marksSupport vector machineAnswer

    What is support vector? How do you evaluate the accuracy of a classifier? Describe. [5]

    Support Vector and Classifier Accuracy Evaluation


    Part 1: What is a Support Vector?

    A Support Vector Machine (SVM) is a supervised learning algorithm used primarily for classification problems. It works by finding the optimal hyperplane that best separates data points belonging to different classes.

    Support Vectors are the data points that lie closest to the decision boundary (hyperplane). These are the critical points that directly influence the position and orientation of the hyperplane.

    Key points about support vectors:

    • They are the boundary data points from each class
    • The margin is defined as the distance between the hyperplane and the nearest support vectors from each class
    • SVM aims to maximize this margin to achieve the best separation
    • Removing support vectors would change the position of the hyperplane
    • SVM takes input data points and outputs the hyperplane (decision boundary) that separates the classes

    Example: If the hyperplane equation is 2x - 2 = 0, the support vectors are the points that satisfy the margin constraints on either side of this boundary.


    Part 2: Evaluating the Accuracy of a Classifier

    Confusion Matrix

    A Confusion Matrix is an N x N matrix used to evaluate the performance of a classification model, where N is the number of target classes.

    For a binary classification, it is a 2 x 2 matrix:

    Predicted: PositivePredicted: Negative
    Actual: PositiveTP (True Positive)FN (False Negative)
    Actual: NegativeFP (False Positive)TN (True Negative)

    Definitions:

    • TP (True Positive): Both actual and predicted classes are positive (correctly classified positive)
    • TN (True Negative): Both actual and predicted classes are negative (correctly classified negative)
    • FP (False Positive): Predicted positive but actually negative -- called Type I Error
    • FN (False Negative): Predicted negative but actually positive -- called Type II Error

    Four Performance Measures

    1. Accuracy

    $$\text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN}$$ Measures the overall proportion of correctly classified instances.

    2. Precision

    $$\text{Precision} = \frac{TP}{TP + FP}$$ Measures how many of the predicted positives are actually positive.

    3. Recall (Sensitivity)

    $$\text{Recall} = \frac{TP}{TP + FN}$$ Measures how many of the actual positives were correctly identified.

    4. F1-Score

    $$\text{F1-Score} = \frac{2 \times \text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}$$ The harmonic mean of Precision and Recall; useful when there is a class imbalance.


    Summary Table

    MeasureFormulaWhat it tells
    Accuracy(TP+TN)/(TP+TN+FP+FN)Overall correctness
    PrecisionTP/(TP+FP)Quality of positive predictions
    RecallTP/(TP+FN)Coverage of actual positives
    F1-Score2PR/(P+R)Balance between precision and recall

    These four measures together provide a comprehensive evaluation of a classifier's performance beyond simple accuracy.

  9. 95 marksClustering techniquesAnswer

    Differentiate between k-means and k-medoids clustering algorithm. [5]

    Both K-Means and K-Medoids are partitioning-based clustering algorithms that divide a dataset of n objects into k clusters. However, they differ significantly in how they represent cluster centers and handle data. --- Feature K-Means K-M...

  10. 105 marksOLAP operation in multidimensional data moAnswer

    List any two OLAP operations with example. How do you compute rule coverage and rule accuracy? [5]

    OLAP Operations and Rule Coverage/Accuracy


    Part 1: Two OLAP Operations with Examples

    1. Roll-up (Consolidation / Aggregation)

    Roll-up generates a summary from a lower level to a higher level. It can be performed by:

    • Reducing dimensions, or
    • Climbing up a concept hierarchy

    Example: In the location dimension, data can be rolled up from city level to country level. For instance, sales data for cities like Kathmandu, Pokhara, and Biratnagar are aggregated (summed) to give total sales for Nepal as a country.

    Concept hierarchy: Day → Week → Month → Year


    2. Drill-down (Opposite of Roll-up)

    Drill-down navigates from higher-level summarized data to lower-level detailed data. It is the reverse of roll-up, either by stepping down a concept hierarchy or by introducing additional dimensions.

    Example: If we have quarterly sales data for a country, we can drill down to view monthly sales for each city within that country. This gives more granular, detailed information.


    Part 2: Rule Coverage and Rule Accuracy

    The quality of an association/classification rule R is measured by two factors: coverage and accuracy.

    Let the rule be of the form:

    R: A => B (If condition A, then class/consequence B)

    Let:

    • n = total number of tuples in the dataset
    • nA = number of tuples satisfying condition A (antecedent)
    • nA,B = number of tuples satisfying both A and B

    Rule Coverage

    Coverage measures how often the rule applies to the dataset. It is the fraction of tuples that satisfy the antecedent of the rule.

    $$\text{Coverage}(R) = \frac{n_A}{n}$$

    Example: If a dataset has 1000 tuples and 400 tuples satisfy condition A, then:

    $$\text{Coverage}(R) = \frac{400}{1000} = 0.4 = 40%$$


    Rule Accuracy (Confidence)

    Accuracy measures how often the rule is correct. It is the fraction of tuples satisfying the antecedent that also satisfy the consequent.

    $$\text{Accuracy}(R) = \frac{n_{A,B}}{n_A}$$

    Example: Out of the 400 tuples satisfying condition A, suppose 300 also satisfy B, then:

    $$\text{Accuracy}(R) = \frac{300}{400} = 0.75 = 75%$$


    Summary Table

    MeasureFormulaMeaning
    Coverage$n_A / n$How often the rule's condition applies
    Accuracy$n_{A,B} / n_A$How often the rule is correct when it applies

    A good rule should have both high coverage (applies to many tuples) and high accuracy (is correct most of the time).

  11. 115 marksLink miningAnswer

    Define link mining. What are the roles of epsilon and MinPts in DBSCAN. [5]

    --- Link mining is a research area that focuses on extracting useful knowledge and patterns from data that contains links or relationships between objects. Unlike traditional data mining which considers only the attributes of individual ...

  12. 125 marksData cleaningAnswer

    Describe any two methods of handling noisy data. [5]

    Noisy data refers to data that contains errors, outliers, or random variance that deviates from the expected values. Real world data tends to be incomplete, noisy, and inconsistent. Data cleaning attempts to smooth out noise and correct ...