CSC420 · TU past paper
Data Warehousing and Data Mining 2081 question paper
The complete TU 2081 exam paper for Data Warehousing and Data Mining (CSC420), all 12 questions with solved model answers written to the mark scheme.
Tap a question to open its answer.
- 110 marksA multidimensional data modelHideAnswer
When do we prefer trim mean for statistical description of data? Justify with an example. Describe about multi-dimensional data model and conceptual modeling of data warehouse.[10]
Trim Mean for Statistical Description & Multidimensional Data Model / Conceptual Modeling of Data Warehouse
PART 1: Trim Mean - When to Prefer It and Justification
Definition of Trim Mean
The trimmed mean (trim mean) is a measure of central tendency calculated by removing a specified percentage of the smallest and largest values from a dataset before computing the arithmetic mean. If we trim p% from each end, it is called a p% trimmed mean.
Formula:
$$\bar{x}{trim} = \frac{1}{n - 2k} \sum{i=k+1}^{n-k} x_{(i)}$$
where:
- $x_{(i)}$ are the sorted (ordered) values
- $k = \lfloor p \times n / 100 \rfloor$ is the number of values trimmed from each end
- $n$ is the total number of observations
When Do We Prefer Trim Mean?
We prefer the trimmed mean over the ordinary arithmetic mean in the following situations:
-
Presence of Outliers: When the dataset contains extreme values (outliers) that distort the arithmetic mean, the trim mean provides a more robust and representative central value.
-
Skewed Distributions: When data is heavily skewed (positively or negatively), the ordinary mean is pulled toward the tail. The trim mean reduces this effect.
-
Noisy or Erroneous Data: In real-world data collection, some values may be recorded incorrectly or represent measurement errors. Trimming removes these suspicious extreme values.
-
When Median is Too Extreme: The median ignores all values except the middle one, while the arithmetic mean is too sensitive to extremes. The trim mean offers a balance between the two - it is more robust than the mean but uses more data than the median.
-
Data Mining and Statistical Description: In data preprocessing for data mining, when describing large datasets with potential noise, the trim mean gives a cleaner summary statistic.
Justification with Example
Example:
Consider the monthly salaries (in thousands) of 10 employees:
$${12, 14, 15, 16, 17, 18, 19, 20, 21, 150}$$
Step 1: Compute the Arithmetic Mean
$$\bar{x} = \frac{12 + 14 + 15 + 16 + 17 + 18 + 19 + 20 + 21 + 150}{10} = \frac{302}{10} = 30.2$$
The mean is 30.2, but 9 out of 10 employees earn between 12 and 21. The value 150 is clearly an outlier (perhaps the CEO's salary), and it has inflated the mean significantly.
Step 2: Compute the 10% Trimmed Mean
- Trim 10% from each end: $k = 0.10 \times 10 = 1$ value removed from each side
- Remove the lowest value: 12
- Remove the highest value: 150
- Remaining values: ${14, 15, 16, 17, 18, 19, 20, 21}$
$$\bar{x}_{10%} = \frac{14 + 15 + 16 + 17 + 18 + 19 + 20 + 21}{8} = \frac{140}{8} = 17.5$$
Conclusion: The trimmed mean of 17.5 is far more representative of the typical employee salary than the arithmetic mean of 30.2. This justifies using the trim mean when outliers are present.
PART 2: Multidimensional Data Model
Definition
A multidimensional data model is a data model that organizes data in multiple dimensions to facilitate complex analytical queries and reporting. It is the foundation of OLAP (Online Analytical Processing) systems and data warehouses.
- Data is viewed as a data cube with multiple dimensions.
- Each dimension represents a perspective or entity of interest (e.g., Time, Location, Product).
- The model is typically organized around a central theme (e.g., Sales), represented by a fact table.
- Facts are numerical measures such as sales amount, units sold, profit, etc.
- Each dimension may have a dimension table associated with it.
Key Components
Component Description Fact Table Central table containing numerical measures (facts) Dimension Table Tables describing the dimensions (e.g., Time, Product, Location) Measures/Facts Quantitative data: sales, revenue, units sold Dimensions Qualitative context: who, what, when, where
OLAP Operations in Multidimensional Data Model
1. Roll-Up (Consolidation / Aggregation)
- Generates summary from lower level to higher level.
- Can be performed by:
- Reducing dimensions
- Climbing up a concept hierarchy
- Concept Hierarchy: A system of grouping things based on their order level. For example, from daily summary, weekly summary is generated; from weekly, monthly summary is generated.
- Example: Roll-up on location from cities to countries - individual city sales are aggregated into country-level totals.
2. Drill-Down
- The reverse of roll-up; navigates from summary data to more detailed data.
- Example: From yearly sales, drill down to quarterly, then monthly sales.
3. Slice
- Selects one particular dimension from a data cube, resulting in a sub-cube.
- Example: Slice on Time = "Q1 2024" to get all sales data for Q1 only.
4. Dice
- Selects two or more dimensions from a data cube to produce a sub-cube.
- Example: Dice on Location = {Nepal, India} AND Time = {Q1, Q2}.
5. Pivot (Rotate)
- Rotates the data axes to provide an alternative presentation of data.
- Example: Swapping rows and columns in a cross-tabulation.
PART 3: Conceptual Modeling of Data Warehouse
Definition
A **conceptual
- 210 marksNumericalFinding frequent itemsetHideAnswer
How do you generate strong association rules? From the following dataset find the frequent item set using FP growth algorithm using 3 as minimum support.
Transaction ID Items T1 {K, E, M, O, Y} T2 {K, E, O, Y} T3 {K, E, M} T4 {K, M, Y} T5 {K, E, O} [10]
Strong Association Rules and FP-Growth
Given Data
Transactions:
- T1: {K, E, M, O, Y}
- T2: {K, E, O, Y}
- T3: {K, E, M}
- T4: {K, M, Y}
- T5: {K, E, O}
Minimum support = 3
Part 1: Generating Strong Association Rules
A strong association rule $X \Rightarrow Y$ satisfies both thresholds:
- $\text{support}(X \cup Y) \geq \text{min_support}$
- $\text{confidence}(X \Rightarrow Y) = \dfrac{\text{support}(X \cup Y)}{\text{support}(X)} \geq \text{min_confidence}$
Procedure:
- Find all frequent itemsets satisfying min support (via Apriori or FP-Growth).
- For each frequent itemset $l$, generate every non-empty proper subset $s$.
- For each such $s$, form the rule $s \Rightarrow (l - s)$ and compute confidence $= \dfrac{\text{support}(l)}{\text{support}(s)}$.
- Output the rule only if confidence $\geq$ min_confidence. These are the strong rules.
Part 2: FP-Growth (min_support = 3)
Step 1: Item frequency (1st scan)
Item Count K 5 E 4 M 3 O 3 Y 3 All satisfy min support. Order (descending, ties alphabetical): K:5, E:4, M:3, O:3, Y:3
Step 2: Reordered transactions
TID Ordered T1 K, E, M, O, Y T2 K, E, O, Y T3 K, E, M T4 K, M, Y T5 K, E, O Step 3: FP-Tree
null └── K:5 ├── E:4 │ ├── M:1 │ │ └── O:1 │ │ └── Y:1 │ └── O:2 │ └── Y:1 └── M:2 └── Y:1Note on M placement:
- T1 (KEMOY) and T3 (KEM) go under E-branch: E→M appears twice (M:2 under E).
- T4 (KMY) goes under K directly: K→M (M:1 under K, no E).
So M under E has count 2, and M under K (no E) has count 1. Total M = 3. ✓
Corrected FP-Tree:
null └── K:5 ├── E:4 │ ├── M:2 (from T1, T3) │ │ └── O:1 (T1) │ │ └── Y:1 │ └── O:2 (T2, T5) │ └── Y:1 (T2) └── M:1 (T4) └── Y:1Take care with the M counts: the correct placement is M:2 under E and M:1 under K, not the other way round.
Header table:
Item Support Node links K 5 K:5 E 4 E:4 M 3 M:2, M:1 O 3 O:1, O:2 Y 3 Y:1, Y:1, Y:1 Step 4: Mining (bottom-up)
Item Y (support 3): Conditional pattern base:
- {K,E,M,O}:1 (T1)
- {K,E,O}:1 (T2)
- {K,M}:1 (T4)
Counts: K=3, E=2, M=2, O=2. Only K:3 ≥ 3. Frequent: {K,Y}:3
Item O (support 3): Conditional pattern base:
- {K,E,M}:1 (T1)
- {K,E}:2 (T2, T5)
Counts: K=3, E=3, M=1. Survive: K:3, E:3. Conditional FP-tree: K→E (3). Frequent: {K,O}:3, {E,O}:3, {K,E,O}:3
Item M (support 3): Conditional pattern base:
- {K,E}:2 (T1, T3)
- {K}:1 (T4)
Counts: K=3, E=2. Only K:3 survives. Frequent: {K,M}:3
Item E (support 4): Conditional pattern base:
- {K}:4
Frequent: {K,E}:4
Item K: root, no conditional patterns.
Final Frequent Itemsets (support ≥ 3)
1-itemsets:
- {K}:5, {E}:4, {M}:3, {O}:3, {Y}:3
2-itemsets:
- {K,E}:4
- {K,M}:3
- {K,O}:3
- {E,O}:3
- {K,Y}:3
3-itemsets:
- {K,E,O}:3
The maximal/most useful frequent itemset is {K, E, O} with support 3.
- 310 marksNumericalID3 as attribute selection algorithmHideAnswer
Decision Tree Classification with ID3 Algorithm
Overfitting and Underfitting + ID3 Decision Tree
Step 1 - EXTRACT: Given Data
TID Age Car Type Class 1 ≤30 Family High 2 ≤30 Sports High 3 >30 Sports High 4 >30 Family Low 5 >30 Truck Low 6 ≤30 Family High - Total tuples = 6
- Class High: TID 1, 2, 3, 6 → 4 tuples
- Class Low: TID 4, 5 → 2 tuples
- Attributes: Age (≤30, >30), Car Type (Family, Sports, Truck)
Part 1: Definitions
Overfitting
Overfitting occurs when a model learns the training data too well, capturing noise and random fluctuations rather than the underlying pattern. It shows high accuracy on training data but poor accuracy on unseen/test data. In decision trees, this happens when the tree grows too deep.
Underfitting
Underfitting occurs when a model is too simple to capture the underlying structure of the data. It shows poor accuracy on both training and test data. In decision trees, this happens when the tree is too shallow.
Part 2: ID3 Algorithm
Step 1: Entropy of the whole dataset
$$Entropy(S) = -\frac{4}{6}\log_2\frac{4}{6} - \frac{2}{6}\log_2\frac{2}{6}$$
$$= -\tfrac{2}{3}(-0.585) - \tfrac{1}{3}(-1.585) = 0.390 + 0.528 = 0.918$$
Step 2: Gain for Age
Age = ≤30: TID 1, 2, 6 → High:3, Low:0 → $Entropy = 0$
Age = >30: TID 3, 4, 5 → High:1, Low:2
$$Entropy(>30) = -\tfrac{1}{3}\log_2\tfrac{1}{3} - \tfrac{2}{3}\log_2\tfrac{2}{3} = 0.528 + 0.390 = 0.918$$
$$Info_{Age}(S) = \tfrac{3}{6}(0) + \tfrac{3}{6}(0.918) = 0.459$$
$$Gain(Age) = 0.918 - 0.459 = 0.459$$
Step 3: Gain for Car Type
Family: TID 1, 4, 6 → High:2, Low:1 → $Entropy = 0.918$ Sports: TID 2, 3 → High:2, Low:0 → $Entropy = 0$ Truck: TID 5 → High:0, Low:1 → $Entropy = 0$
$$Info_{CarType}(S) = \tfrac{3}{6}(0.918) + \tfrac{2}{6}(0) + \tfrac{1}{6}(0) = 0.459$$
$$Gain(CarType) = 0.918 - 0.459 = 0.459$$
Step 4: Select Root
Attribute Gain Age 0.459 Car Type 0.459 Both are tied at 0.459. Selecting Age as root (either is valid).
Step 5: Split on Age
Age / \ ≤30 >30 [1,2,6] [3,4,5] H:3,L:0 H:1,L:2 PURE→High impure- Age ≤30: all High → leaf Class = High
- Age >30: TID 3 (High), 4 (Low), 5 (Low) → not pure, split further.
Step 6: Split the (Age > 30) subset
Remaining attribute: Car Type. Subset = {3, 4, 5}.
- Sports: TID 3 → High → pure → High
- Family: TID 4 → Low → pure → Low
- Truck: TID 5 → Low → pure → Low
Entropy of each subset = 0, so Car Type perfectly classifies this subset.
Final Decision Tree
Age / \ ≤30 >30 | | [High] Car Type / | \ Sports Family Truck | | | [High] [Low] [Low]Classification Rules
- If Age ≤ 30 → High
- If Age > 30 AND Car = Sports → High
- If Age > 30 AND Car = Family → Low
- If Age > 30 AND Car = Truck → Low
This tree classifies all 6 training tuples correctly.
The completed tree splits the (Age > 30) subset on Car Type, with entropy = 0.918, both gains = 0.459 and Age chosen as the root, the left branch being pure High.
- 45 marksNumericalClustering techniquesHideAnswer
Using k-means++ algorithm and Euclidean distance, find the initial 3 cluster centroids from A1 = (3, 11), A2 = (3, 6), A3 = (9, 5), A4 = (6, 9), A6 = (7, 5), A7 = (2, 3), A8 = (5, 10). Choose (3, 11) as one of the initial centroids. [5]
K-Means++ Initial Centroid Selection
Step 1 - Given Data
Data points:
Point Coordinates A1 (3, 11) A2 (3, 6) A3 (9, 5) A4 (6, 9) A6 (7, 5) A7 (2, 3) A8 (5, 10) Number of clusters: $k = 3$ First centroid (given): $C_1 = (3, 11)$
Step 2 - Solve
K-Means++ Rule
- First centroid is fixed as $C_1 = (3,11)$.
- For each point compute $D^2$ = squared Euclidean distance to the nearest chosen centroid.
- Choose the point with the largest $D^2$ (highest selection probability) as the next centroid.
Selecting $C_2$: distances to $C_1 = (3,11)$
$$D^2 = (x-3)^2 + (y-11)^2$$
Point Calculation $D^2$ A2 (3,6) $0 + 25$ 25 A3 (9,5) $36 + 36$ 72 A4 (6,9) $9 + 4$ 13 A6 (7,5) $16 + 36$ 52 A7 (2,3) $1 + 64$ 65 A8 (5,10) $4 + 1$ 5 Total $= 25+72+13+52+65+5 = 232$
Probability $= D^2 / 232$:
Point P A2 0.108 A3 0.310 A6 0.224 A7 0.280 A4 0.056 A8 0.022 Largest $D^2$ is A3 (72).
$$\boxed{C_2 = (9, 5)}$$
Selecting $C_3$: nearest distance to ${C_1, C_2}$
$D^2(C_2)$ with $C_2 = (9,5)$:
Point $D^2(C_1)$ $D^2(C_2)$ $\min$ A2 (3,6) 25 $36+1=37$ 25 A4 (6,9) 13 $9+16=25$ 13 A6 (7,5) 52 $4+0=4$ 4 A7 (2,3) 65 $49+4=53$ 53 A8 (5,10) 5 $16+25=41$ 5 Total $= 25+13+4+53+5 = 100$
Probability $= \min D^2 / 100$:
Point P A2 0.25 A4 0.13 A6 0.04 A7 0.53 A8 0.05 Largest $\min D^2$ is A7 (53).
$$\boxed{C_3 = (2, 3)}$$
Final Result
The 3 initial centroids chosen by K-Means++ are:
Centroid Coordinates $C_1$ (3, 11) $C_2$ (9, 5) $C_3$ (2, 3) - 55 marksGeneral strategies for cube computationHideAnswer
Explain the general strategies for cube computation. [5]
Data cube computation is an essential task in data warehouse implementation. The precomputation of all or part of a data cube can greatly reduce response time and enhance the performance of OLAP. However, it is challenging because it may...
- 65 marksAttribute oriented induction for data charHideAnswer
Distinguish between data characterization and data discrimination. What are the challenges of multimedia mining? [5]
--- Both are descriptive data mining functions used to summarize and compare data. They are distinguished as follows: Aspect Data Characterization Data Discrimination --------- Definition Summarization of the general characteristics or f...
- 75 marksSigned networkHideAnswer
Define graph mining. Discuss the conflict between theory of balance and theory of status. [5]
Graph Mining and Conflict Between Theory of Balance and Theory of Status
Part 1: Definition of Graph Mining (1 mark)
Graph Mining is a subfield of data mining that focuses on discovering useful patterns, structures, and knowledge from graph-structured data. It applies data mining techniques to graphs, where data is represented as nodes (vertices) representing entities (such as users in a social network) and edges (links) representing relationships between those entities.
Graph mining is closely related to Social Network Analysis (SNA), which is the process of investigating social structures through the use of networks and graph theory. It characterizes networked structures in terms of:
- Nodes: Individual actors or entities
- Edges/Links: Relationships connecting them (which can be positive or negative)
Graph mining tasks include link prediction, community detection, classification, and clustering over graph data.
Part 2: Conflict Between Theory of Balance and Theory of Status (4 marks)
In social networks, relationships can be either positive (trust, friendship) or negative (distrust, opposition). Two major theories attempt to predict the sign of links in such signed networks: Balance Theory and Status Theory. These two theories sometimes come into direct conflict with each other.
Theory of Balance
Balance theory is based on the idea that relationships in a social network tend toward consistency. In simple terms:
- A positive link from A to B means A considers B a friend.
- A negative link from A to B means A considers B an enemy.
The classic rule is: "A friend of my friend is my friend."
Theory of Status
In the Theory of Status, a signed directed link from A to B is interpreted in terms of relative status:
- A positive directed link from A to B means A regards B as having higher status than A.
- A negative directed link from A to B means A regards B as having lower status than A.
These relative levels of status can then be propagated along multi-step paths of signed links, often leading to different predictions than balance theory.
The Conflict: A Concrete Example
Consider the following scenario:
User A links positively to User B, and B links positively to User C. If C then forms a link to A, what sign should we expect this link to have?
Prediction by Balance Theory:
- A is a friend of B (positive link A -> B)
- B is a friend of C (positive link B -> C)
- Therefore, C is a friend of A's friend
- Balance theory predicts: C should link POSITIVELY to A
Prediction by Status Theory:
- A links positively to B => A regards B as having higher status than A
- B links positively to C => B regards C as having higher status than B
- Therefore, C has higher status than B, who has higher status than A
- So C should regard A as having low status
- Status theory predicts: C should link NEGATIVELY to A
Summary of Conflict
Aspect Balance Theory Status Theory Positive link A->B means A and B are friends B has higher status than A Prediction for C->A Positive (friend of friend) Negative (A has low status) Basis Friendship/enmity consistency Relative social status propagation This conflict shows that the same network configuration (A->B positive, B->C positive) leads to opposite predictions depending on which theory is applied. Machine learning algorithms and matrix factorization approaches have been proposed to handle both theories together for predicting positive and negative links in social networks.
- 85 marksSupport vector machineHideAnswer
What is support vector? How do you evaluate the accuracy of a classifier? Describe. [5]
Support Vector and Classifier Accuracy Evaluation
Part 1: What is a Support Vector?
A Support Vector Machine (SVM) is a supervised learning algorithm used primarily for classification problems. It works by finding the optimal hyperplane that best separates data points belonging to different classes.
Support Vectors are the data points that lie closest to the decision boundary (hyperplane). These are the critical points that directly influence the position and orientation of the hyperplane.
Key points about support vectors:
- They are the boundary data points from each class
- The margin is defined as the distance between the hyperplane and the nearest support vectors from each class
- SVM aims to maximize this margin to achieve the best separation
- Removing support vectors would change the position of the hyperplane
- SVM takes input data points and outputs the hyperplane (decision boundary) that separates the classes
Example: If the hyperplane equation is 2x - 2 = 0, the support vectors are the points that satisfy the margin constraints on either side of this boundary.
Part 2: Evaluating the Accuracy of a Classifier
Confusion Matrix
A Confusion Matrix is an N x N matrix used to evaluate the performance of a classification model, where N is the number of target classes.
For a binary classification, it is a 2 x 2 matrix:
Predicted: Positive Predicted: Negative Actual: Positive TP (True Positive) FN (False Negative) Actual: Negative FP (False Positive) TN (True Negative) Definitions:
- TP (True Positive): Both actual and predicted classes are positive (correctly classified positive)
- TN (True Negative): Both actual and predicted classes are negative (correctly classified negative)
- FP (False Positive): Predicted positive but actually negative -- called Type I Error
- FN (False Negative): Predicted negative but actually positive -- called Type II Error
Four Performance Measures
1. Accuracy
$$\text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN}$$ Measures the overall proportion of correctly classified instances.
2. Precision
$$\text{Precision} = \frac{TP}{TP + FP}$$ Measures how many of the predicted positives are actually positive.
3. Recall (Sensitivity)
$$\text{Recall} = \frac{TP}{TP + FN}$$ Measures how many of the actual positives were correctly identified.
4. F1-Score
$$\text{F1-Score} = \frac{2 \times \text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}$$ The harmonic mean of Precision and Recall; useful when there is a class imbalance.
Summary Table
Measure Formula What it tells Accuracy (TP+TN)/(TP+TN+FP+FN) Overall correctness Precision TP/(TP+FP) Quality of positive predictions Recall TP/(TP+FN) Coverage of actual positives F1-Score 2PR/(P+R) Balance between precision and recall These four measures together provide a comprehensive evaluation of a classifier's performance beyond simple accuracy.
- 95 marksClustering techniquesHideAnswer
Differentiate between k-means and k-medoids clustering algorithm. [5]
Both K-Means and K-Medoids are partitioning-based clustering algorithms that divide a dataset of n objects into k clusters. However, they differ significantly in how they represent cluster centers and handle data. --- Feature K-Means K-M...
- 105 marksOLAP operation in multidimensional data moHideAnswer
List any two OLAP operations with example. How do you compute rule coverage and rule accuracy? [5]
OLAP Operations and Rule Coverage/Accuracy
Part 1: Two OLAP Operations with Examples
1. Roll-up (Consolidation / Aggregation)
Roll-up generates a summary from a lower level to a higher level. It can be performed by:
- Reducing dimensions, or
- Climbing up a concept hierarchy
Example: In the location dimension, data can be rolled up from city level to country level. For instance, sales data for cities like Kathmandu, Pokhara, and Biratnagar are aggregated (summed) to give total sales for Nepal as a country.
Concept hierarchy: Day → Week → Month → Year
2. Drill-down (Opposite of Roll-up)
Drill-down navigates from higher-level summarized data to lower-level detailed data. It is the reverse of roll-up, either by stepping down a concept hierarchy or by introducing additional dimensions.
Example: If we have quarterly sales data for a country, we can drill down to view monthly sales for each city within that country. This gives more granular, detailed information.
Part 2: Rule Coverage and Rule Accuracy
The quality of an association/classification rule R is measured by two factors: coverage and accuracy.
Let the rule be of the form:
R: A => B (If condition A, then class/consequence B)
Let:
- n = total number of tuples in the dataset
- nA = number of tuples satisfying condition A (antecedent)
- nA,B = number of tuples satisfying both A and B
Rule Coverage
Coverage measures how often the rule applies to the dataset. It is the fraction of tuples that satisfy the antecedent of the rule.
$$\text{Coverage}(R) = \frac{n_A}{n}$$
Example: If a dataset has 1000 tuples and 400 tuples satisfy condition A, then:
$$\text{Coverage}(R) = \frac{400}{1000} = 0.4 = 40%$$
Rule Accuracy (Confidence)
Accuracy measures how often the rule is correct. It is the fraction of tuples satisfying the antecedent that also satisfy the consequent.
$$\text{Accuracy}(R) = \frac{n_{A,B}}{n_A}$$
Example: Out of the 400 tuples satisfying condition A, suppose 300 also satisfy B, then:
$$\text{Accuracy}(R) = \frac{300}{400} = 0.75 = 75%$$
Summary Table
Measure Formula Meaning Coverage $n_A / n$ How often the rule's condition applies Accuracy $n_{A,B} / n_A$ How often the rule is correct when it applies A good rule should have both high coverage (applies to many tuples) and high accuracy (is correct most of the time).
- 115 marksLink miningHideAnswer
Define link mining. What are the roles of epsilon and MinPts in DBSCAN. [5]
--- Link mining is a research area that focuses on extracting useful knowledge and patterns from data that contains links or relationships between objects. Unlike traditional data mining which considers only the attributes of individual ...
- 125 marksData cleaningHideAnswer
Describe any two methods of handling noisy data. [5]
Noisy data refers to data that contains errors, outliers, or random variance that deviates from the expected values. Real world data tends to be incomplete, noisy, and inconsistent. Data cleaning attempts to smooth out noise and correct ...