2075

CSC420 · TU past paper

Data Warehousing and Data Mining 2075 question paper

The complete TU 2075 exam paper for Data Warehousing and Data Mining (CSC420), all 9 questions with solved model answers written to the mark scheme.

Tap a question to open its answer.

  1. 110 marksMultimedia data miningAnswer

    List some issues of multimedia mining. Describe how back propagation is used in classification.[10]

    Issues of Multimedia Mining and Back Propagation in Classification


    Part 1: Issues of Multimedia Mining

    Multimedia database is a collection of interrelated multimedia data that includes text, sketches, drawings, images, animations, video, audio, and hypertext. Multimedia data mining is an interdisciplinary field that integrates image processing and understanding, computer vision, data mining, and pattern recognition.

    The key issues in multimedia data mining are as follows:

    • Retrieving multimedia data based on its actual content (e.g., color, texture, shape in images) rather than metadata or keywords is extremely challenging.
    • Defining meaningful similarity measures across different media types (audio, video, image) is non-trivial.

    2. High Dimensionality

    • Multimedia data such as images and video are represented by high-dimensional feature vectors.
    • Mining in high-dimensional spaces suffers from the "curse of dimensionality," making distance-based methods less effective.

    3. Large Data Volume and Scalability

    • Multimedia files (video, audio) are very large in size.
    • Storing, indexing, and mining such large volumes of data efficiently is a major challenge.

    4. Heterogeneity of Data Types

    • Multimedia data combines multiple modalities (text, image, audio, video).
    • Integrating and mining across these heterogeneous data types requires specialized techniques for each modality.

    5. Feature Extraction and Representation

    • Raw multimedia data must be converted into meaningful features before mining.
    • Extracting relevant low-level features (color histograms, MFCC for audio, motion vectors for video) that capture high-level semantics is difficult.

    6. Semantic Gap

    • There is a significant gap between low-level features automatically extracted from multimedia data and the high-level semantic meaning perceived by humans.
    • Bridging this gap remains an open research problem.

    7. Uncertainty and Noise

    • Multimedia data often contains noise (blurry images, background noise in audio).
    • Handling uncertainty and incompleteness in such data during mining is challenging.

    8. Mining Methodology

    • It includes issues in mining various and new kinds of knowledge, mining knowledge in multidimensional space, and handling uncertainty, noise, or incompleteness of data (Under Issues in Data Mining).

    Part 2: Back Propagation Used in Classification

    Background

    Back Propagation (Backprop) is a supervised learning algorithm used to train Artificial Neural Networks (ANNs). It is widely used for classification tasks where the goal is to assign input data to one of several predefined classes.


    Architecture of a Neural Network for Classification

    A typical neural network used for classification consists of:

    LayerDescription
    Input LayerReceives input features x₁, x₂, ..., xₙ
    Hidden Layer(s)Performs non-linear transformations
    Output LayerProduces class probabilities or class labels

    How Back Propagation Works in Classification

    The back propagation algorithm works in two phases:


    Phase 1: Forward Pass

    1. Input features are fed into the network.
    2. Each neuron computes a weighted sum of its inputs:

    $$net_j = \sum_{i} w_{ij} \cdot x_i + b_j$$

    1. An activation function (e.g., sigmoid, ReLU, softmax) is applied:

    $$o_j = f(net_j)$$

    1. The output layer produces predicted class probabilities.
    2. For classification, softmax is commonly used at the output layer:

    $$P(class_k) = \frac{e^{net_k}}{\sum_{j} e^{net_j}}$$


    Phase 2: Backward Pass (Error Propagation)

    1. Compute the Error (Loss):
      • The difference between the predicted output and the actual class label is computed using a loss function.
      • For classification, Cross-Entropy Loss is commonly used:

    $$E = -\sum_{k} y_k \cdot \log(\hat{y}_k)$$

    where yₖ is the true label (1 or 0) and ŷₖ is the predicted probability.

    1. Compute Error at Output Layer:
      • The error signal (delta) at the output neuron k is:

    $$\delta_k = (y_k - o_k) \cdot f'(net_k)$$

    where f'(net_k) is the derivative of the activation function.

    1. Propagate Error to Hidden Layers:
      • The error is propagated backward through the network:

    $$\delta_j = f'(net_j) \cdot \sum_{k} \delta_k \cdot w_{jk}$$

    1. Update Weights:
      • Weights are updated using gradient descent to minimize the error:

    $$w_{ij} \leftarrow w_{ij} + \Delta w_{ij}$$

    $$\Delta w_{ij} = \eta \cdot \delta_j \cdot o_i$$

    where:

    • η = learning rate
    • δⱼ = error signal at neuron j
    • oᵢ = output of the previous neuron i
    1. Update Biases similarly:

    $$b_j \leftarrow b_j + \eta \cdot \delta_j$$


    Iterative Training for Classification

    • The forward pass and backward pass are repeated for all training samples over multiple epochs.
    • The network gradually adjusts its weights to correctly classify training examples.
    • Training stops when:
      • The error falls below a threshold, OR
      • A maximum number of epochs is reached.

    Summary of Back Propagation Steps for Classification

    1. Initialize weights randomly
    2. For each training sample:
       a. Forward pass: compute output
       b. Compute classification error (loss)
       c. Backward pass: compute deltas at output and hidden layers
       d. Update all weights and biases
    3. Repeat until convergence
    4. Use trained network to classify new/unseen data
    

    Example

    Consider a network trained to classify a scanned handwritten digit into one of the ten classes 0 to 9. The input layer holds one neuron for each pixel of the 28 by 28 image, a single hidden layer holds a smaller number of sigmoid neurons, and the output layer holds ten softmax neurons, one per class.

    During the forward pass each hidden neuron computes $net_j = \sum_i w_{ij} x_i + b_j$ and applies the sigmoid, and the output layer turns the ten net values into class probabilities with the softmax formula given above. Suppose the image is really a 7 but the network gives the highest probability to class 1: the cross-entropy loss is large because the probability assigned to the true class is small.

    The backward pass then repairs that. The output deltas $\delta_k = (y_k - o_k) f'(net_k)$ are positive for the neuron of class 7, whose target is 1, and negative for the neuron of class 1, whose target is 0. Each hidden neuron receives its share of the blame through $\delta_j = f'(net_j) \sum_k \delta_k w_{jk}$, and every weight is then nudged by $\Delta w_{ij} = \eta \delta_j o_i$, so the pixels that actually distinguish a 7 gain influence on the class 7 output and lose influence on the class 1 output.

    Repeating this over the whole training set for many epochs drives the loss down until the network assigns the highest probability to the correct digit, and the trained weights can then classify digits it has never seen.


    Conclusion

    Multimedia mining is difficult because the data is large, high dimensional, heterogeneous, noisy, and separated from human meaning by the semantic gap. Back propagation is one of the standard answers to that difficulty: it learns the mapping from raw features to class labels automatically by alternating a forward pass that produces a prediction with a backward pass that distributes the error over every weight, and it repeats the pair until the network classifies unseen data correctly.

  2. 210 marksComponents of data warehouseAnswer

    Describe how bitmap and join indexing are used to represent OLAP data. Explain the different components of data warehouse.[10]

    --- Definition: A bitmap index is a special type of index that uses bit vectors (arrays of 0s and 1s) to represent the presence or absence of a value for each row in a table. It is highly efficient for OLAP queries involving low-cardinal...

  3. 310 marksNumericalFinding frequent itemsetAnswer

    Association Rules and Apriori Algorithm Analysis

    • Support threshold (min support count) = 2 - Confidence threshold = 60% - Total transactions = 6 TID Items ------------ T1 HotDogs, Buns, Ketchup T2 HotDogs, Buns T3 HotDogs, Coke, Chips T4 Chips, Coke T5 Chips, Ketchup T6 HotDogs, Coke...
  4. 45 marksTypes of data in cluster analysisAnswer

    What is the purpose of cluster analysis in data mining? Explain. [5]

    Cluster analysis is a data mining technique in which data objects are grouped into clusters based on their similarity. Objects within the same cluster have high similarity to one another, while objects belonging to different clusters are...

  5. 55 marksKDDAnswer

    How does KDD differ with data mining? Describe the stages of data mining. [5]

    KDD (Knowledge Discovery from Data) is the overall process of discovering useful knowledge from raw data. Data Mining is just one step within this larger KDD process. Aspect KDD Data Mining --------- Scope Entire process from raw data to...

  6. 65 marksOLAP operation in multidimensional data moAnswer

    Explain OLAP operations with examples. [5]

    OLAP (Online Analytical Processing) allows users to analyze multidimensional data from multiple perspectives. Data is organized in a data cube with dimensions (e.g., location, time, product) and facts (e.g., sales amount, units sold). Th...

  7. 75 marksData mining primitivesAnswer

    Explain the primitives of data mining query language. [5]

    A data mining query is used to specify a data mining task and is input to the data mining system. It is defined in terms of data mining primitives, which allow the user to interactively communicate with the data mining system in order to...

  8. 85 marksA multidimensional data modelAnswer

    How different schema are used to model data warehouse? Explain. [5]

    A schema is a logical description of the entire data warehouse, including the name and description of records and aggregates. The goal of conceptual data warehouse modeling is to develop a schema for logical representation of data stored...

  9. 95 marksEfficient method for data cube computationAnswer

    Describe the significances of pre-computation of data cube. [5]

    Data cube computation is an essential task in data warehouse implementation. Pre-computation refers to computing all or part of a data cube in advance (offline), so that query results can be retrieved quickly during online analytical pro...