🎓 Computer Science & Engineering Portal

Master Engineering Disciplines with Structured Notes

Comprehensive academic lecture notes, exam-oriented unit summaries, laboratory manuals, and previous year question papers designed strictly for university students.

📑

University Syllabi

AKTU & AICTE aligned semester credit guidelines.

📝

Exam Question Papers

Previous 5 years solved university semester papers.

💡

Lab Manuals & Viva

Practical codes with outputs and interview questions.

Core Subjects & Units Hub

Click on any specific unit to immediately view its lecture notes below

🎨 Computer Graphics (CG)

Scan Conversion, Bresenham Line & Circle, 2D/3D Transformations, Viewing & Clipping.

Unit 1: Raster Scan, DDA & Bresenham →
Unit 2: 2D & 3D Transformations →
Unit 3: Sutherland-Hodgman & Clipping →
Unit 4: Hidden Surface Elimination →

🧠 Machine Learning (ML / MLT)

Supervised/Unsupervised Learning, Regression, Decision Trees, SVM, Neural Nets & Clustering.

Unit 1: Linear & Logistic Regression →
Unit 2: Decision Trees & Support Vector (SVM) →
Unit 3: K-Means & Dimensionality Reduction →
Unit 4: Neural Networks & Gradient Descent →

🤖 Artificial Intelligence (AI)

Search Algorithms, First Order Logic, Probabilistic Reasoning, Expert Systems & Robotics.

Unit 1: Propositional Logic & Connectives →
Unit 2: Probabilistic Reasoning & Uncertainty →
Unit 3: State Space Search & Heuristics →
Unit 4: First Order Predicate Logic (FOL) →

🗄️ Database Management (DBMS)

ER-Modeling, Relational Algebra, SQL Queries, Normalization (1NF-BCNF) and ACID Transactions.

Unit 1: ER Model, Entities & Attributes →
Unit 2: Functional Dependencies & 1NF to BCNF →
Unit 3: ACID Properties & Concurrency Control →
Unit 4: Relational Algebra & Complex SQL Joins →

🌲 Data Structures & Algorithms

Arrays, Linked Lists, Stacks, Queues, Binary Trees, Graphs, Sorting & Asymptotic Analysis.

Unit 1: Arrays, Matrices & Recursion →
Unit 2: Stacks, Queues & Infix-to-Postfix →
Unit 3: Binary Trees, BST & AVL Rotations →
Unit 4: Graphs (BFS, DFS, Dijkstra, MST) →

⚡ Operating Systems

Process Scheduling, Deadlocks, Synchronization, Virtual Memory, Paging and Disk Management.

Unit 1: Process States, PCB & Multi-Threading →
Unit 2: CPU Scheduling (FCFS, SJF, RR) →
Unit 3: Deadlocks, Semaphores & Banker's Algo →
Unit 4: Virtual Memory, Paging & Disk Scheduling →

🌐 Computer Networks

OSI & TCP/IP Models, Error Detection, IPv4 Subnetting, Routing Protocols and TCP Handshake.

Unit 1: OSI vs TCP/IP Protocol Architectures →
Unit 2: Data Link Layer, Framing & Sliding Window →
Unit 3: IPv4 Addressing, Subnetting & Routing →
Unit 4: Transport Layer (TCP 3-Way Handshake) →

⚙️ Design of Algorithms (DAA)

Asymptotic Notations, Divide & Conquer, Dynamic Programming, Greedy Approach & Backtracking.

Unit 1: Time Complexity, Master's Theorem →
Unit 2: 0/1 Knapsack & Dynamic Programming →
Unit 3: Greedy Methods & Graph Algorithms →
Viewing All Lectures

ISSUES IN DECISION TREE LEARNING

1.      Avoiding Overfitting the Data

Reduced error pruning 

Rule post-pruning

2.      Incorporating Continuous-Valued Attributes

3.      Alternative Measures for Selecting Attributes

4.      Handling Training Examples with Missing Attribute Values

5.      Handling Attributes with Differing Costs

1.  Avoiding Overfitting the Data

  • The ID3 algorithm grows each branch of the tree just deeply enough to perfectly classify the training examples but it can lead to difficulties when there is noise in the data, or when the number of training examples is too small to produce a representative sample of the true target function. This algorithm can produce trees that overfit the training examples.
  •  Definition - Overfit: Given a hypothesis space H, a hypothesis h ∈ H is said to overfit the training data if there exists some alternative hypothesis h' ∈ H, such that h has smaller error than h' over the training examples, but h' has a smaller error than h over the entire distribution of instances.

The below figure illustrates the impact of overfitting in a typical application of decision tree learning.

  • The horizontal axis of this plot indicates the total number of nodes in the decision tree, as the tree is being constructed. The vertical axis indicates the accuracy of predictions made by the tree.
  • The solid line shows the accuracy of the decision tree over the training examples. The broken line shows accuracy measured over an independent set of test example
  • The accuracy of the tree over the training examples increases monotonically as the tree is grown. The accuracy measured over the independent test examples first increases, then decreases.

 How can it be possible for tree h to fit the training examples better than h', but for it to perform more poorly over subsequent examples?

  1. Overfitting can occur when the training examples contain random errors or noise
  2. When small numbers of examples are associated with leaf nodes.

 Noisy Training Example

  • Example 15: <Sunny, Hot, Normal, Strong, ->
  • Example is noisy because the correct label is +
  • Previously constructed tree misclassifies it


Approaches to avoiding overfitting in decision tree learning
  • Pre-pruning (avoidance): Stop growing the tree earlier, before it reaches the point where it perfectly classifies the training data
  • Post-pruning (recovery): Allow the tree to overfit the data, and then post-prune the tree

Criterion used to determine the correct final tree size

  • Use a separate set of examples, distinct from the training examples, to evaluate the utility of post-pruning nodes from the tree
  • Use all the available data for training, but apply a statistical test to estimate whether expanding (or pruning) a particular node is likely to produce an improvement beyond the training set
  • Use measure of the complexity for encoding the training examples and the decision tree, halting growth of the tree when this encoding size is minimized. This approach is called the Minimum Description Length

                 MDL – Minimize : size(tree) + size (misclassifications(tree))


Labels: ,

Discussion & Queries (<$I18NNumComments$>):

<$CommentPager$>
<$I18NAtCommentTimeWithPermalink$>, <$I18NCommentAuthorSaid$>

<$BlogCommentBody$>

<$BlogCommentDeleteIcon$>
<$CommentPager$>