How The System Works
An architectural breakdown of the natural language processing, TF-IDF feature extraction, and machine learning classification pipeline.
System Architecture Pipeline
1. Raw Email
Subject & Body
2. Preprocessing
Clean & Lemmatize
3. TF-IDF Vectors
10,000 Features
4. ML Inference
Supervised Model
5. Prediction
Spam or Ham
1
Data Preprocessing
- Emails are cleaned by stripping HTML tags, web URLs, email addresses, numbers, and special characters.
- Text is converted to lowercase and tokenized into individual word tokens.
- Common English stop words (e.g. "the", "is", "at") are filtered out.
- Words are reduced to their base root form using NLTK WordNet lemmatization.
2
Feature Extraction (TF-IDF)
- Term Frequency-Inverse Document Frequency (TF-IDF) converts text into numerical vectors.
- Important spam-related words receive high weights; frequent generic words receive low scores.
- Extracts unigrams and bigrams (single words and word pairs) for rich contextual understanding.
- Configured with a maximum vocabulary size of 10,000 features.
3
Machine Learning Models
- Multinomial Naive Bayes: Fast probabilistic classifier for text baselines.
- Support Vector Machine (SVM): Geometric hyperplane classifier optimized for high dimensions.
- Logistic Regression: Interpretable linear logit model selected for deployment.
- Random Forest: Robust decision-tree ensemble algorithm.
4
Evaluation & Deployment
- Trained using 5-fold cross-validation with 80% train, 10% validation, and 10% test split.
- Evaluated on Accuracy, Precision, Recall, F1-Score, ROC Curves, and Confusion Matrices.
- The best-performing model (Logistic Regression, F1: 0.9896) was selected for final web deployment.
- Predictions complete in under 10ms with confidence score calculations.
Enron Spam Dataset Specifications
Dataset Size
33,716 Emails
51% Spam (17,171) / 49% Ham (16,545)
Data Split Strategy
80% / 10% / 10%
Training / Validation / Testing
Preprocessed Clean Corpus
30,483 Unique Emails
3,222 duplicates removed