BOUESTI Logo
BOUESTI SPAM DETECTORDept. of Computing & Info Science

How The System Works

An architectural breakdown of the natural language processing, TF-IDF feature extraction, and machine learning classification pipeline.

System Architecture Pipeline

1. Raw Email
Subject & Body
2. Preprocessing
Clean & Lemmatize
3. TF-IDF Vectors
10,000 Features
4. ML Inference
Supervised Model
5. Prediction
Spam or Ham
1

Data Preprocessing

  • Emails are cleaned by stripping HTML tags, web URLs, email addresses, numbers, and special characters.
  • Text is converted to lowercase and tokenized into individual word tokens.
  • Common English stop words (e.g. "the", "is", "at") are filtered out.
  • Words are reduced to their base root form using NLTK WordNet lemmatization.
2

Feature Extraction (TF-IDF)

  • Term Frequency-Inverse Document Frequency (TF-IDF) converts text into numerical vectors.
  • Important spam-related words receive high weights; frequent generic words receive low scores.
  • Extracts unigrams and bigrams (single words and word pairs) for rich contextual understanding.
  • Configured with a maximum vocabulary size of 10,000 features.
3

Machine Learning Models

  • Multinomial Naive Bayes: Fast probabilistic classifier for text baselines.
  • Support Vector Machine (SVM): Geometric hyperplane classifier optimized for high dimensions.
  • Logistic Regression: Interpretable linear logit model selected for deployment.
  • Random Forest: Robust decision-tree ensemble algorithm.
4

Evaluation & Deployment

  • Trained using 5-fold cross-validation with 80% train, 10% validation, and 10% test split.
  • Evaluated on Accuracy, Precision, Recall, F1-Score, ROC Curves, and Confusion Matrices.
  • The best-performing model (Logistic Regression, F1: 0.9896) was selected for final web deployment.
  • Predictions complete in under 10ms with confidence score calculations.

Enron Spam Dataset Specifications

Dataset Size
33,716 Emails
51% Spam (17,171) / 49% Ham (16,545)
Data Split Strategy
80% / 10% / 10%
Training / Validation / Testing
Preprocessed Clean Corpus
30,483 Unique Emails
3,222 duplicates removed