BOUESTI Logo
BOUESTI SPAM DETECTORDept. of Computing & Info Science

Model Performance & Results

Evaluation metrics and performance plots derived from training 4 machine learning models on 33,716 Enron emails with 5-fold cross-validation.

Selected Best Model for Deployment

Logistic Regression

Selected for production deployment based on achieving the highest overall F1-Score (0.9896) and highest Accuracy (99.00%) with low false negative rates.

99.00%
Accuracy
0.9896
F1-Score

4-Model Algorithm Performance Comparison

Test Split: 6,097 Emails
Model AlgorithmAccuracyPrecisionRecallF1-ScoreTraining TimeStatus
Logistic Regression99.00%98.57%99.35%0.98965.18sDEPLOYED
Support Vector Machine98.95%98.57%99.25%0.98919.15sEVALUATED
Random Forest98.44%97.73%99.04%0.9838265.48sEVALUATED
Multinomial Naive Bayes98.39%98.35%98.28%0.98327.04sEVALUATED

Python Training Visualizations & Plots

Original graphical plot figures generated directly from the scikit-learn model evaluation pipeline.

1. Model Metrics Comparison

Comparison of Accuracy, Precision, Recall, and F1-Score across all 4 algorithms.

1. Model Metrics Comparison

2. Confusion Matrices

True Positive, True Negative, False Positive, and False Negative counts for each model.

2. Confusion Matrices

3. ROC Curves & AUC Score

Receiver Operating Characteristic curves showing false positive vs true positive rate tradeoffs.

3. ROC Curves & AUC Score

4. Training Time Comparison

Execution time in seconds for model training and 5-fold cross-validation.

4. Training Time Comparison

5. Feature Importance & Top Words

Top TF-IDF feature weights that indicate spam vs ham classification.

5. Feature Importance & Top Words