1. Target Project: [Waste_Classification_Multiclass_CNN_Pytorch] https://github.com/pankajsen29/Waste_Classification_Multiclass_CNN_Pytorch
this is the main multiclass waste image classifier model which is used as the target/victim by the attacker/user, whose weights and parameters are hidden, but it returns prediction results to user queries. With these results attacker intends to build a copycat model, which is termed as surrogate model.
2. Surrogate (i.e., extracted model) Project: [Waste_Classification_Model_Extraction] https://github.com/pankajsen29/Waste_Classification_Model_Extraction
this is the copycat model attacker builds using the prediction results obtained via querying the target model.
3. Evaluation (of extracted model) Project: [Waste_Classification_Model_Extraction_Evaluation] [ -> CURRENT REPOSITORY] https://github.com/pankajsen29/Waste_Classification_Model_Extraction_Evaluation
this is the independent evaluation project for the extracted waste classification model which evaluates:
i) Fidelity or Agreement rate: Target Prediction vs Surrogate Prediction,
ii) confidence similarity: Validation KL Divergence (Target score vector vs Surrogate score vector) etc.
Test images are predicted both by the target and the surrogate models and below metrices are computed for the surrogate prediction results by comparing the corresponding prediction from the target.
Test Dataset (images from test folder of the dataset): https://www.kaggle.com/datasets/saimonv/trashbox
Output files: target_results.jsonl surrogate_results.jsonl
Steps:
1) iterates all the images from DATA_DIR,
2) queries target API for each of the images for prediction,
3) then surrogate target API for each of the images for prediction,
4) Computes the below metrices:
- Fidelity,
- KL Divergence,
- Mean Squared Error (MSE),
- Root Mean Squared Error (RMSE).
Display Output:
Target query results is stored in:
C:\xxx\Waste_Classification_Model_Extraction_Evaluation\data\Query_Results\target_results.jsonl
Surrogate query results is stored in:
C:\xxx\Waste_Classification_Model_Extraction_Evaluation\data\Query_Results\surrogate_results.jsonl
Waste Classification Model Extraction Evaluation
----------------------------------------
Fidelity : 0.7037
Avg KL Div : 0.3837
Avg MSE : 0.0175
Avg RMSE : 0.1321
Conclusions on the evaluation results:
-
Fidelity = agreement rate, (i.e., percentage of queries where surrogate predicts exactly the same class as the target)
- Conclusions:
- Better than random guessing
- Evidence that extraction succeeded to some extent
- Not a high-fidelity clone
- Conclusions:
-
KL Divergence = Measures how close the full probability distributions are. Lower is better.
- Example:
- target (Plastic: 0.95, Paper: 0.03, Glass: 0.01, Metal: 0.01),
- surrogate (Plastic: 0.60, Paper: 0.20, Glass: 0.10, Metal: 0.10)
- Both predict "Plastic", but the confidence profiles differ considerably, which increases KL divergence.
- Conclusions:
- Even when they predict the same class, confidence levels can differ substantially.
- Example:
-
MSE Average squared difference between class probabilities. Lower is better.
-
RMSE: 0.1321 means on average, the surrogate's predicted probability for a class differs from the target's probability by about 13%
Notes:
- Accuracy on true labels is not checked as there are class mismatch between training (RealWaste) and test datasets (TrashBox).
- But it doesn’t matter because these metrices compare the target's outputs against the surrogate's outputs, not against TrashBox labels. And this is exactly is our goal.