H₁: There is a statistically significant difference in cyber-threat detection performance among the selected Machine Learning algorithms.
Where scientifically justified, formulate additional hypotheses concerning:
Class imbalance
Feature selection
Model performance
Do not create hypotheses for purely qualitative or descriptive research questions.
Clearly identify:
Independent variables
Dependent variables
Control variables
Measurement variables
11. SIGNIFICANCE OF THE STUDY
Explain the significance separately for:
11.1 Academic Contribution
Contribution to research on ML-based cyber-threat detection in Ethiopia and comparable low-resource environments.
11.2 Technical Contribution
Potential contribution to:
ML model comparison
Feature selection
Cyber-threat classification
Dataset evaluation
Lightweight detection
11.3 Ethiopian Contribution
Explain how findings may inform future cybersecurity research or system development in:
Financial institutions
Government organizations
Telecommunications
Universities
Other organizations
Do not claim direct national-security impact unless supported by the actual research design.
11.4 Student Contribution
Explain the practical skills developed in:
Python
Data analysis
Machine Learning
Cybersecurity
Experimental research
Statistical evaluation
12. SCOPE OF THE STUDY
Define a strict and manageable scope.
Geographic Scope
Ethiopia.
Technical Scope
Machine Learning-based early cyber-threat detection.
Threat Scope
Only threats represented in the selected dataset.
Do not claim to study every cyber threat affecting Ethiopia.
Dataset Scope
Specify whether the research uses:
Public international datasets
Ethiopian datasets, if legally available
Synthetic data
A combination
Algorithm Scope
Select approximately 3–5 algorithms.
Consider:
Logistic Regression
Decision Tree
Random Forest
Support Vector Machine
XGBoost
Select the final algorithms based on literature, dataset characteristics, interpretability, and computing requirements.
Do not include deep learning merely because it is currently popular.
13. LIMITATIONS OF THE STUDY
Discuss realistic limitations, particularly:
Limited Ethiopian cybersecurity datasets
Dependence on public datasets
Lack of live institutional validation
Dataset imbalance
Dataset distribution differences
Limited computing resources
Limited research duration
Limited access to institutional cybersecurity data
For every major limitation, explain an appropriate mitigation strategy.
14. CONCEPTUAL FRAMEWORK
Develop a conceptual framework connecting:
INPUT
Cybersecurity dataset
Network traffic
Network features
Attack labels
↓
PREPROCESSING
Data cleaning
Missing-value treatment
Duplicate removal
Encoding
Scaling
Feature selection
Class balancing
↓
ML MODELS
Model A
Model B
Model C
Model D, if justified
↓
OUTPUT
Normal traffic
Malicious traffic
Threat category
Prediction probability
↓
EVALUATION
Precision
Recall
F1-score
False-positive rate
ROC-AUC
PR-AUC
Computational cost
↓
ETHIOPIAN APPLICABILITY
Dataset transferability
Distribution shift
Computing requirements
Data availability
Institutional constraints
Privacy considerations
Explain the framework in academic prose.
15. RESEARCH METHODOLOGY
Recommend a quantitative experimental and comparative research design if supported by the research questions.
Explain:
15.1 Research Approach
Why quantitative experimental research is appropriate.
15.2 Research Design
Explain the comparative ML experiment.
15.3 Research Process
Use:
> Problem Definition → Literature Review → Dataset Selection → Data Exploration → Preprocessing → Feature Selection → Model Training → Validation → Testing → Performance Comparison → Transferability Analysis → Conclusion
16. DATASET STRATEGY
This section must directly answer RQ4.
Investigate credible cybersecurity datasets, including where appropriate:
CICIDS2017
UNSW-NB15
NSL-KDD
Other newer and credible datasets
For each dataset provide:
August 15, 2026 238 8