Conference paper · 2024

Exploring Cybersecurity Risks in Higher Education Environments with Machine Learning

Kenneth David StrangiD & Narasimha Rao VajjhalaiD

2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN), IEEE · Published

ScopusIEEE XploreWeb of Science

Summary

What question does this paper answer?

How can universities assess their risk of a cybersecurity breach by analyzing hypertext data from their websites with unsupervised machine learning?

What did the study find?

Hypertext data (URLs, embedded XML and HTTP scripts) were extracted from about 848 records of U.S. higher education institution websites, containing over 4,000 potential cybersecurity indicators, and analyzed with t-SNE, radial analysis and correspondence analysis. A combined 'control' feature scored breach likelihood from 1 to 3; no malware breaches were found in the sample, but higher risk scores concentrated in parts of the visualization, and some links did not resolve to working sites.

Why does it matter?

Universities run open, complex and decentralized IT infrastructures that make them targets for phishing, ransomware, DDoS attacks and insider threats. The authors argue that website-based machine learning risk scoring lets decision-makers in educational institutions spot anomalies that may indicate cybercrime or impending attacks, while method and data triangulation are needed to verify the reliability and validity of machine learning outputs.

Key findings

  1. The study extracted hypertext data from roughly 848 records of U.S. higher education institution websites, yielding over 4,000 elements identified as potential cybersecurity indicators, including URLs, XML embeddings and HTTP scripts.
  2. Unsupervised machine learning (t-distributed stochastic neighbor embedding, t-SNE) was used to train a model identifying likely indicators of cybersecurity breach risk on university websites, supported by radial analysis and correspondence analysis.
  3. Radial analysis identified the top five website features related to potential cybersecurity breaches, and a combined 'control' feature summarizing the other indicators was the most useful feature for flagging potentially very high-risk components.
  4. University websites were scored for cybersecurity breach likelihood on a 1–3 scale, revealing a diverse range of security postures across the sampled higher education institutions.
  5. No instances of malware breaches were identified within the university sample, but higher risk scores were concentrated towards the middle and top of the t-SNE visualization, suggesting a link between website feature complexity and security vulnerability.
  6. Some university website links could not be verified because the URL did not resolve to a working site (dead-end sites or orphaned links), which the analysis flagged separately.
  7. The authors conclude that machine learning for cybersecurity in higher education raises data reliability and model validity challenges that require method and data triangulation to verify outputs.

Source: Strang & Vajjhala (2024), 2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN), IEEE.

Study at a glance

Design and results of Exploring Cybersecurity Risks in Higher Education Environments with Machine Learning
Research questionHow can sampled universities assess the risk of a cybersecurity breach by analyzing their websites?
DesignComparative, exploratory study applying unsupervised machine learning to publicly available website data from U.S. higher education institutions.
DataRoughly 848 records from educational websites with over 4,000 elements identified as potential cybersecurity indicators (URLs, XML embeddings, HTTP scripts); a medium-sized subset of 26 was narrowed to two private and two public universities in Pennsylvania, Connecticut, New Jersey and California.
MethodsPython-based machine learning to scan websites, t-SNE to train a model of breach-risk indicators, and radial and correspondence analysis to visualize risk signals.
Main resultA combined 'control' feature scored breach likelihood on a 1–3 scale; no malware breaches were detected, but higher risk scores clustered in the central and upper parts of the t-SNE map.
ImplicationDecision-makers can use such machine learning risk maps to target cybersecurity interventions, provided outputs are verified through method and data triangulation.
CitationStrang & Vajjhala (2024)

Abstract

This paper uses unsupervised machine learning techniques (ML) to explore the risk of cybersecurity breaches or attacks in the higher education sector, including schools, universities, and support organizations. A large sample of higher education institutions (N=848) was analyzed by extracting hypertext data from the websites of educational institutions to identify links that may signal a future cybersecurity breach or attack. ML T-distributed Stochastic Neighbor Embedding (t-SNE) was used to train a model for identifying likely indicators of cybersecurity breach risks. The authors used additional techniques, including radial analysis and correspondence analysis, to visualize the cybersecurity breach signals in the data. In this way, decision-makers should be able to spot anomalies that indicate cybercrime activity or potential cyber-attacks. This paper also exposes the problems of using ML and how to use methods and data triangulation to check reliability and validity.

Abstract as published in 2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN).

Keywords: unsupervised machine learning; cybersecurity breaches; higher education sector; hypertext data analysis; t-distributed stochastic neighbor embedding; pervasive computing; cyber-attack indicators; radial analysis; correspondence analysis; anomaly detection; data triangulation

Key terms

t-SNE (t-distributed stochastic neighbor embedding)
A machine learning technique that reduces high-dimensional data to a lower-dimensional map by minimizing the Kullback-Leibler divergence between pairwise similarity distributions, keeping similar points close and dissimilar points apart.
Unsupervised machine learning
Machine learning that finds structure or patterns in data without a labeled outcome variable.
Data triangulation
Checking the reliability and validity of findings by comparing results across multiple methods or data sources.

Limitations

  • The sample focused on medium-sized universities; the authors recommend future research include large and small institutions to test whether the identified risks and indicators hold across institution types and sizes.
  • Leveraging machine learning for cybersecurity raises challenges of data reliability and model validity, requiring rigorous verification of machine learning outputs.
  • The authors call for refined feature selection, comparison with deep learning and ensemble methods, additional data such as network traffic and email patterns, and longitudinal studies.

How to cite

Strang, K. D., & Vajjhala, N. R. (2024). Exploring Cybersecurity Risks in Higher Education Environments with Machine Learning. In 2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN). IEEE. https://ieeexplore.ieee.org/document/10607755

BibTeX
@inproceedings{strang2024cybersecurity,
  title = {Exploring Cybersecurity Risks in Higher Education Environments with Machine Learning},
  author = {Strang, Kenneth David and Vajjhala, Narasimha Rao},
  booktitle = {2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN)},
  year = {2024},
  publisher = {IEEE},
  url = {https://ieeexplore.ieee.org/document/10607755}
}
Download citation:BibTeXRISCSL-JSONMarkdown