Research synthesis · Machine learning & data science

Applied machine learning and data science: rigorous evaluation over headline accuracy

Vajjhala and colleagues show that in applied machine learning — from weed classification and weather prediction to medical imaging, software defects and project analytics — multi-seed statistical comparison, metrics beyond accuracy and leakage checks matter more than headline scores, and that simple models often match complex ones.

A short narrative review of research by Narasimha Rao Vajjhala and co-authors in this area. Every claim carries an APA citation that links to the publication’s own page (abstract, key findings, DOI); the full references are listed at the end. Updated .

Synthesis

Across applied studies in agriculture, meteorology, medical imaging, software engineering and business analytics, Narasimha Rao Vajjhala and colleagues have argued that the value of a machine learning model depends less on its headline accuracy than on how rigorously it is evaluated. Kumar et al. (2026) compared six deep learning configurations for nine-class weed classification on the DeepWeeds benchmark and demonstrated that, although fine-tuned DINOv2 and EfficientNet-B0 performed best, the differences among several high-performing models were not consistently statistically significant across random seeds. An exploratory test on Albanian field images exposed a clear geographic domain shift, showing that benchmark rankings do not transfer automatically to new farming regions.

A second recurring result is that simpler models are often sufficient. Using ten years of daily data for Sofia, Berlin and Tokyo, Gigov et al. (2025) found that linear regression was statistically equivalent to random forest for temperature (R² > 0.994) at a fraction of the run time, and clearly better for humidity. The same pattern appears in project analytics: across 439 successful US government IT projects, linear regression predicted project manager commitment tenure better (r² = 0.238) than random forest or support vector machines (Vajjhala & Strang, 2024a; Strang & Vajjhala, 2025b).

Vajjhala and colleagues also insist on metrics and diagnostics beyond accuracy. Saheed et al. (2021) evaluated a seven-model boosting and bagging ensemble for software defect prediction with AUC, F-measure and the Matthews correlation coefficient, reporting outstanding performance for ensemble CatBoost. Haveri and Vajjhala (2026) reported bootstrapped confidence intervals, calibration and CLAIM and TRIPOD-AI compliance for a pneumonia detector that reached an AUC of 0.967 and pneumonia recall of 0.98. Most pointedly, Strang and Vajjhala (2026b) showed that a perfect in-sample classifier (AUC = 1.000) trained on supply chain records was an artefact of label leakage from near-duplicate items, and a companion study found that classifiers learned little from a firm’s own ESG ratings (kNN AUC = 0.497) (Strang & Vajjhala, 2026a).

Where the data are informative, the group has used machine learning to support decisions. Random forest analysis of about 17,430 US government IT projects identified seven failure indicators with 79.9% precision and a ROC area of 0.849 (Strang & Vajjhala, 2023a), and a random forest model best explained the performance gains of Albanian SMEs that piloted data analytics (R² = 0.661) (Vajjhala & Strang, 2024b). Earlier work applied correspondence analysis to 125,087 open global terrorism records (Vajjhala et al., 2015), surveyed visual data mining for collaborative filtering (Biba et al., 2017), and built recommender, web-prefetching and face-recognition systems (Vajjhala et al., 2021b; Ajibesin et al., 2025; Imoh et al., 2023).

Taken together, this body of work supports three practical claims for applied machine learning: compare models across repeated runs with statistical tests rather than single leaderboard scores; check training data for leakage before trusting near-perfect accuracy; and prefer the simpler, more interpretable model when performance is equivalent.

Key claims with sources

  1. Differences among high-performing deep learning models for weed classification were not consistently statistically significant across random seeds (Kumar et al., 2026).
  2. Linear regression matched random forest for hyperlocal temperature prediction and outperformed it for humidity (Gigov et al., 2025).
  3. A perfect in-sample classifier (AUC = 1.000) on organizational records was an artefact of label leakage (Strang & Vajjhala, 2026b).

References

  • Ajibesin, A. A., Vajjhala, N. R., Joel, E., & Rakshit, S. (2025). Predictive Web Prefetching: A Combined Approach Using Clustering Algorithms and WEKA in High-Traffic Settings. In Frank Lin, David Pastor, Nishtha Kesswani, Ashok Patel, Sushanta Bordoloi, Chaitali Koley (Eds.), Artificial Intelligence in Internet of Things (IoT): Key Digital Trends: Proceedings of 8th International Conference on Internet of Things and Connected Technologies (ICIoTCT 2023) (pp. 221–231). Springer Nature Singapore. https://doi.org/10.1007/978-981-97-5786-2_17 Summary & key findings →
  • Biba, M., Vajjhala, N. R., & Nishani, L. (2017). Visual Data Mining for Collaborative Filtering: A State-of-the-Art Survey. In Collaborative Filtering Using Data Mining and Analysis (pp. 217–235). IGI Global. https://doi.org/10.4018/978-1-5225-0489-4.ch012 Summary & key findings →
  • Gigov, L., Vajjhala, N. R., & Stoilov, A. (2025). Hyperlocal Temperature and Humidity Prediction Using Supervised Machine Learning. In 2025 2nd Global AI Summit – International Conference on Artificial Intelligence and Emerging Technology (AI Summit) (pp. 940–945). IEEE. https://doi.org/10.1109/AISummit66170.2025.11411099 Summary & key findings →
  • Haveri, K., & Vajjhala, N. R. (2026). Deep Learning for Medical Image Analysis: CNN-based Pneumonia Detection on Chest X-Rays. In 2026 6th International Conference on Pervasive Computing and Social Networking (ICPCSN) (pp. 156–161). IEEE. https://doi.org/10.1109/ICPCSN68523.2026.11543623 Summary & key findings →
  • Imoh, N., Vajjhala, N. R., & Rakshit, S. (2023). Experimental Face Recognition Using Applied Deep Learning Approaches to Find Missing Persons. In Subhadip Basu, Dipak Kumar Kole, Arnab Kumar Maji, Dariusz Plewczynski, Debotosh Bhattacharjee (Eds.), Proceedings of International Conference on Frontiers in Computing and Systems: COMSYS 2021 (pp. 3–11). Springer Nature Singapore. https://doi.org/10.1007/978-981-19-0105-8_1 Summary & key findings →
  • Kumar, R., Vajjhala, N. R., Ramollari, E., & Joca, E. (2026). A Comparative Evaluation of Deep Learning Architectures for Weed Classification, with an Exploratory Out-of-Distribution Analysis of Albanian Field Images. Information, 17(8), 751. https://doi.org/10.3390/info17080751 Summary & key findings →
  • Saheed, Y. K., Longe, O., Baba, U. A., Rakshit, S., & Vajjhala, N. R. (2021). An Ensemble Learning Approach for Software Defect Prediction in Developing Quality Software Product. In Mayank Singh, Vipin Tyagi, P. K. Gupta, Jan Flusser, Tuncer Ören, V. R. Sonawane (Eds.), Advances in Computing and Data Sciences: 5th International Conference, ICACDS 2021, Nashik, India, April 23–24, 2021, Revised Selected Papers, Part I (pp. 317–326). Springer International Publishing. https://doi.org/10.1007/978-3-030-81462-5_29 Summary & key findings →
  • Strang, K. D., & Vajjhala, N. R. (2023a). Mining Project Failure Indicators From Big Data Using Machine Learning Mixed Methods. International Journal of Information Technology Project Management, 14(1), 1–24. https://doi.org/10.4018/IJITPM.317221 Summary & key findings →
  • Strang, K. D., & Vajjhala, N. R. (2025b). Exploring project manager commitment using machine learning on fuzzy big data. International Journal of Project Organisation and Management, 17(2), 135–152. https://doi.org/10.1504/IJPOM.2025.146727 Summary & key findings →
  • Strang, K. D., & Vajjhala, N. R. (2026a). Machine Learning and Data Science for ESG Compliance Measurement in Financial Software Engineering Projects: A Single-Case Socio-Technical Systems Analysis. Systems, 14(8), 990. https://doi.org/10.3390/systems14080990 Summary & key findings →
  • Strang, K. D., & Vajjhala, N. R. (2026b). Regulatory Volatility in Digital Supply Chains: An Information Systems Analytics Study of Decision-Maker Risk Perceptions and Sustainability-Related Outcomes. Sustainability, 18(18), 9359. https://doi.org/10.3390/su18189359 Summary & key findings →
  • Vajjhala, N. R., & Strang, K. D. (2024a). An Exploratory Big Data Approach to Understanding Commitment in Projects. In Á. Rocha (Ed.), WorldCIST 2024 (World Conference on Information Systems and Technologies) (pp. 66–75). Springer Nature Switzerland AG. https://doi.org/10.1007/978-3-031-60227-6_6 Summary & key findings →
  • Vajjhala, N. R., & Strang, K. D. (2024b). Profitability, effectiveness, operational efficiency, and market growth of SMEs in Albania after piloting data analytics. International Journal of Services and Standards, 14(1), 51–64. https://doi.org/10.1504/IJSS.2024.140078 Summary & key findings →
  • Vajjhala, N. R., Rakshit, S., Oshogbunu, M., & Salisu, S. (2021b). Novel User Preference Recommender System Based on Twitter Profile Analysis. In Samarjeet Borah, Ratika Pradhan, Nilanjan Dey, Phalguni Gupta (Eds.), Soft Computing Techniques and Applications: Proceeding of the International Conference on Computing and Communication (IC3 2020) (pp. 85–93). Springer Singapore. https://doi.org/10.1007/978-981-15-7394-1_7 Summary & key findings →
  • Vajjhala, N. R., Strang, K. D., & Sun, Z. (2015). Statistical Modeling and Visualizing Open Big Data Using a Terrorism Case Study. In 2015 3rd International Conference on Future Internet of Things and Cloud (FiCloud), Rome, Italy (pp. 489–496). IEEE. https://doi.org/10.1109/FiCloud.2015.15 Summary & key findings →

How to cite this synthesis

Please cite the original publications above for specific findings. To cite this overview itself:

Vajjhala, N. R. (2026, September 25). Applied machine learning and data science: rigorous evaluation over headline accuracy. Narasimha Rao Vajjhala. https://www.narasimharao.net/research/focus-areas/machine-learning-data-science/

Markdown version