Slhoub, Khaled
Khaled Ali Slhoub
Associate Professor | College of Engineering and Science: Department of Electrical Engineering and Computer Science
Program Chair
Program Chair | Computer Information Systems and Human Centered Design
Contact Information
Expertise
Personal Overview
Dr. Khaled Slhoub is an Associate Professor in the Department of Electrical Engineering and Computer Science at Florida Institute of Technology. He also serves as the Program Chair of Computer Information Systems & Human-Centered Design. Dr. Slhoub earned his Ph.D. in Computer Science from Florida Tech, focusing on developing a standard framework for formalizing the analysis process of agent-based systems. His academic journey includes an M.Sc. from the University of New Brunswick in Canada and a B.Sc. from Benghazi University in Libya.
His primary research interests lie in software engineering, with a current focus on the evaluation, testing, and trustworthy use of large language models in software development. His work investigates hallucination detection in AI-generated software artifacts, self-evaluation and confidence calibration in generative AI systems, and multi-agent architectures for improving the reliability of LLM outputs. He also studies agentic AI frameworks for mining and analyzing open-source software repositories, and continues a long-standing research line on detecting and modeling the behavior of social bots in networking platforms, including the emerging challenge of LLM-driven social agents. This research builds on his earlier contributions in agent-oriented methodologies, software quality, and requirements engineering. His teaching reflects these interests, including a graduate course on software engineering with AI-assisted tools.
Educational Background
Ph.D., Computer Science, Florida Institute of Technology, USA 2018
M.Sc., Computer Science, University of New Brunswick, Canada 2008
B.Sc., Computer Science, University of Benghazi, Libya 1999
Current Courses
Fall 2026:
CSE 4425 - Software Testing
CSE 5625 - Advanced Software Testing
CSE 1101 - Computer Disciplines and Careers
Previously Taught:
Software Development with AI-assisted tools, Intro. to Software Engineering, Software Testing, Advanced Software Testing, Database Systems, Software Quality and Metrics, Software Development, C++ Programming, Data Structures in C, OOP Concepts (Java), Introduction to Computer Science, Introduction to Computer Programming, Operating Systems, Concepts of Programming Languages, Computer Disciplines and Careers, and Introduction to Engineering
Selected Publications
- A. Alharbi, K. Slhoub, A Systematic Review of Artificial Intelligence for Software Requirements in Open-Source Software, in IEEE-Access, vol. 14, pp. 111509-111532 (2026).
- D. Yohn, L. Flancher, M. Islam, K. Slhoub, Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment. arXiv preprint 2026 arXiv:2606.11672.https://ieeexplore.ieee.org/abstract/document/11614799.
- N. Alkathiri, K. Slhoub, Challenges in Machine Learning-based Social Bot Detection: A Systematic Review, Discover Artificial Intelligence 5, 214 (2025). https://doi.org/10.1007/s44163-025-00448-w
- C. Miskell, K. Slhoub, N. Mensi, Performance and Self-evaluation of Free-Access Large Language Model Chatbots: A Multi-metric Comparative Study, In Proceedings of the Future Technologies Conference (FTC) 2025, Volume 4. Lecture Notes in Networks and Systems, vol 1678. Springer, Cham. https://doi.org/10.1007/978-3-032-07992-3_17
- B. Sherifi, K. Slhoub, F. Nembhard, The Potential of Large Language Models in Automating Software Testing: From Generation to Reporting, In Proceedings of Software and Data Engineering Conference. IEEE-SEDE 2025. Communications in Computer and Information Science, vol 2720 . Springer, Cham. https://doi.org/10.1007/978-3-032-08649-5_13
- P. Pochiraju, V. Uppalapati, N. Mensi, K. Slhoub, "YOLOv10 for Real-Time Emergency Vehicles Detection in Intelligent Traffic Systems: A Comprehensive Comparative Analysis, 2025 IEEE International Conference on Advanced Systems and Emergent Technologies (IC_ASET), Mammamet-Yasmine, Tunisia, 2025, pp. 1-6, doi: 10.1109/IC_ASET65966.2025.11232047.
- C. Miskell, R. Diaz, P. Ganeriwala, K. Slhoub, and F. Nembhard, Automated Framework to Extract Software Requirements from Source Code, accepted and will be published into ACM-NLPIR 2023 conference proceedings, South Korea, December 15-17, 2023.
- F. Nembhard, K. Slhoub, and M. Carvalho, An Agent-Based Approach Toward Smart Software Testing, accepted and will be published into Future Technologies Conference 2023, San Francisco, USA, November 2-3, 2023.
- J. Brenner, C. Sen, R. Weaver, K. Hamed, R. Mesa-Arango, K. Demoret, K. Slhoub, and M. Gaal, Reframing Making to Integrate Entrepreneurially Minded Learning (EML), 7th International Symposium on Academic Makerspaces (ISAM23), Pittsburgh, USA, October 18-20, 2023.
- A. Averza, K. Slhoub and S. Bhattacharyya, "Evaluating the Influence of Twitter Bots via Agent-Based Social Simulation," in IEEE Access, vol. 10, pp. 129394-129407, 2022.
- F. Alsuliman, S. Bhattacharyya, K. Slhoub, N. Nur, & C. Chambers, "Social Media vs. News Platforms: A Cross-analysis for Fake News Detection Using Web Scraping and NLP", In Proceedings of the 15th International Conference on PErvasive Technologies Related to Assistive Environments (PETRA '22), Association for Computing Machinery (ACM), USA, 190–196, 2022.
- B. Wood & K. Slhoub, "Detecting Amazon Bot Reviewers Using Unsupervised and Supervised Learning", (Conf Best Paper Award ), 2022 IEEE World AI IoT Congress (AIIoT), pp. 01-08, 2022.
- K. Slhoub, F. Nembhard & M. Carvalho, "A Metrics Tracking Program for Promoting High-Quality Software Development", IEEE SoutheastCon2019, pp. 1-8, 2019.
- K. Slhoub, M. Carvalho & F. Nembhard, "Evaluation and Comparison of Agent- Oriented Methodologies: A Software Engineering Viewpoint", IEEE SYSCON2019), 2019.
- K. Slhoub & M. Carvalho, "Towards Process Standardization for Requirements Analysis of Agent-Based Systems", Advances in Science, Technology and Engineering Systems Journal, June 2018.
- K. Slhoub, M. Carvalho & W. Bond, "Recommended Practices for the Specification of Multi-Agent Systems Requirements", 8th IEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conf. (IEEE UEMCON 2017), Columbia University, NY, USA, Oct 2017.
- Khaled Slhoub, "A Software Quality Resource Tool That Improves Quality Management of Scaled-Down Development Environments", Proc. 11th IASTED International Conf. on Software Engineering (IASTED SE 2012), International Association of Science and Technology for Development (IASTED), Crete, Greece, June 2012.
- Khaled Slhoub, "A Strategy That Improves Quality of Software Engineering Projects in Classroom", Proc. 10th International Arab Conf. on Information Technology (ACIT), University of Benghazi, Libya, Dec 2010.
- Khaled Slhoub, "Managing Software Quality in Educational and Small Business Environments (Poster)", Proc. 4th Annual Research Exposition on Information Technology (UNB CS), University of New Brunswick, Faculty of Computer Science, Fredericton, Canada, April 2007.
Recognition & Awards
2026 - Recognized for excellence in service
- Recipient of EECS Faculty Excellence Award for Service, Florida Tech
2025 - Recipient of FY25 COES Institutional Research Incentive (IRI)
- Seed grant funds - Project: Web Application for Aqualab Sensor Monitoring and Analysis, Florida Tech
2022 - Recipient of FY23 COES Institutional Research Incentive (IRI)
- Received IRI funding for the software engineering research - extracting software requirements directly from source code, USA
2022 KEEN Rising Star Award
- Named as Florida Tech's 2022 Campus Rising Star by the Kern Entrepreneurial Engineering Network National Organization (KEEN). Also be submitted for review and consideration for the National KEEN Rising Star award https://engineeringunleashed.com/content/2022-campus-keen-rising-stars
2013 Recipient of A Government Scholarship
- Selected by the Ministry of Higher Education and Scientific Research to undertake graduate studies in Computer Science in United State of America, Libya
Research
I am currently accepting new graduate students (Masters and Doctoral). Please reach out to me if you are interested and have SE/CS or other related background.
----------------------------
Trustworthy AI and LLM Evaluation
Evaluation and Testing of Large Language Models (LLMs)
The goal of this project is to explore and develop methods for evaluating and testing large language models (LLMs) to ensure their accuracy, fairness, and robustness across different domains. A major focus of the project will be on testing the performance of LLMs through various scenarios and stress-testing techniques. This includes the design of test cases for handling edge cases, adversarial inputs, and ambiguous queries to assess how well these models generalize and respond to challenging conditions.
Testing will involve both quantitative and qualitative approaches. Quantitative testing will use performance metrics like perplexity, accuracy, and response diversity, while qualitative testing will focus on human evaluation to determine how well the model responds to real-world, contextually complex questions. In addition to functional performance, tests will assess ethical concerns like bias detection, fairness, and handling of sensitive or inappropriate content. The project will also explore adversarial testing techniques, including metamorphic testing approaches that check consistency across related inputs, to evaluate the model's ability to resist attacks or manipulation, ensuring its robustness in unpredictable environments.
The project aims to produce reliable evaluation techniques and guidelines to make LLMs more reliable, robust, and ethically sound, helping improve their performance in diverse, real-world applications.
Multi-Agent Architectures for Hallucination Prevention in LLMs
Large language models can generate fluent but factually unsupported content, which limits their safe use in software engineering and other high-stakes domains. Single-model self-checking has shown only partial success in catching these errors.
This project investigates whether multi-agent architectures, in which dedicated verifier and critic agents review and challenge the output of a generator agent, can reduce hallucination rates more effectively than single-model approaches. The research examines agent orchestration strategies, disagreement resolution, and the cost of added inference overhead. The goal is to produce empirically validated design guidelines for building multi-agent pipelines that improve the factual reliability of LLM outputs.
Self-Evaluation and Confidence Calibration in Generative AI Systems
Generative AI systems often produce responses with high confidence, even when the information is incorrect or unsupported. This overconfidence can make it difficult for users to recognize errors, especially when AI systems are used in complex domains such as software engineering, research, or decision support.
This project investigates methods to evaluate and improve the self-evaluation capabilities of generative AI systems. The research explores whether AI models can accurately assess the correctness and reliability of their own outputs. The project will examine techniques for detecting cases where AI systems produce confident but incorrect responses and develop approaches to improve confidence calibration and reliability assessment in AI-generated outputs. The goal is to enhance the trustworthiness of generative AI systems and support safer adoption of AI tools in real-world applications.
Detecting Hallucinations in AI-Generated Software Requirements
Large Language Models are increasingly used to assist in generating and refining software requirements. However, these systems may produce hallucinated requirements, which are requirements that are not supported by the original problem description or stakeholder needs. Such inaccuracies can lead to incorrect system design and development decisions.
This project investigates techniques for detecting and reducing hallucinated requirements generated by AI systems. The research will explore methods for verifying whether AI-generated requirements are grounded in the source specification using techniques such as traceability analysis, consistency checking, and natural language processing. The goal is to develop evaluation frameworks and automated tools that improve the reliability and trustworthiness of AI-assisted requirements engineering.
Cost-Aware Evaluation of Local and Open-Weight Language Models
Organizations increasingly deploy open-weight language models locally for privacy, cost, and control reasons, yet published evaluations rarely account for the computational cost of achieving reported accuracy. A model that scores highest on a benchmark may be impractical to run in a small lab, classroom, or business setting.
This project evaluates local language models on reasoning tasks while jointly measuring accuracy and inference cost, including runtime, memory, and energy considerations. The research asks whether the most accurate local model is also the most expensive to operate, and how accuracy-per-cost trade-offs should inform model selection. The goal is to provide practical selection guidance for resource-constrained environments, extending earlier work on software quality management in scaled-down development settings.
AI in Software Engineering Practice and Education
Software Metrics for AI-Assisted Development
Software teams increasingly rely on AI coding assistants to generate a substantial share of their code, yet established quality metrics were designed and validated for human-written software. It remains unclear whether measures such as defect density, complexity, and maintainability behave the same way on AI-generated code, or how teams should measure quality in mixed human-AI codebases.
This project investigates how classical software metrics perform on AI-assisted projects and develops new metrics suited to this setting, including AI contribution ratio, prompt-to-code traceability, and verification effort. The research combines empirical analysis of human and AI-generated code with the design of a metrics tracking program for AI-assisted development teams. The goal is to give practitioners and educators reliable, evidence-based measures for managing quality in modern development environments, extending earlier work on metrics programs for promoting high-quality software development.
Trust Calibration in Human-AI Pair Programming
Developers working with AI coding assistants must constantly decide whether to accept, modify, or reject generated suggestions. Evidence suggests that both over-trust and under-trust are common, leading to defects slipping through or productivity gains being lost. How developers form and adjust trust in these tools is not well understood.
This project studies trust calibration in human-AI pair programming, with a focus on students and early-career developers. The research examines whether interface interventions, such as displaying model confidence or highlighting uncertain code regions, help users decide when to verify suggestions rather than accept them. The goal is to produce design guidance for AI coding tools and educational practices that foster appropriately calibrated trust, connecting software engineering with human-centered design.
LLM Agents as Teammates in Classroom Software Projects
AI assistants are typically used by student teams as ad hoc helpers, with little structure around their role in the development process. This leaves open the question of what happens when an LLM agent is treated as an actual team member with defined responsibilities, such as tester, reviewer, or documentation lead.
This project studies the integration of LLM agents into student software teams as structured teammates. The research measures the effects on product quality, process metrics, individual learning outcomes, and team dynamics, and examines how role assignment shapes how students engage with AI-generated contributions. The goal is to develop evidence-based models for incorporating AI agents into project-based software engineering education, supporting both learning effectiveness and responsible AI use.
Agentic AI for Conflict Detection in Open-Source Software Repositories
Doctoral student: Amal Alharbi
Open-source software development involves many contributors working asynchronously, which produces conflicts in code, requirements, and project decisions that are costly to detect and resolve manually. Repository histories contain rich signals about these conflicts, but mining them at scale requires more than static analysis.
This project develops an agentic AI framework in which specialized agents mine repository artifacts, including commits, issues, and pull request discussions, to detect and classify emerging conflicts. The research addresses agent coordination, grounding agent conclusions in repository evidence, and empirical validation against human-labeled conflict data. The goal is a framework that helps maintainers identify disruptive patterns early and improve the health of open-source projects.
Towards a Framework to Extract Software Requirements Directly from Source Code - Funded by FIT-IRI
This research project aims to develop an intuitive framework for analyzing source code files and generating requirement specifications to show the functions the code is trying to accomplish. Instead of attempting to comprehend and manually construct requirements of open-source systems or legacy systems, our proposed framework will extract requirements directly from source code. While some attempts exist, a fully functional system that efficiently fulfills this task has not been created. This research effort will help us contribute to the ever-growing body of knowledge concerning requirements elicitation for open-source software or legacy systems.
Social Agents and Computational Propaganda
A Framework to Model and Analyze the Behavior of Social Agents (Bots)
Co-PI: Dr. Siddhartha Bhattacharyya
The influence of social media in our daily lives has increased significantly over the last decade. This is evident from the change in stock prices guided by Reddit or the increase in participation of protesters/activists for a cause or the increase in new job opportunities, known as influencers. On platforms such as X (formerly Twitter), bot software can post content, respond to others, repost content, create relationships with other users, and direct message other accounts. One of the key factors in this increasing influence of social media is the encouragement or information provided by recommending social agents; the phenomenon is known as Computational Propaganda. Computational Propaganda summarizes automated systems that spread fake/false information and make use of data-driven methods to shape public opinion. Politicians, for example, misuse bots in social media platforms to increase voting participation, influence voters to support them, and promote their campaign platform. It has been demonstrated through societal outcomes that Computational Propaganda can have lasting and overreaching influence that might lead to dangerous or unsafe societal outcomes. The rise of LLM-driven social agents has made such automated behavior more fluent and substantially harder to detect.
This project's goal is to conduct research on developing a framework that leads to a better understanding, monitoring, and managing of the behavior of social agents to assure the integrity and prevent unsafe and unethical actions. This will help us contribute to the ever-growing body of knowledge concerning content manipulation by social agents that use social media to influence the stock market and society in general.
/prod01/fit-cdn-pxl/media/fit-website/site-assets/images/FT-Horiz_crimson-gold.png)