As organizations increasingly rely on open source software, proactively securing software dependencies has become a major cybersecurity challenge. In this master’s thesis, you will research how AI techniques, including Large Language Models (LLMs), can predict vulnerabilities before they emerge. You will develop a predictive framework that combines source code analysis with project and ecosystem metadata, enabling organizations to identify high-risk packages and improve the resilience of their software supply chain.
These packages are continuously updated, forked, and maintained by contributors across the world. While they significantly accelerate software development, they also introduce security risks. Vulnerabilities in dependencies can propagate throughout entire software ecosystems, leading to large-scale security incidents such as the Shai-Hulud worm in 2025 and the axios supply-chain attack in March 2026.
Another well-known example is the XZ package hack in 2024, in which a malicious backdoor was introduced into a widely used Linux package. Fortunately, the attack was discovered by a Microsoft engineer before it could cause any damage. But this begs the question, how can future hacks be prevented?
Current vulnerability detection methods focus mainly on static analysis and dependency scanning using CVE databases. However, these approaches detect vulnerabilities after disclosure, rather than proactively. With the rapid pace of open source development, organizations lack tools to predict which packages are likely to become vulnerable in the near future.
Recent advancements in Large Language Models (LLMs) and other AI techniques create new opportunities to address this challenge. By analyzing source code, contributor behavior and project metadata in new ways. For example, signals such as sudden maintainer changes, irregular commit history, or risky coding practices could serve as early-warning indicators of potential vulnerabilities. This could even be used to prevent vulnerabilities from ending up in the software.
The Assignment
Develop a predictive framework that combines code analysis and ecosystem metadata to estimate the risk of vulnerabilities emerging in a software package. Validate the framework using real-world open source datasets and compare its predictive performance with existing vulnerability detection methods. Deliver a proof of concept that demonstrates how organizations can use these predictions to proactively strengthen their supply chain security.
Datasets:
-
CVE databases
-
GitHub metadata
-
Package feeds (NPM, NuGet, PyPI, Maven)
About Info Support
Info Support specializes in custom software, data/AI solutions, management, and training and is active in the Finance, Industry, Agriculture, Food & Retail, Mobility & Public, and Healthcare sectors. We provide solid and innovative solutions for complex and critical software issues. Our headquarters are located in Veenendaal (NL) and Mechelen (BE). At present, approximately 500 employees are employed by Info Support.
Info Support’s working method is characterized by a number of core values: solidity, integrity, craftsmanship, and passion. These core values are intertwined in our work and the way we interact with each other.
To ensure that all employees are always up to date with the latest developments, Info Support has an in-house IT Academy that eagerly satisfies the hunger for more or different knowledge and skills.
B2 language proficiency in Dutch is required.























