This blog was originally published by Ryan Chaplin on the Raxis blog May 5, 2026

AI is on everyone’s minds today. As a penetration tester, AI is of specific interest for several reasons. It can allow everyone (including malicious hackers) to accomplish more in less time, but it can also make mistakes, miss things, and unintentionally cause harm. Simply put, AI on its own can miss things that humans are much better at, and it can be tricked into doing unintended things. That gap between what AI produces and what an attacker can actually do is exactly why our penetration testing services still put a human behind every exploit.

That’s what I’d like to examine here. AI models are often programmed not to do illegal or malicious tasks, but there are ways to get around that, and the details and results are interesting.

The Scenario

Today we will be bypassing security restrictions on GPT-OSS-120B, which is ChatGPT’s open-source model.