I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
As a part of my work I do security research for various apps and websites. I wanted to see if LLMs could reproduce a common class of exploits I’ve found in multiple apps.
I made a fake React Native app in Expo and a backend in Python. It’s a book review app and the goal is to find a flag in a user’s private reviews.
If you would like to try solving it yourself before I spoil it, here’s a ZIP of the APK and challenge description each LLM was fed.
Starting with the models that got 10 full runs:
Let’s go per model and then we’ll dig into the ones that didn’t get full 10 runs:
I also tried a few other models but due to the costs getting so high I didn’t do ten full runs of them, including them for completion’s sake: