4,500 break-in attempts, nothing taken: is software built with AI safe?
Case study
Late on a Saturday night, someone started trying to break into an app we had built with AI for one of our clients, a digital marketing agency. They kept at it for four days, made more than 4,500 password guesses, changed disguises every time we blocked them, and left with nothing.
Here is how those four days went, and what they say about a question we hear a lot: is software built with AI actually safe?
Saturday night: a stranger at the door
The app is an internal tool. It holds the agency’s advertising data for all of its clients, and a few of its buttons move real money: raise a budget, change a bid, pause an ad.
At 21:38 Budapest time, a rented server across the Atlantic started talking to it, skipping the website and going straight to the database. That is less alarming than it sounds: most modern apps let your browser talk to their database directly, the way anyone can ring the bell of an office building. What matters is what happens when a stranger rings.
On Sunday the agency’s owner messaged me: staff were receiving sign-in and password-reset emails nobody had asked for. Was there a security hole?
What the attacker tried
The logs read like a burglar’s checklist.
- They mapped the building. By asking for rooms that don’t exist, they got the database to suggest the real names, so they soon had a floor plan, without a single thing from inside.
- They tried every door, plus a hundred guessed ones like “secrets” and “payment keys”. Every answer was an empty list.
- They guessed passwords, hundreds on the first night.
- They tried a con, forging password resets to send the link to their own website. The system ignored it.
- They changed disguises, moving to an anonymising network and a new address for almost every guess once we limited attempts per address.
This was no random bot: within seconds, their word list named one of the agency’s own clients, a payment company. They were most likely after its payment keys, through a supplier they hoped was less guarded. Those keys were never stored in the app.
Why nothing came out
I built the sign-in and permission layer of this app with Claude as my AI coding partner (here is how we work with it). Before a line of code existed, we wrote down two rules.
First, the database checks every single record against who is asking: not just whether you may enter the building, but whether you may see this particular record. A stranger matches nothing, so a stranger gets an empty list.
Second, the screens are not the security. Hiding a button from the screen doesn’t stop anyone from triggering what’s behind it with a simple program, so every rule that matters lives in the database itself.
Over four days, the attacker did not read, change or delete a single record, or get into a single account.
The gap we found ourselves
To see what an attacker could do with a real login, we stopped reading our rules and tried to break them. A customer account that should only see its own reports turned out to be able to change a few of its own settings that the screens never offer. Nobody else’s data was exposed, and the attacker never had a login to try it with.
We closed the gap that afternoon and added an automated check that tries everything each kind of user must not do, so a gap like this gets caught before it goes live. Reading the settings had not revealed it, but trying the forbidden thing did.
Sunday: closing the doors
The AI did the legwork, reading the logs of three systems and safely rehearsing every change, while I approved each step (why that matters). In one day we:
- required stronger passwords, checked against lists of leaked ones;
- added a “prove you’re not a robot” check to the sign-in page;
- shut strangers out completely, so instead of an empty list they get no answer at all;
- removed the stored password from every account that didn’t need one. Those users now sign in with an emailed link, and nobody can guess a password that doesn’t exist.
The robot check went live at 16:49, and within five minutes the attack went quiet.
At one point the AI reported that the attack had stopped. I asked how it knew, and a second look at a longer stretch of the logs showed the attacker was still going. That question is the part of the work no AI does for you.
Sunday night and Monday: the comeback
Three hours later they were back with a robot driving a real web browser. It opened our sign-in page, passed the robot check and fired off 3,629 guesses overnight, about 80% of them getting past the check. Every one failed, because there were no passwords left to find.
On Monday the logs showed how it was done. Anyone who passes the robot check gets a short-lived proof from our sign-in page that they are human, and a password can only be tried with that proof. The robot simply kept reopening our page and collecting fresh proofs. On Sunday I had written off blocking them as pointless, since a blocked attacker simply rents another address. Now blocking had a purpose: we blocked their whole hosting network from our website, so the robot could no longer open the page or collect new proofs, and we reported the server to its provider.
Tuesday, 04:25: the last try
They came back before dawn from two directions: their original server and a new one behind a commercial VPN. All 168 requests for the data they had mapped on Saturday were refused outright, and none of their 25 password guesses got past the robot check. It could no longer open our page, so it had no proof to show.
So, is software built with AI safe?
That is the wrong question. The popular database service behind this app has for years left newly created data open to anonymous visitors by default, unless the builder switches on those record-by-record checks. It is changing that default this year, but everything built before keeps the open setting.
So the well-known stories of AI-built apps leaking their whole database are rarely about bad code. They are about a door nobody told the builder to lock. An AI will build you a beautiful app with an open back door if you never mention back doors, and so would a junior developer, or a senior one on a deadline. What kept this app safe were rules written down before the first line, and checked against what actually happened:
- Check every record against who is asking, from day one.
- Don’t treat the screens as security. People will go straight to your database.
- Give strangers nothing, not even an empty room.
- Don’t store what can be stolen. No passwords means nothing to guess.
- Test by breaking. Sign in as every kind of user and try the forbidden thing.
- Never rely on one lock. Our robot check was beaten, and it didn’t matter.
- Read the logs before calling something pointless. The block I had dismissed is what stopped the guessing.
AI made this app faster to build and the incident faster to fix: hours, not weeks. But the safety came from knowing what to ask for, and software built with AI is exactly as safe as the person who knows what to ask it.
If you are building with AI, or had something built and aren’t sure what’s behind the front door, we’re happy to take a look. Read about our services here.
This article was published on the Andronia blog. Andronia helps Hungarian businesses grow with AI automation. Read about our services here.