← All articles

Can AI hack a server or a website? What it can already do and where it stumbles

Topics: AI, Security, Small business

Server racks in the dark and a glowing amber padlock with lines of data reaching towards it

The short answer: AI already helps with individual steps of an attack. At the same time, in the 2024 Cybench study, agents without hints solved only short tasks from hacking competitions and did not solve the long multi-step ones. Defence still relies on the same measures as without AI: updates, passwords, rules for checking payments. Below we look at what exactly the models can do, who is at risk and where to start.

What AI really does in attacks

Reconnaissance. Before an attack, people gather information about a company: employees, email addresses, which software it uses, which services are exposed. A language model helps sort through what was found and suggests whom to target and how.

Phishing and deception. A model writes a polished email in many languages and can add details from the victim's public social media pages. That is why well-written text without mistakes does not, on its own, prove that an email is genuine.

Faked voice and video. In early 2024, an employee at the Hong Kong office of the engineering company Arup transferred about 25 million dollars to fraudsters after a video call in which the "chief financial officer" and "colleagues" were fakes. The same technique works for the "relative in trouble" call: a familiar voice can be imitated by software.

Finding vulnerabilities. Models can read code and find errors in it. In 2024, Google's AI agent Big Sleep found a previously unknown dangerous bug in a very widely used program, the SQLite database; Google itself called this the first public example of its kind. The bug was fixed before the release. In the Cybench study (2024), agents built on the models of that time, without hints, solved tasks from hacking competitions (CTF) that took human teams up to 11 minutes. Without hints, they did not solve longer tasks, including those that took people almost a full day. This is a snapshot at the time of the study, not a final limit.

Automating routine work. Agents can run scanners, read their output, choose the next step and write a report.

Malicious code. Major model developers forbid their models to help with malware, but these restrictions are not absolute, and open models can be run without them.

Where AI still stumbles

  • Long multi-step tasks. In the same Cybench study, agents without hints did not solve the longest tasks: the more steps, the more chances to make a mistake along the way.
  • Well-configured protection. Sign-in with a second confirmation step, up-to-date software and splitting the network into zones make an attack harder, and sign-in logs help to notice it in time. This reduces risk; it is not a guarantee: a successful attack remains possible even with good protection.

The practical conclusion for a small company: the most obvious scenarios listed above are deceiving employees by email, phone call or video, so that is where protection should start.

Who is at risk

Who Typical scenario What increases the risk
Small business Phishing the accountant, swapped bank details There may be no security specialist and no rules for checking payments
Ordinary people A call imitating the voice of someone close, a "safe account" Trust in a familiar voice, panic
Companies with their own website Searching for forgotten admin panels and old plugins Rare updates, forgotten services
Developers Keys and passwords in public code A leaked key can be found within minutes

The last row is no exaggeration: in researchers' experiments, a cloud key published on GitHub started being used one to two minutes after it was published.

The other side: AI helps defenders too

The same abilities work on the defence side: models help analyse event logs, look for unusual behaviour, check code before release and fix the bugs found. At the final of the AI Cyber Challenge run by the US agency DARPA in August 2025, autonomous AI systems found 54 of the 63 vulnerabilities planted by the organisers and fixed 43, and some of these systems were released publicly for defenders.

What to do in practice

For an ordinary person:

  1. Agree on a code word with the people close to you for "urgent" calls. Voice and video can be faked, so on their own they are not proof.
  2. Check any request for money by calling back a number you already knew.
  3. Turn on sign-in with a second confirmation step, preferably through a code generator app rather than SMS.
  4. Get a password manager and use a unique password for every service.

For a website owner or a small company. You do not have to do the technical items yourself: you can hand them to whoever looks after your website and computers, an in-house specialist or a contractor. Your job is to set the task and get confirmation that it has been done.

  1. Updates. The system your website runs on and its add-ons (plugins) are updated regularly. Among the data breaches analysed in Verizon's 2025 report, roughly one in five began with an attacker exploiting a vulnerability, that is, a flaw in software, and this share grew over the year. Updates close flaws of this kind that are already known. Confirmation: a list of what is installed on the website, with versions and the result of checking whether security updates are available for them.
  2. Nothing unnecessary exposed. Website admin panels, databases and test copies of the website should not be open to anyone on the internet. Confirmation: a list of what is reachable from outside and an explanation of why each item is open.
  3. No passwords or keys in program code. A key here is a secret code that a program uses to access a service, for example a cloud or a payment system. A key that has already become public must be revoked and replaced: removing it from the code is not enough. Confirmation: the result of the latest run of an automated check for keys in your code, with the date and what was checked, and, if anything was found, confirmation that those keys have been revoked and replaced.
  4. The two-person rule for money. Any payment to new bank details, or any change of bank details, is prepared by one person and confirmed by a second through a separate channel, for example by a call to a known number.
  5. Backups. Kept separately from the main network. Confirmation: the date of the latest trial restore from a backup in a separate test environment, and that restore is recent.
  6. Sign-in alerts. A sign-in to email, cloud services or the website panel from an unusual place sends a notification to the person responsible. Confirmation: a test sign-in from a new device, after which the notification arrived.

If there is no contractor or the settings are out of your reach, start with what can be done without technology: the two-person rule, a code word, and sign-in with a second confirmation step for email and cloud services.

For a developer: use AI as a second reviewer of your code, but do not trust its output blindly: it can both find a bug and write insecure code itself.

How to check the two-person rule. Walk through a made-up situation without money and without the payment system: at a meeting or in a chat, describe an "urgent request from the director" to transfer money to made-up bank details, and ask the person who prepares payments and the person who confirms them to explain step by step what they would do. See whether the rule comes through in their answers: preparation by one person, confirmation by a second through a separate channel. This check covers payments only; the other items are checked with the confirmations from the list above.

The legal side

Before testing any system other than your own for vulnerabilities, get written permission from its owner with agreed limits for the test. Without it you are taking a risk: in Russia, unlawful access to computer information can lead to liability under Article 272 of the Russian Criminal Code, and it is not you who decides where the line of what is acceptable lies. Legal options: your own test environments, dedicated platforms with hacking challenges (CTF), bug bounty programmes with clear rules, and audits under a contract.

Summary

AI has not made hacking magical; it helps with individual steps of an attack, from reconnaissance to a convincing email. In the cases discussed above, the deciding factor was trust or an overlooked detail: people believed a video call, a key was left in public code. That is why it makes sense to start defence with the simple measures from the list above and with rules for checking payments, and only then decide whether your company needs something more complex. Even if an email or a voice seems genuine, the rule "money only after checking through a second channel" works regardless.

If you would like to see how I approach automation and AI in my projects, or to discuss your own task, take a look at my portfolio: there you will find my projects and a way to contact me directly.

Sources