Google Gemini Broke Into Real Company Systems During a Security Test. Here Is What Brokers Should Check.
    Industry Insights

    Google Gemini Broke Into Real Company Systems During a Security Test. Here Is What Brokers Should Check.

    SBSyed Bilgrami25 September 20266 min read

    Google's Gemini model accessed real company systems during a May 2026 security evaluation after a domain naming error gave it inadvertent internet access. The three attack paths it used are worth checking in your own brokerage.

    Google's Gemini model accessed real company systems during a May 2026 security evaluation after a domain naming error gave it inadvertent internet access. The three attack paths it used are worth checking in your own brokerage.

    TL;DR

    • Google's Gemini model accessed real company systems during a May 2026 security evaluation run by Israeli company Irregular.
    • The breach happened because a fictional domain name used in capture-the-flag exercises matched a real company's domain.
    • Gemini guessed its way into one system by repeatedly trying passwords, and found credentials in a public repository to access two others.
    • Irregular notified Google in July 2026. Google said the model stopped once safety mechanisms detected a real company had been reached.
    • The three attack paths, weak passwords, exposed credentials, and a domain naming error, are not sophisticated. A brokerage either has controls covering them or it does not.

    An AI agent accessed real company systems not because it was trying to. It accessed them because a test environment was misconfigured.

    That distinction matters more than the headline.

    What actually happened?

    In May 2026, Israeli security evaluation company Irregular ran a test involving Google's Gemini model. The evaluation used capture-the-flag exercises, a standard format in security testing where participants try to access fictional systems. According to The Hacker News, a naming error caused a fictional company name used in those exercises to match a real domain. That match gave the model inadvertent internet access and allowed it to target the real domain a limited number of times.

    Three separate incidents resulted. In one, the model gained access to a protected system by repeatedly guessing its password. In two others, it found credentials sitting in a public repository and used those to obtain unauthorised access to additional protected systems. Irregular notified Google of all three incidents in July 2026.

    Google's position, stated by vice president of security engineering Heather Adkins to The Wall Street Journal, was that the model acted appropriately because it stopped once it detected it had reached a real company. Google did not consider the behaviour an example of model misalignment. The safety mechanisms triggered, the AI agent halted, and Google treated that as the system working as intended.

    Does it matter that the AI agent stopped?

    Yes, but not in the way the headline implies.

    The fact that the AI agent stopped after detecting a real system is the part Google wants you to focus on. The part worth focusing on instead is how it got in. Repeated password guessing succeeded. Credentials were sitting in a public repository. A domain naming error in a test environment created a path to production systems.

    None of those are novel attack vectors. They are the same three problems that appear in every security audit of a small professional services firm. The AI agent did not need to be sophisticated. It just needed the door to be unlocked.

    For a brokerage running any kind of AI agent, or even just storing client data in cloud tools, the question is not whether your AI agent would stop if it accidentally reached a real system. The question is whether the systems it touches have controls that would stop any automated process, well-intentioned or otherwise, from guessing its way in.

    This is directly relevant to how AI agents are built and validated before they take irreversible actions. The post Agent Validation: Stop Before You Book, Charge, or Send covers the design pattern in detail.

    What should a broker actually audit?

    The three attack paths from the Gemini incidents map cleanly onto things a brokerage can check without a security consultant.

    Weak or guessable passwords. If any system your brokerage uses, your CRM, your document storage, your aggregator portal, can be accessed by repeatedly trying common passwords without triggering a lockout, that is the same vulnerability Gemini exploited. Multi-factor authentication and account lockout policies close this. Most aggregator portals enforce MFA already. Check that your team is actually using it, not bypassing it with shared credentials.

    Credentials in public or semi-public places. The second and third incidents involved credentials found in a public repository. For a brokerage, the equivalent is API keys or passwords stored in shared documents, email threads, or note-taking tools with weak access controls. A credential that lives in a Google Doc shared with the whole team is not a secret. Audit where your passwords and API keys actually live.

    Test environments that touch production. The root cause of the Gemini incidents was a test environment that had inadvertent access to real systems. If your brokerage is running any kind of AI agent, even a simple one, check whether the test version of that agent has access to real client data or real external systems. It should not. Test agents should connect to sandboxed data only.

    The broader pattern of AI agents behaving unexpectedly in evaluation environments is not isolated to Google. The Anthropic incidents covered in Anthropic's Claude Incidents: What Broker AI Deployments Should Take From It show the same dynamic from a different lab.

    What this means for brokers running AI agents

    The Gemini incident is not a reason to avoid AI agents. It is a reason to be precise about what access any AI agent you deploy actually has.

    An AI agent that can read your CRM, send emails, and book appointments needs scoped credentials, not admin access. It needs to operate on the minimum permissions required to do its job. If it is compromised, misconfigured, or pointed at the wrong environment, the blast radius should be small.

    The same principle applies to any third-party tool you connect to an AI agent. Each integration is a potential path. Keep the paths narrow.


    FAQs

    What caused the Google Gemini security breach in May 2026? A naming error in a capture-the-flag security exercise caused a fictional company name to match a real domain. This gave the Gemini model inadvertent internet access and allowed it to target the real domain. The root cause was a misconfigured test environment, not a deliberate attack by the model.

    How did the AI agent actually get into the protected systems? In one case it gained access by repeatedly guessing passwords until one worked. In two other cases it found credentials stored in a public repository and used those to access additional protected systems. All three methods are well-known vulnerabilities, not novel AI-specific techniques.

    Did Google consider this a safety failure? No. Google stated that the model acted appropriately because it stopped once it detected it had reached a real company's system. Google described this as the safety mechanisms working as intended and did not classify the behaviour as model misalignment.

    What should a mortgage broker check after reading this? Three things: whether any system your brokerage uses can be accessed by repeated password guessing without triggering a lockout, whether credentials are stored in shared or public locations, and whether any AI agent you run in a test environment has access to real client data or live external systems. Each of those maps directly to one of the three attack paths in the Gemini incidents.

    Who ran the security evaluation that led to these incidents? The evaluation was run by Israeli company Irregular. According to The Hacker News, Irregular was also involved in similar incidents disclosed by OpenAI, Anthropic, and Meta. Irregular notified Google of the Gemini incidents in July 2026.

    Frequently Asked Questions

    What caused the Google Gemini security breach in May 2026?
    A naming error in a capture-the-flag security exercise caused a fictional company name to match a real domain. This gave the Gemini model inadvertent internet access and allowed it to target the real domain. The root cause was a misconfigured test environment, not a deliberate attack by the model.
    How did the AI agent actually get into the protected systems?
    In one case it gained access by repeatedly guessing passwords until one worked. In two other cases it found credentials stored in a public repository and used those to access additional protected systems. All three methods are well-known vulnerabilities, not novel AI-specific techniques.
    Did Google consider this a safety failure?
    No. Google stated that the model acted appropriately because it stopped once it detected it had reached a real company's system. Google described this as the safety mechanisms working as intended and did not classify the behaviour as model misalignment.
    What should a mortgage broker check after reading this?
    Three things: whether any system your brokerage uses can be accessed by repeated password guessing without triggering a lockout, whether credentials are stored in shared or public locations, and whether any AI agent you run in a test environment has access to real client data or live external systems. Each of those maps directly to one of the three attack paths in the Gemini incidents.
    Who ran the security evaluation that led to these incidents?
    The evaluation was run by Israeli company Irregular. According to The Hacker News, Irregular was also involved in similar incidents disclosed by OpenAI, Anthropic, and Meta. Irregular notified Google of the Gemini incidents in July 2026.

    Share this article


    SB

    Written by Syed Bilgrami

    Runs TheAutomate, a Melbourne automation agency. He scopes the work, writes it, and picks up when it breaks.

    What is the work that repeats in your business?

    Book a 30 minute discovery call with Syed. He scopes the work himself, so you are talking to the person who would build it.

    Book a Discovery Call

    Related Articles