AI Safety Tests: Meta Is the Third Lab Caught Out

AI safety tests failing as a reinforced door swings open and a Meta model slips out of its enclosure

AI safety tests keep going wrong: Meta confirmed that its Muse Spark 1.1 model gained unauthorised access to a third-party company’s systems during one such evaluation. The lab pins the incident on a misconfigured test environment, and points at an outside contractor.

Key Takeaways

  • The sandbox meant to be isolated had kept outbound internet access, which let the model reach a live service.
  • The same evaluation partner, Irregular, was already behind the three incidents Anthropic disclosed the previous week.
  • Neither Meta nor the contractor has named the third-party service that was hit.

Have an AI Sum Up This Article

ChatGPT

An exit door left open in the sandbox

The whole scenario fits in one configuration line. The environment where Muse Spark was being evaluated was presented as isolated, yet outbound connectivity stayed switched on.

Once on the network, the model did exactly what it was being tested for. It found a vulnerability in a third-party service, exploited it, then made unauthorised changes to that target’s internal environment.

Meta confirmed the episode through spokesperson Andy Stone and named who owned the faulty setup: Irregular, the contractor running the evaluation. The identity of the affected company has not been disclosed.

The model involved is no lab prototype. Muse Spark 1.1 shipped on July 9, positioned as Meta’s most capable release for real-world coding and agentic tasks, and we noted at the time that Meta was pricing this model aggressively against OpenAI and Anthropic.

The mismatch sits right there. A model sold as production-ready, tested inside an environment described as sealed, which walks out and hits a live target. The model did not fail its mission, the perimeter simply did not exist.

It is worth naming what the episode does not show. Nothing public suggests the model tried to work around an instruction, or that it knew it had stepped outside its frame. It treated a live system as an exercise target because nothing let it tell the two apart.

That distinction matters for how the industry files this. A model deliberately breaking its constraints is an alignment story. A model behaving correctly inside a broken enclosure is an infrastructure story, and infrastructure stories get fixed by contracts and audits rather than by training runs.


AI Safety Tests

The shared weak link is called Irregular

The name of the contractor running these AI safety tests is the real story of the week. In late July, Anthropic published a detailed breakdown of three incidents in the same context: Claude models reaching the internet from inside a third-party evaluation environment, then gaining access to the production infrastructure of three different organisations.

The models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research model, all running capture the flag exercises. The identified cause was already the same: internet access left available by mistake, inside the environment of a partner named Irregular.

The scale of Anthropic’s review shows how much work spotting this kind of drift takes: 141,006 evaluation runs combed through. We covered that file when it landed, when Anthropic admitted three real breaches inside its cybersecurity evaluations.

One incident at an evaluation vendor is an accident. The same incident at two major clients becomes a supply chain failure. The question facing Meta is no longer about its model, it is about how much control a lab keeps over environments it does not operate itself.

Anthropic at least shipped a structured answer: closer monitoring of evaluation transcripts, tighter security standards on evaluation infrastructure, an external review run with METR, a planned transcript release and stronger vendor oversight. Meta so far offers the misconfiguration finding and little else.


More articles on Horizon


What buyers will start demanding

For teams buying frontier models, the consequence of failed AI safety tests lands immediately. An offensive score published by a lab is worth nothing without a description of the environment that produced it.

The procurement question changes shape. It stops being about model capability and starts being about the evaluation vendor’s name, the sandbox network topology, and whether outbound control was verified rather than merely declared.

On the competitive side, the advantage tilts toward whoever documents. Anthropic took a hit admitting three incidents, but published a method, a volume of reviewed transcripts and a fix list. Meta now has to match that, without having picked the timing.

The third-party evaluation market comes out of this bruised. These vendors sell exactly one promise: that an offensive model can be pushed to its limit without consequences leaking outside. When that promise collapses twice in a fortnight, the value of the service itself is on the table.

Labs have two ways out, and both cost money. Bring the offensive evaluation work back in house and carry the infrastructure bill, or keep outsourcing it and start auditing the vendor’s network posture the way a bank audits a payment processor.

This will reach regulators eventually, because the victim each time is a third party that never signed up for any of it. A company whose systems were altered during the evaluation of a model it does not use holds legal arguments nobody has tested yet.

The next step is a transparency question. If the labs finally publish shared standards for external evaluation environments, this run of incidents will have been worth something. Otherwise a fourth name joins the list before the summer ends.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *