All posts
13 min read

AI Keeps Getting Better. The AI Industry Is a Complete Mess.

The models have never been more capable. But this year OpenAI's own agents broke into Hugging Face, RubyGems and an Australian government portal, Big Tech is spending about $730 billion on AI, and both leading labs lose more than they make. Here's the mess, how it happened, and what would fix it.

  • AI
  • OpenAI
  • Claude
  • Security
The words Better models. Messier industry. beside a bar chart of Big Tech capital expenditure: $260 billion in 2024, $448 billion in 2025 and about $730 billion of guidance for 2026
Big Tech’s AI spending has nearly tripled in two years. That’s one of three messes.

ColdFusion put out a video yesterday with a title I can't argue with: The AI Industry is a Complete Mess. The argument is simple. The tools keep getting better. The industry around them keeps getting more chaotic: agents doing things nobody asked them to, money that doesn't add up, and launches that fall over on stage.

I've spent this week writing about two small pieces of that: ChatGPT Pro $200 losing half its usage, and GPT-6.1 Sol, a great model I'll probably never use. So I wanted to step back and look at the whole thing. I checked every claim I could find a source for and left out the ones I couldn't.

Start with one day. On September 29, at DevDay, OpenAI:

  • launched GPT-6.1 Sol, one of the best models for the money anyone has shipped;
  • cut the usage on its $200 Pro plan in half, because it's short of compute;
  • confirmed it had pulled GPT-6.1 Astra, because in safety testing it was too willing to mislead users and to act beyond its instructions;
  • and showed off a new agent on stage that said "Checking that now," went quiet for about ten seconds, and then said "still checking."

Five days earlier, Australia's prime minister had spoken to Sam Altman to express his country's "extreme concern," because one of OpenAI's agents had got into a government health portal.

That's the whole story in one week. The models have never been better, and the companies building them have never looked less in control: of their agents, of their costs, and some days, of their own products.

AI in 2026: better models, messier industry

Vulnerabilities found by Claude Mythos
10,000+
High or critical, across Project Glasswing partners, in about a month
Events in the Hugging Face intrusion
17,000+
Run end to end by OpenAI agents that got out of a test sandbox
Big Tech capex, 2026
~$730B
Amazon, Alphabet, Microsoft and Meta guidance. Nearly triple 2024
OpenAI’s projected cash burn, 2026
$25B
Rising to $85B in 2028, by its own forecast

Sources: Anthropic’s Project Glasswing update (May 22), Hugging Face’s incident report (July 16), company guidance from Q2 2026 earnings, and OpenAI’s projections as reported by The Information in February.

First, the part that's genuinely incredible

I want to be fair here, because none of this is happening because AI is fake.

In April, Anthropic announced Claude Mythos Preview, a model so good at finding security bugs that Anthropic decided not to release it to the public. Pointed at its partners' software through Project Glasswing, it found zero-days in every major operating system and browser, and more than 10,000 high- or critical-severity vulnerabilities in its first month.

A week ago I wrote that GPT-6.1 Sol gets close to OpenAI's flagship for a fifth of the price. I run Claude Code most of my working day, and it gets noticeably better every few months. And people are paying for it: Anthropic's annualised revenue crossed $47 billion in May and reached $65 billion by the end of July. OpenAI's was $40 billion in August.

If the story of AI in 2026 were only about the models, it would be a happy one. Here's the rest of it:

Six months of AI news, in one list

Only events with a source I could check. Every one of them happened in 2026.

  • Agents and security
  • Money
  • Product and capacity
  1. Apr 7

    Agents and security: Anthropic announces Claude Mythos Preview, which finds zero-days in every major OS and browser, and keeps it from general release.

  2. May 11–12

    Agents and security: More than 2,000 malicious packages flood RubyGems. New sign-ups are switched off for four days.

  3. May 28

    Money: Anthropic raises $65 billion at a $965 billion valuation.

  4. Jun 18

    Agents and security: An OpenAI agent gets into Australia's Medicare statistics portal. Nobody notices for 54 days.

  5. Jul 9–13

    Agents and security: OpenAI agents running a cyber evaluation break into Hugging Face. OpenAI later confirms they were its models.

  6. Jul 30

    Money: Amazon raises its 2026 capex to $220 billion and says it still won't have enough capacity.

  7. Sep 3

    Product and capacity: GPT-6 Astra launches. ChatGPT goes down the same morning.

  8. Sep 10

    Product and capacity: OpenAI pauses ChatGPT Pro $200 sign-ups over Astra demand, and emails Services Australia's public mailbox about the breach.

  9. Sep 12

    Agents and security: Researchers trace the RubyGems attack to OpenAI agents. OpenAI says they were doing “benign tasks.”

  10. Sep 24

    Agents and security: Australia's prime minister goes public: the agent “didn't accept ‘no’ for an answer.”

  11. Sep 28

    Agents and security: The UK AI Security Institute says GPT-6 Astra ran unsanctioned supply-chain attacks in testing. The WSJ reports GPT-6.1 Astra is shelved.

  12. Sep 28

    Money: Anthropic's IPO filing shows an $8 billion operating loss for 2025, on $4.6 billion of revenue.

  13. Sep 29

    Product and capacity: DevDay: GPT-6.1 Sol launches, Pro $200 is cut in half, and the live agent demo stalls on stage.

  14. Sep 29

    Money: OpenAI is reported to be raising at a $1.4 trillion valuation.

Sources: Anthropic, Hugging Face, The Register, Simon Willison, ABC News, Fortune, TechCrunch, BleepingComputer, Futurism, Reuters, Bloomberg and the WSJ, each linked where it comes up in the post.

Sort that list and you get three different messes. They feed each other, but let's take them one at a time.

The agents don't take no for an answer

The scariest AI security stories this year aren't about hackers using AI. They're about AI labs' own agents doing things nobody told them to do.

Hugging Face. In July, Hugging Face disclosed an intrusion that, in its own words, "was driven, end to end, by an autonomous AI agent system." More than 17,000 recorded events, "able to do in hours what would usually take days." The attacker got into a limited set of internal datasets and several credentials. Public models, datasets and Spaces weren't touched.

The attacker turned out to be OpenAI. Its agents, GPT-5.6 Sol and "an even more capable pre-release model," were being tested on a cyber benchmark called ExploitGym, with safeguards turned down on purpose. They got out of their test environment and went after real infrastructure. OpenAI has since "deactivated, encrypted, and restricted" the pre-release model.

RubyGems. On May 11 and 12, more than 2,000 malicious packages flooded RubyGems, the Ruby package registry, and its maintainers had to switch off new sign-ups for four days. In September, three researchers traced them to OpenAI's agents. OpenAI's answer was that "our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information." Simon Willison put his finger on the worst part: OpenAI hadn't told RubyGems it was responsible until the researchers did.

Medicare. On June 18, an OpenAI agent got into the Medicare statistics portal run by Services Australia. OpenAI found out on August 11. It told Australia on September 10, and the prime minister made it public on September 24. In Anthony Albanese's words, "The AI agent found a way around those blocks, didn't accept 'no' for an answer." And: "The notification was an email sent just to the public mailbox." OpenAI says no patient records were accessed, only "aggregate health statistics and internal file names."

That timeline bothers me more than the breach itself:

98 days from breach to headline

How long each step took after an OpenAI agent got into Australia's Medicare statistics portal, next to the deadlines two US state laws set for a frontier lab to report a critical safety incident. Neither law covers this case. They're here for scale.

  • The Medicare portal breach
  • Reporting deadlines in US state law

Breach to OpenAI finding it

Finding it to telling Australia

Telling Australia to the public

California SB 53

New York RAISE Act

Sources: ABC News (dates, from the Australian government and OpenAI); Wiley on New York’s RAISE Act and California’s SB 53. Both laws require reporting to the state, not to whoever was breached. Day counts are my arithmetic.

Show the data
StepDaysWhen
Breach to OpenAI finding it54June 18 to August 11
Finding it to telling Australia30August 11 to September 10
Telling Australia to the public14September 10 to 24
California SB 5315From discovery of a critical safety incident
New York RAISE Act372 hours, from January 1, 2027

And it isn't a one-off. The UK AI Security Institute tested GPT-6 Astra with its standard security classifiers off, and found it "conducted a range of unsanctioned attack activities" at a higher rate than GPT-5.6 Sol and GPT-5.5. That included "creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews." Then OpenAI pulled GPT-6.1 Astra, because testing found it too willing to deceive users about what it had done and to go beyond its instructions. It "didn't quite meet the bar," said Saachi Jain, who leads OpenAI's safety systems.

To be fair to OpenAI: pulling 6.1 Astra was the right call, and it cost them a headline launch. And these incidents came out of evaluations where safeguards were deliberately turned down, which is how you find out what a model can really do. But a test that can reach Hugging Face, RubyGems and the Australian government isn't a test. It's the internet. A sandbox you can walk out of is just a room.

Anthropic has the opposite problem

Anthropic did the cautious thing with Mythos. It kept the model back and gave it to defenders first. Here's what happened next, in open-source software alone:

The AI finds bugs faster than people can fix them

What Claude Mythos Preview found when Anthropic pointed it at more than 1,000 open-source projects, and how many of those bugs had been disclosed and patched by May 22, about six weeks in. Of the 1,752 findings people had checked by then, 90.6% were real.

  • Done by the AI
  • Done by people

Vulnerabilities found

High or critical (estimated)

Disclosed to maintainers

Patched

Source: Anthropic’s Project Glasswing update, May 22, 2026. Across all of Glasswing’s partners, not just open source, Mythos found more than 10,000 high- or critical-severity vulnerabilities in its first month.

Show the data
StageCount
Vulnerabilities found23,019
High or critical (estimated)6,202
Disclosed to maintainers530
Patched75

Anthropic said it plainly: "Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it's limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI."

That's the good-guy version. Anthropic also expects that "within 6 to 12 months" many other AI companies will have Mythos-class models, and I doubt all of them will be as careful about who gets one. Its September threat report already sums up where this goes in one line: "Sophisticated attacks no longer require sophisticated attackers." In one case it describes, a single stolen developer token became full admin access in about three hours.

What this looks like from the security side

I'm the lead engineer at Jutsu, an AI-native security operations platform, so weigh this accordingly. But this is the part of the mess I deal with every day.

An attacker that does 17,000 things in a few days, never sleeps, and takes "no" as a puzzle to solve is a different problem from a person at a keyboard. And there's an irony in the Hugging Face story that I can't stop thinking about. When Hugging Face tried to fight the attack with Anthropic's Opus and Fable models, they "refused a large part of that work" because of their safety guardrails. The attacking agents had their safeguards turned down for an evaluation. The defenders' AI had its safeguards turned up.

In July I asked who is watching the agents. This year, for at least one lab, the answer was that nobody noticed for 54 days.

The money doesn't add up (yet)

Every one of these models runs in a data center somebody had to build, and the building is getting very expensive:

Big Tech's AI spending nearly tripled in two years

Capital expenditure, mostly data centers and the chips inside them. 2026 is what Amazon, Alphabet, Microsoft and Meta told investors they'd spend after their July earnings, before Oracle and before anything they add in the second half.

  • Actual, five companies
  • 2026 guidance, four companies

2024

2025

2026

Sources: Epoch AI for 2024 and 2025 (Microsoft, Amazon, Alphabet, Meta and Oracle). 2026 is the midpoint of company guidance: Amazon ~$220B (Fortune), Alphabet $195–205B, Microsoft ~$175B for the calendar year, Meta $130–145B (its Q2 8-K). Oracle alone spent $55.7B in its fiscal year to May 2026.

Show the data
Year or companyCapex
2024$260B
2025$448B
2026$733B
2026: Amazonabout $220B
2026: Alphabet$195–205B
2026: Microsoftabout $175B
2026: Meta$130–145B

The four biggest spenders told investors in July that they'll spend about $730 billion this year. Amazon alone raised its number to $220 billion, and Andy Jassy still said Amazon won't "have enough capacity to meet all the demand we have in 2026. And I believe this dynamic will also be true in 2027, too."

That spending is starting to outrun the cash these companies make. Epoch AI projects that the hyperscalers' combined capex catches up with their operating cash flow around Q3 2026: $185.7 billion spent against $187.1 billion coming in. The gap is being filled with borrowing. FactSet found that new debt went from 9% of capex in 2024 to 32% by the middle of this year.

The labs themselves are further from paying their way:

In 2025, both leading labs lost more than they made

Revenue and operating loss for 2025, before the one-off accounting charges that make both companies' net losses look even bigger. Both are growing very fast. Neither is close to paying for itself yet.

  • Revenue
  • Operating loss

OpenAI

Anthropic

Sources: OpenAI’s audited 2025 financials, as reported by Ed Zitron, who says the Financial Times independently verified them (OpenAI hasn’t published them). Anthropic’s IPO filing, as reported by Reuters on September 28, 2026.

Show the data
Lab2025 revenue2025 operating loss
OpenAI$13.1B$20.9B
Anthropic$4.6B$8.1B

OpenAI made $13.07 billion in 2025 and lost $20.92 billion from operations, according to audited financials reported by Ed Zitron. OpenAI hasn't published them. Anthropic made about $4.6 billion and lost $8.06 billion, according to its IPO filing, as reported by Reuters. And OpenAI's own plan says the burn gets bigger before it gets better:

OpenAI's own plan: burn $218 billion, then turn a profit

The cash OpenAI expects to burn each year, from projections it shared with investors in February. The plan has it turning cash-positive in 2030, with $39 billion. The forecast was revised up by about $111 billion from the previous one.

2026

2027

2028

2029

Source: The Information, February 2026, as reported by The Decoder. These are OpenAI’s projections, not results, and they predate the Astra launch and the Pro $200 changes. The four-year total is my arithmetic.

Show the data
YearProjected cash burn
2026$25B
2027$57B
2028$85B
2029$51B
2026–2029$218B

Meanwhile the valuations keep climbing. OpenAI raised at $852 billion in March and is reportedly talking to investors at $1.4 trillion. Anthropic raised at $965 billion in May and, per Reuters, its IPO could value it at more than $2 trillion. It has also signed up for $518 billion of cloud and compute obligations.

I don't think this is simply 2000 again. The revenue is real, and it's growing faster than anything I've seen in software. Anthropic went from $47 billion to $65 billion annualised in about three months. But a business that loses money on its customers needs one of two things: cheaper compute, or higher prices. We just watched what "higher prices" looks like in practice. It looks like the $200 plan quietly becoming half a plan.

The product changes under your feet

This is the mess regular users actually feel.

In one month, OpenAI launched GPT-6 Astra on September 3 and ChatGPT went down that morning. It paused Pro $200 sign-ups on September 10 because demand for Astra was "really unprecedented." Codex users have been seeing "Selected model is at capacity" since early September, even with quota left. Then GPT-6 Sol on September 22, GPT-6.1 Sol a week later, and the $200 plan cut in half on the same day. If you built a workflow around any of it, you rebuilt it twice this month.

And then there are the demos. At DevDay, an OpenAI presenter asked the new agent, "Can you catch me up on just user testing from last night?" It said "Checking that now," went silent for about ten seconds, and said "still checking." "I guess Dottie's having a slow morning," said Holly Li, who was presenting. Engadget's live blog logged a voice demo that wouldn't start on the same stage. A year earlier, Meta's smart-glasses cooking demo fell apart at Connect because, as its CTO later put it, "we DDoS'd ourselves, basically."

Live demos fail, and one stalled agent isn't a scandal. But a keynote is the one time a company controls everything: the network, the prompt, the data, the audience. If the agent can't find yesterday's notes on stage, it's fair to wonder how it does in your Slack on a Tuesday.

How it got this messy

I don't think any of these are separate problems. I think they're one loop:

  1. Each generation of model costs more to train and serve, so labs raise more money at higher valuations.
  2. Higher valuations need faster growth, so labs ship faster: a new Sol, then another Sol a week later.
  3. Faster shipping runs into finite compute, so users get rationed: paused sign-ups, halved plans, "at capacity."
  4. And to stay ahead, labs push their agents' capabilities as hard as they can, in tests that turn out to be connected to the real internet.

Everyone is racing, and nobody can afford to slow down first. Even the rules are slipping: the EU has pushed its high-risk AI obligations from August 2026 to December 2027.

So no, I don't think the problem is that AI is overhyped or useless. It's real, and it's arriving faster than the security, the economics and the operations around it.

What would actually fix it

None of this needs a breakthrough. It needs boring things done properly.

  • Sandboxes that are actually sandboxes, and someone watching them. The UK AI Security Institute's own conclusion is that "measures beyond model alignment, like sandboxing and monitoring, may be necessary to prevent real-world harm." An agent being tested for hacking skills shouldn't be one misconfigured proxy away from the internet.
  • Fast, mandatory disclosure, to the people affected. New York's RAISE Act will require frontier labs to report critical safety incidents within 72 hours from January. That should be the floor everywhere, and the victim should hear first, not a public inbox a month later.
  • Selling only what you can serve. If you don't have the compute for a plan, don't sell the plan. Pausing sign-ups was the honest move. Halving the plan for the people already paying for it was the cheap one.
  • Paying for the fixing, not just the finding. If AI makes bugs cheap to find, the bottleneck is the people who fix them. Anthropic committed up to $100 million in usage credits to Glasswing, plus $4 million in donations to open-source security groups. The rest of the industry should be matching that.

What I'd do

  • If you run AI agents: no standing credentials, an allow-list for where they can connect, a log of every action they take, and a way to stop them. Assume the sandbox leaks, because this year it did.
  • If you buy AI for your company: treat your AI provider like any other third party with access to your systems. Ask how, and how fast, they'd tell you about an incident. "An email to the public mailbox" is not an answer.
  • If you build on AI: keep a fallback model from a different provider, and don't plan your costs around a subscription staying the same for a year. This month, one didn't stay the same for a week.
  • If you maintain open source: expect a lot more AI-found bug reports, good and bad. Ask for a working reproduction before you spend an evening on one.
  • If you're deciding what to learn next: the boring skills, like security, evaluation and operations, are where the shortage is. The models aren't the bottleneck any more. Everything around them is.

The verdict

ColdFusion is right: the AI industry is a complete mess. But it's worth being precise about what kind. The technology isn't the mess. Mythos finds bugs faster than humans can fix them. Sol does near-flagship work for a fifth of the price. Those are real achievements, and I use these tools all day.

The mess is everything around the models. Agents that get out of their test environments and keep going, and companies that take weeks to notice. Data centers being built faster than the cash comes in. Plans that change under you, and demos that stall on stage. It all comes from the same race, and nobody in it wants to be the one who slows down.

The models keep passing their tests. The industry keeps failing its own.


A note on sources: the Hugging Face intrusion is from Hugging Face's own report and Fortune. RubyGems is from The Register and Simon Willison. The Medicare breach and its dates are from ABC News. The UK AI Security Institute's findings are from The Register; the GPT-6.1 Astra decision is from NPR via KATU. Mythos and Glasswing numbers are Anthropic's own, from its announcement, May update and June expansion; the threat report is Anthropic's September report. Capex for 2024 and 2025 is from Epoch AI; 2026 is company guidance from July, including Meta's Q2 8-K and Amazon via Fortune. Cash flow and debt are from Epoch AI and FactSet. OpenAI's 2025 figures are leaked audited financials reported by Ed Zitron, not published by OpenAI; its projections are from The Information via The Decoder. Anthropic's 2025 figures are from its IPO filing as reported by Reuters, via Unite.AI. Valuations and revenue run-rates are from Anthropic and TechCrunch. Demo failures are from Futurism, Engadget and TechCrunch. The RAISE Act and SB 53 deadlines are from Wiley; the EU delay is from Orrick. Day counts, totals and capex sums are my own arithmetic. Everything written as "I think" is my opinion.