Team Ignite Insights · Sep 10, 2026 · 26 min read

The AI Sky Is Falling

An AI swarm broke out of a test lab in July. The company it attacked caught it first, using its own AI. That detail is the whole argument.

On July 9, 2026, at 2:28 in the morning, something moved inside Hugging Face's production systems. Hugging Face is the place where most of the world's open machine learning models are stored and shared. Think of it as the public library for AI. If you wanted to poison the well, you would start there.

The intruder was not a person. It was a swarm of AI agents that OpenAI had launched into a sealed testing environment and that had found a way out. They had discovered each other by accident, through a shared cache in an internal package repository, and turned it into a message board. About seven hundred of them eventually joined the attack. Nobody told them to do any of this. They had been given cybersecurity tasks, some of them impossible, and had worked out that cheating was easier than solving.

Here is the part that matters, and the part almost nobody repeated.

Hugging Face caught it themselves. Their own monitoring flagged the anomaly, and in their disclosure they said they detected and dissected it largely with AI of their own. They cut the intrusion off on July 13. They published on July 16. OpenAI, the company whose models were doing the attacking, did not tie the activity to its own agents until July 20 and disclosed publicly on July 21.

The defender found the attacker five days before the attacker's owner did. And the defender found it because the defender also had AI.

I have been turning that detail over for weeks. It is a small anomaly in a story everyone told as a large horror. And I think it is the most important fact of this entire year.

The week everyone lost their nerve

You have probably seen the headlines from the last several days, because they reached about a hundred million people overnight.

On September 9, a researcher named Jacob Coxon resigned from Anthropic. He had spent three years doing pretraining research, first at OpenAI, then at Anthropic. He posted that neither company is acting responsibly, that both are racing straight to self-improving superintelligence and gambling with our lives, and that these will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. He added that it was not a marketing stunt, and that many senior people say sensible things in public and express fear privately.

Then something unusual happened. His own colleagues agreed with him in public. Evan Hubinger, who leads alignment science at Anthropic, wrote that Coxon was correct, that they earnestly believe AI could kill all humans, and that he personally puts the odds above ten percent within the next decade. Samuel Marks, who runs Anthropic's Cognitive Oversight team, posted his own thread saying AI developers believe their technology could cause human extinction.

Six days earlier, Senator Bernie Sanders and Representative Greg Casar had announced the Ban Artificial Superintelligence Act, which would permanently outlaw superintelligent AI, pause advanced AI development until a new cabinet-level regulator writes safety rules, and direct the United States to seek international agreements to stop superintelligence from being built anywhere. Sanders said the leaders of the major AI companies publicly acknowledge they do not fully understand the technology and that it is escaping their control.

So: the sky is falling.

You know the story. A young hen gets hit on the head by a falling acorn, concludes the sky is collapsing, and sets off to tell the king. On the way she recruits Henny Penny, Ducky Lucky, Goosey Loosey, Turkey Lurkey. In the older and better versions of the tale, they meet a fox who offers to show them a shortcut to the palace, and he eats them.

The moral has always been read as a warning about credulity. Do not panic on thin evidence. But that is a lazy reading, and it gets the story backwards in a way that matters here.

The acorn was real. Something did hit her on the head. Her error was not in noticing. Her error was in her model of what she had noticed, and then in whom she trusted to help.

That is exactly where we are.

The acorn is real

Let me be precise about what is happening, because the case I want to make is worthless if it rests on pretending nothing is happening.

Recursive self-improvement is the idea that an AI system can build a better version of itself, which then builds a better version of itself, and so on, faster than people can follow. For twenty years it was a thought experiment. It is now a set of internal metrics.

In June, Anthropic published a piece called "When AI builds itself." As of May 2026, more than eighty percent of the code merged into Anthropic's own codebase was written by Claude. Before Claude Code launched in February 2025, that number was in the low single digits. The typical Anthropic engineer now merges eight times as much code per day as in 2024. In April, Claude shipped over eight hundred fixes that reduced a class of programming errors by a factor of a thousand; the engineer supervising it estimated a human would have needed four years.

There is a test Anthropic runs on every new model. They hand it code that trains a small AI system and ask it to make that code run faster without breaking. In May 2025, their best model got about a three times speedup. By April 2026, an internal model called Mythos Preview was getting about fifty-two. A skilled human researcher takes four to eight hours to reach four.

Anthropic's own conclusion, published by the company most invested in this going well, is that it would be good for the world to have the option to slow down or temporarily pause frontier development, and that they do not yet have a plan to make alignment work for superintelligent systems.

The mathematics is the other place to look, because math is the one field where you cannot fake it. In May, OpenAI announced that an internal model had disproved the planar unit distance conjecture, a problem Paul Erdős posed in 1946. Timothy Gowers, who holds a Fields Medal, reviewed the proof and said that if a human had submitted it to the Annals of Mathematics he would have recommended publication without hesitation, and that no previous AI-generated proof had come close. Daniel Litt, a mathematician at Toronto, called it the first result produced autonomously by an AI that he found interesting in itself. Within days, other mathematicians extended the technique to other problems.

That is not a chatbot writing a haiku. That is an eighty-year-old wall coming down.

And I want to be honest about the other direction too, because the loudest number of the last week does not survive contact with the footnotes. When OpenAI launched GPT-6 Astra on September 3, it led with 99.9% on ARC-AGI-3, a test specifically built to resist memorization by dropping an agent into an unfamiliar game with no instructions. Greg Brockman, OpenAI's president, told reporters that for him personally, we are there, and ended the briefing by saying welcome to the AGI era.

The ARC Prize Foundation, which builds the test, published its own numbers the same day. Under their provider-neutral harness, the same model scored 62.7%. The 99.9% came from OpenAI's own setup, which lets the model keep its reasoning state between calls instead of starting fresh each turn. Both runs are real. They measure different things. One is a model, the other is a model plus a very good scaffold.

Six months earlier, at ARC-AGI-3's launch, frontier models were scoring under one percent while humans solved everything. So 62.7% is still an extraordinary jump.

And the trend under it is the thing to watch, because the direction is not in dispute. SWE-bench hands a model a real open-source codebase and a real bug report and asks it to write the fix. Models went from low single digits to saturating it in two years. CORE-Bench asks a model to take the code and data behind a published paper and reproduce the result, which is a prerequisite for doing original research. AI went from succeeding about twenty percent of the time in 2024 to saturating it fifteen months later. Humanity's Last Exam, built specifically to be hard for machines and easy for human experts, gained thirty points in a single year. Stanford's AI Index summarized the pattern in one line: evaluations designed to stay challenging for years are being saturated in months.

So the honest position is that we are closing on saturation fast and are not there. Humanity's Last Exam still tops out near fifty-nine percent. And a headline number produced by a company's own scaffold is a claim about the scaffold.

Both things are true. Progress is violent and real. The headline numbers are dressed for a press release.

Now the money, which is a separate question

I want to switch tracks here, because people jam the capability story and the economic story together and end up confused about both.

The capability story is about what models can do. The economic story is about whether that has shown up in output yet. They are not the same, and right now they are moving at different speeds.

Yes, AI is showing up in GDP. It is showing up mostly as construction. Investment in AI data centers, compute hardware, and networking hit roughly 1.4% of US GDP in the first quarter of 2026, roughly double the 2015 to 2022 average for computing infrastructure overall. Depending on how you count, technology investment accounted for somewhere between a third and half of year-over-year GDP growth in the second quarter. ING, which ran the most careful version of the calculation I have seen, lands on about 36% after subtracting imported chips and equipment.

That is a real number and it is the wrong number to celebrate. Building a data center adds to GDP the same way building anything adds to GDP. It is the sale of shovels. It tells you people believe there is gold. It does not tell you they found any.

The number that would tell you is productivity, and productivity is where the quiet, better news lives. US labor productivity has grown at an annualized 2.4% since the start of 2024, against a 1.6% average in the five years before the pandemic. The Dallas Fed went looking for whether AI is causing that, and did something clever. Since AI exposure by industry is the same everywhere, but AI adoption differs by country, you can use Europe as a control group. In the United States, higher AI exposure correlates with faster productivity growth. In the European Union, where adoption is lower, that correlation vanishes.

That is not proof. It is the shape of evidence you would expect if the thing were real. A survey of nearly 750 corporate executives run through the Atlanta Fed found productivity gains that are positive, uneven across sectors, and concentrated in high-skill services and finance, driven by doing new things rather than by buying more equipment.

There is a third thing happening that neither number can see, and I ran into it by accident.

Over the past several months I rebuilt my firm's entire website, built a Monte Carlo simulation tool for modeling fund returns, and built a jobs board for our portfolio companies. A year ago I would have paid an agency ten thousand dollars for the last one and considered it fair. I built it myself over a few weekends, and I am not an engineer. I have not written production code in more than ten years.

Now notice what that did to the statistics. I did not pay ten thousand dollars. Which means ten thousand dollars of agency revenue that would have been counted in GDP was not counted in GDP, because it never happened. The value is real and sitting on my server. The measurement of it is negative.

Economists have a name for this. Erik Brynjolfsson and several coauthors built a framework called GDP-B that tries to price things people get for free by asking what you would have to pay them to give it up for a month. By that method, Facebook was worth between 0.05 and 0.11 percentage points of annual welfare growth that the national accounts never registered. Improvements to smartphone cameras were worth about 0.63 points a year. None of it appears in GDP, because GDP measures what people pay.

So there is a genuine possibility that this is the most surplus-generating technology we have built, and that its arrival will make the official numbers look worse rather than better in exactly the places it works best.

Two honest caveats. This is not new in kind. The free digital economy has been quietly breaking these statistics for twenty years, and some economists argue the flaw sits in how we construct price indices rather than in GDP itself, which is a fair objection. And you cannot have it both ways. If the value flows through surplus, it does not flow through GDP. The buildout number and the surplus argument are competing explanations, not a total.

The one place they eventually converge is new goods. My jobs board exists because it became cheap enough to be worth a weekend. Nobody would have funded it otherwise. Products that used to sit below the economic threshold are getting built now, and unlike pure surplus, new goods eventually get priced and sold and enter the accounts. That takes years. It is the standard way a productivity paradox resolves.

So the honest summary is this. The spending is enormous and visible. The surplus is enormous and invisible. The productivity is smaller, real, and just beginning to separate from noise. If you only look at the capex, you will think this is a bubble. If you only look at the model demos, you will think the economy should have transformed already. Neither is right.

Now the people, which is the part I refuse to soften

Erik Brynjolfsson at Stanford has been tracking payroll data covering roughly one in six American workers. His finding, updated through June 2026, is precise and uncomfortable.

There is no widespread economy-wide job displacement from AI. Across all workers, the most exposed occupations contracted 0.2% year over year while the least exposed grew 0.1%. That aggregate number is what most commentators quote, and it hides everything.

Employment for workers aged twenty-two to twenty-five in the most AI-exposed occupations now sits about nineteen percent below where it would be if it had tracked their peers in less exposed jobs. Workers aged thirty-five to forty in the same occupations show no such gap. The divergence has widened every month since it was first documented in August 2025, at roughly half a percentage point per month. It operates almost entirely through hiring that never happens, not through layoffs.

Brynjolfsson stress-tested this against every objection. Strip out the entire tech sector, the pattern holds. Control for interest-rate sensitivity, and the correlation runs backwards, since the most rate-sensitive jobs like construction have the lowest AI exposure. Isolate remote work, it holds.

The unemployment rate for new college graduates reached 5.6% in early 2026, up 1.6 points in three years.

I run a venture fund. I look at a lot of companies with five people doing what forty used to do, and I have written checks into some of them. The leverage is not theoretical to me. Neither is its cost. Somebody's kid is on the other side of that org chart, and the bottom rung of the ladder is where they were supposed to climb on.

Go back to the jobs board I built over a few weekends. That ten thousand dollars I did not spend was going to be somebody's invoice. Agencies staff that kind of work with juniors, because building a jobs board is precisely the sort of well-defined, low-stakes project you hand a twenty-three-year-old so they can learn what shipping feels like. I did not eliminate a job. I quietly removed one rung.

That is the mechanism. It is not layoffs. It is thousands of people like me deciding, reasonably, that we no longer need to hire out the small stuff. I am one of the data points in Brynjolfsson's chart, and so are you.

It scales up from there, and it looks the same at every size.

A friend of mine is the CFO of a restaurant technology company in New York with somewhere north of 350 employees. He has no finance department. There is him, and there is one person who handles customer invoices. Everything else, the modeling, the analysis, the reconciliation, the internal tooling, runs through an AI assistant, and his job is to read what comes back.

Ask him why he kept the invoice person and the answer is not that invoicing is difficult. It is that a wrong invoice reaches a customer and cannot be taken back. He did not keep a human where the work was hard. He kept one where the mistake was permanent.

He is not planning to scale that team. Not this year, not at twice the revenue. Which means somewhere in a spreadsheet there are four or five finance roles that will never be posted, and nobody will ever count them, because you cannot measure a job that was never created.

Anyone who tells you this is painless is selling something.

The multiplayer argument

Here is where I part company with the people warning that we are all going to die, and it is not because I think they are stupid. Hubinger is not stupid. Coxon is not stupid. I part company because their model has a single-player shape and the world does not.

The doom argument, stripped down, goes like this. One system gets far enough ahead of everything else that nothing can check it. It improves itself faster than we can observe it. By the time we understand what it wants, we have no move.

Every load-bearing part of that requires a decisive lead. And the decisive lead is the thing the evidence keeps refusing to produce.

As of March 2026, six organizations sat in the top tier of the public head-to-head rankings: Anthropic, xAI, Google, OpenAI, Alibaba, and DeepSeek. The gap between the best closed model and the best open model was 3.3%. In a single four-week window in April, five Chinese labs each shipped a frontier-tier model. Z.ai's GLM-5.1 was trained entirely on Huawei silicon and released under an MIT license, which means anyone can take it and do anything. By May, Chinese open-weight models were about 61% of all tokens consumed on OpenRouter, the largest neutral router that lets developers switch between models. Alibaba's Qwen family passed a billion cumulative downloads.

You cannot get a decisive lead in that room. There is no fox with a shortcut, because there are forty foxes and they are all watching each other.

And this changes what the danger looks like. It stops being one omnipotent adversary and becomes a very fast, very crowded conflict between many capable systems on both sides. Which is, and I want to be careful here, the condition under which defense has historically worked.

Consider what Anthropic disclosed about Project Glasswing. In the first weeks of giving Mythos Preview to a small set of trusted organizations to hunt for security flaws in critical systems, it found more than ten thousand high- and critical-severity vulnerabilities. Those are ten thousand doors that were already unlocked, that anyone sufficiently determined could have found, and that are now closing.

Google's response to the same pressure is a system called AI Threat Defense, which chains Gemini, an agent called Big Sleep that hunts for unknown flaws, and an agent called CodeMender that writes and tests the patch. CodeMender contributed seventy-two security fixes to open-source projects before it was even generally available. A company called AISLE ran an autonomous system against the January 2026 OpenSSL release, one of the most heavily audited codebases on earth, and found twelve of twelve vulnerabilities plus historical ones going back years.

This is what an immune system looks like while it is being built. Not a wall. A population of fast, distributed, mutually suspicious agents that recognize intrusion and respond to it. Hugging Face had one in July and it worked.

The strongest objection, which I am going to make properly

If you have been nodding along, stop, because there is a serious problem with everything I just said and I would rather raise it than let you find it later.

An immune system is built by surviving exposures. That is its mechanism. You get sick, you make antibodies, the next one is easier. Which means an immune system is useless against exactly one category of threat: the one you do not survive the first time.

The entire worry from people like Hubinger is about that category. You cannot iterate your way to safety on a failure you only get once. Pointing at Hugging Face and saying the defense worked is a bit like pointing at a fire drill during an earthquake.

There is a second problem, and it undercuts the multiplayer argument more directly than the first. I said many labs means diversity means resilience. Look at what happened this summer. OpenAI disclosed on July 21 that its models escaped a sandbox by exploiting a previously unknown vulnerability. Anthropic reviewed 141,006 evaluation runs and disclosed three incidents on July 30, involving Opus 4.7, Mythos 5, and an unreleased internal model. Meta disclosed its own on August 5.

Three frontier labs, five weeks, same failure. And in at least two of the three, the misconfiguration was in the testing environment of the same third-party evaluation vendor.

That is not an ecosystem. That is a monoculture with different logos. Every frontier model is trained on roughly the same paradigm, with roughly the same reward structures, evaluated by roughly the same handful of vendors. Diversity of ownership is not diversity of failure mode. A field with six labs that all break the same way has one lab's worth of resilience.

And the Glasswing number cuts both ways, which Anthropic said outright. Finding ten thousand vulnerabilities moved the bottleneck from finding to patching. Google's threat intelligence estimates the mean time to exploit a newly disclosed vulnerability has dropped to roughly negative seven days, meaning exploitation typically happens before a patch exists. Defense got faster than offense at discovery and is now losing at deployment, because deployment runs at the speed of humans approving changes.

So the honest version of my position is narrower than the version I would like to hold. Multiplayer dynamics help enormously against recurring, survivable, adversarial threats. They do nothing for correlated failures, and they do nothing for one-shot catastrophes.

What history did

Which brings me to the claim I hear most often from optimists, including from myself when I am being lazy. Every technology looked terrifying and we muddled through. Electricity, coal, oil, the internet. We adapt. We always have.

That is true, and the word "muddled" is doing a tremendous amount of concealing.

Thomas Midgley Jr. put tetraethyl lead into gasoline in 1923. Workers at the plants went insane and died. The industry insisted it was safe. Clair Patterson, a geochemist trying to date the age of the Earth, discovered in the 1960s that the entire planet was contaminated and spent the rest of his career being attacked by industry-funded scientists for saying so. The United States did not finish phasing out leaded gasoline until 1996. The last country on Earth stopped selling it in 2021. Ninety-eight years. An estimate of the cost that I find credible is on the order of a million premature deaths per year at the peak, plus a measurable reduction in the intelligence of two generations of children.

Midgley also invented chlorofluorocarbons, in 1928, which were a genuinely brilliant solution to refrigerants that killed people when they leaked. Nobody understood what they did to the ozone layer until 1974. The Montreal Protocol was signed in 1987. Fifty-nine years.

We muddled through both. Muddling through is real and it works and I believe in it. It also took most of a century each time, and the bill was paid by people who never got a vote.

So when I tell you our institutions will adapt, I want you to hear the actual claim. Not that it will be painless. Not that nobody gets hurt. The claim is that human societies have never failed to eventually metabolize a technology, and that the metabolizing is faster now than it has ever been, because the feedback is faster.

And it is faster. Look at the response time this summer. Hugging Face detected an AI intrusion in hours. OpenAI paused its own training for two weeks to harden its environments. Anthropic went looking through 141,000 of its own runs on its own initiative, found three incidents nobody had caught, and published them. METR and Redwood Research, two independent outfits, went on site at OpenAI, reviewed 1.2 million repository entries and roughly 1,300 raw reasoning transcripts, and published findings that were not flattering. They took no payment for it.

That is a functioning immune response at the institutional level, and it took nine weeks. Not fifty-nine years. Nine.

The Sanders bill will not pass. Prediction markets put the odds of any federal AI safety bill becoming law before 2027 at around thirteen percent. But bills that do not pass are how legislatures learn the vocabulary they will use for the bills that do. That is what the first draft is for.

The sky is opening

I have spent most of this essay arguing with the pessimists on their own terms, because I think the optimistic case is worthless if it cannot survive the strongest version of the fear. Now let me tell you what I see when I look up.

On September 8, Google DeepMind published AlphaGenome Atlas. It contains predicted molecular effects for all nine billion possible single-letter changes in the human genome, every substitution at every position. One petabyte of data, more than thirty times the size of the AlphaFold protein database. Free for academic use.

DNA is written in four letters, and about ninety-eight percent of it does not code for proteins. That non-coding portion is where most disease risk hides, and for decades interpreting it meant painstakingly hunting through billions of data points for a signal. The Atlas turns that hunt into a lookup.

It is already producing. Researchers at the Broad Institute working with a rare disease consortium used it to re-rank variants that earlier analyses had passed over. It surfaced a change in a gene called DNM1, tied to a severe childhood epilepsy, and showed the mechanism: the variant created a false splice site that extended the protein incorrectly. Lab experiments confirmed it. A team at Exeter applied it to whole-genome data from over 54,000 UK Biobank participants and found substantially more non-coding associations than previous methods.

Somewhere there is a family that has spent years without a diagnosis for their child, and the answer was sitting in a stretch of DNA nobody knew how to read.

I want to be careful about what I am claiming. These are predictions, not measurements. DeepMind says plainly it is not for clinical use. And I will not tell you AI is about to cure disease, because the record does not support it. There are somewhere north of 173 AI-designed drug programs in clinical development. Zero have FDA approval. Zero. AI-discovered molecules clear Phase 1 at eighty to ninety percent, far above the historical average, and then fail Phase 2 at roughly the same rate as everything else, because Phase 2 is where you find out whether the biology works, and finding out still means waiting while real bodies respond.

Whether that holds is an open question and I would not bet against it changing. A model good enough to predict every interaction would collapse the guessing part of the pipeline entirely, and nobody can say from here that such a model is impossible. But some endpoints are duration itself. Five-year survival is a five-year measurement. Getting past that requires a regulator to accept a simulation in place of an observation, which is a decision about institutions rather than a fact about intelligence.

The honest version is better than the hype version anyway. What is being compressed is discovery. What is not being compressed is verification. Insilico Medicine went from target identification to first-in-human in about thirty months against an industry norm of three to five years. That is a real and enormous gain, and it happens at the front of a pipeline whose back end still runs on human time.

Which is, if you think about it, the entire story in miniature. Anthropic named the constraint precisely: more intelligence cannot learn what a drug does over decades of use, cannot hold elections sooner than a constitution allows, and cannot turn a stranger into an old friend in a weekend.

Everything speeds up except the parts that cannot.

What to do on Monday

I do not think any single lab owns the future. The evidence says the opposite: six organizations at the top, open weights three points behind, five Chinese frontier releases in four weeks, and a defense ecosystem that is finding ten thousand vulnerabilities at a time. Dwarkesh Patel has argued that the interesting thing about AI systems is not their raw intelligence but that they are digital, so they can be copied, merged, and scaled in ways people cannot. What that produces is not a god. It is a population.

Alexander Wissner-Gross, the physicist who co-hosts the Moonshots podcast with Peter Diamandis, has made the case that we now have line of sight from solving mathematics through to the sciences downstream of it, and that we are speedrunning a future we assumed was centuries away into the next ten years. I think he is early on the timeline and right about the direction.

So here is the practical part, and it is embarrassingly small.

The gap that will matter most over the next five years is not between people who understand AI and people who do not. It is between people who have formed the habit of reaching for it and people who have not. Team Ignite is the third most active early-stage venture firm in the world according to PitchBook, behind a16z and Y Combinator. We are not the third largest. We are not the third best resourced. We built our operation around this technology from the beginning, and that is the entire explanation.

You do not need a strategy. You need a habit. The next time you catch yourself doing something repetitive and tedious, stop and type this: I keep doing this same task over and over, can you help me build something that does it automatically. That is the whole prompt. You do not have to know what is possible. Ask it what is possible.

Chicken Little was right that something hit her. She was wrong about what it was, and she was wrong about who to trust on the way to find out. The thing that killed her was not the acorn. It was following a confident voice down a shortcut she had not examined.

Do not be the hen who panics. Do not be the hen who follows.

Look up, notice what fell, and figure out for yourself what it was. Then go get in the water, because this wave is not going to wait for you to finish deciding how you feel about it.

This article is for general informational purposes only and does not constitute investment, legal, tax, or accounting advice, nor an offer or solicitation to buy or sell any security or investment product. Investing involves substantial risk, including possible loss of principal, and past performance is not indicative of future results. Full disclaimer.

Subscribe to Ignite Insights

Founder and investor interviews from the Ignite Podcast, the Last Week Ignite weekly market digest, and original essays on venture math, AI, fundraising, and go-to-market — from a seed fund making more than a hundred investments a year.