The Future of Life Institute's AI Safety Index Summer 2026 graded 9 AI companies with 7 experts across 6 domains: none scored an A or a B. The best was Anthropic with a C+ (2.66); xAI (0.65), DeepSeek (0.47), and Mistral (0.33) failed with an F.
Picture the end-of-year report cards for an entire industry, and not a single honor roll name shows up. No top marks. Not even a solid B. That is exactly what happened when the Future of Life Institute published its AI Safety Index Summer 2026: nine of the most powerful AI companies on the planet sat the exam, and none of them scored an A or a B. The top of the class, Anthropic, scraped a pass with a C+. And at the other end, three labs flunked outright. When the star student barely clears the pass mark, the problem isn't one student: it's the whole syllabus.
## How to build a report card the industry can't ignore
What gives this report its weight isn't the grade, it's the method. The Future of Life Institute isn't a rival company throwing stones: it's an independent organization that has spent years pushing for safe AI. For the index it assembled a panel of seven outside experts who reviewed each company's public evidence through June 3, 2026, spread across six broad domains — risk assessment, current harms, safety frameworks, existential safety, governance, and transparency — broken down into 37 concrete indicators. All of that was translated into a GPA-style grading scale, the same one used at U.S. universities. It is, as of today, the only AI safety ranking that is independent and has a public methodology. In other words: this isn't dinner-table opinion, it's the most serious thermometer that exists.
## The honor roll that honors no one
Let's get to the numbers, because they sting. Anthropic topped the table with a 2.66 (C+), leading in five of the six domains evaluated — first in the class, but a first that still has homework to hand in. Behind it, OpenAI scored a 2.28 (C), pulling ahead in the risk assessment domain, and Google DeepMind a 2.01 (C): bare passes. Meta landed at a D+, and Z.ai and Alibaba Cloud tied at a weak D-. And then the pack of failing grades: xAI with a 0.65 (F), DeepSeek with 0.47 (F), and Mistral closing out the list with a 0.33 (F). Translation: most of the sector, including names you use or hear about every day, doesn't even clear the pass mark on safety. And we're talking about the companies building the systems that are increasingly making decisions for us.
## The uncomfortable detail almost no one puts in the headline
There's one data point in the report that you read in passing but that should freeze your smile: the reviewers noted that, between 2024 and 2026, Anthropic, OpenAI, Google DeepMind, and Meta reversed their bans on military use and are now actively seeking defense partnerships. That is, the four best-positioned labs on the table — the ones supposedly leading on responsibility — walked back a red line they had drawn for themselves. This isn't a technicality: it's proof of where gravity pulls once the race tightens. When competition heats up, the safety guardrails are the first thing to loosen, and whoever promises one thing on the corporate website may be signing the opposite in a contract. That's exactly why an independent index matters so much: it measures what gets done, not what gets preached.
## What it means for anyone building with AI
The lesson from this report card is as clear as it is uncomfortable: AI safety doesn't come built in, not even at the labs that lead on it. If the best in the world barely passes and several fail, responsibility can't be handed off entirely to whichever model you're using; it has to be built into the layer where your product and your data live. And we'll say this without any smoke: NeuralOS doesn't build frontier models, nor does it claim to fix the safety of the entire industry. What we do is take guardrails seriously on our own turf — per-tenant data isolation, encrypted credentials, an adversarial audit protocol on every sprint — and stay model-agnostic precisely so we're not tied to whoever scores a C+ today and who-knows-what tomorrow. If this index leaves us with anything, it's that in AI, trusting is easy and verifying is the real work. We prefer the work.