
ℹ️ Quick Answer: Was GPT-6 Astra nerfed? Users say it got worse, and OpenAI confirmed three problems hurting its quality. A week after the September 3 release, people posted same-prompt tests showing weaker output, and one fix turned off a context experiment affecting 4,000 to 5,000 users. Claude Fable 5 took the same slide in July.
📋 WHAT’S INSIDE
- What GPT-6 Astra Users Noticed
- What OpenAI Actually Admitted
- Fable 5 Went Through the Same Thing
- Why It Keeps Looking Like a Pattern
- How to Tell If Your AI Got Worse
- Frequently Asked Questions
Last updated September 12, 2026
When Claude Fable 5 came out on June 9, it was amazing. It was a huge jump from Opus, and I wrote up my first 24 hours with it. Three days later a US export-control order knocked it offline. When Anthropic brought it back on July 1, it wasn’t the same. Then Fable 5.1 shipped on September 1, and it got closer to the original, but it still wasn’t the original.
So when GPT-6 Astra users started posting launch-day versus today comparisons this week, I recognized it. It’s starting to look like a pattern to me. Release a model at full capacity, let people make impressive things with it, let the benchmarks look impressive, and then quietly pull it back so the company can support all the new usage.
What GPT-6 Astra Users Noticed

About a week in, developers started running the exact same prompt against launch-day Astra and the current version, and posting the weaker results side by side.
One post on X showed both outputs with the note “Identical prompt. Same settings. Today’s output is clearly weaker,” and added, “Not saying it’s been nerfed, but something has definitely changed.” Decrypt rounded up a week of posts like it. A thread on OpenAI’s own developer forum is titled “A noticeable drop in quality compared to how it was at launch,” and a bug report on OpenAI’s Codex repo describes the model ending its turn after about 30 seconds and reporting work as finished that was never done. Over on Reddit, an r/codex post with more than 1,000 upvotes put Blender scenes from September 9 and 10 next to what came out after, and one reply with over 100 upvotes laid out the same cycle I described up top.
Dax Raad, who builds the coding tool Opencode, wrote that part of his team had gone back to the older GPT-5.6 Sol because “our effective spend looks doubled so tough to justify.” Not everyone agreed. The most detailed rebuttal Decrypt found argued Astra “is as dumb as it was on launch” and people were simply overhyped that first week.
What OpenAI Actually Admitted
OpenAI confirmed three problems that were dragging Astra’s quality down, fixed them, and reset everyone’s usage limits.
Tibo Sottiaux runs Codex and ChatGPT over there, and he listed all three himself. Skills written for older models were firing too often or stopping Astra from checking its own work. An opt-in context management experiment could cause early stops or replies to older messages, and he wrote that “our rough estimate is that 4-5k users were affected.” They also removed some badly configured engines that were giving a slice of users worse answers.
And that same week, OpenAI was straining to keep up. Sottiaux posted that demand was “really unprecedented” and they were pulling every lever they had, and on September 10 OpenAI paused new $200 Pro sign-ups because that plan puts the most strain on its systems. This also isn’t OpenAI’s first round of it. In July, users said GPT-5.6 Sol’s top reasoning mode had gone shallow, and Sottiaux denied weakening it on purpose while confirming the company had been experimenting with reasoning effort, which is how long the model thinks before it answers.
Fable 5 Went Through the Same Thing

Fable 5 came back from its ban with a stricter safety filter that sent a lot of ordinary coding requests to the older Claude Opus 4.8, and a benchmark caught it the next day.
The June 12 shutdown came from a US export-control order, after Amazon researchers found a way to get Fable 5 to identify software vulnerabilities. When Anthropic brought it back on July 1, it added a classifier that blocks flagged requests and sends them to Opus 4.8 instead. Anthropic said up front that the classifier “comes at the cost of flagging benign requests more often during routine coding and debugging tasks.” In other words, they told everyone normal coding was going to get caught.
BridgeMind reran its BridgeBench coding tests on July 2, and debugging went from 86.2 down to 25.9. Refactoring got cut about in half. A breakdown on Yahoo Tech found that 9 of the 12 debugging tasks never reached Fable 5 at all. They got rerouted to Opus 4.8 and scored as zero. Arena’s blind human voting the same day showed coding down 18 points while document work went up, and BleepingComputer reported Reddit users saying the restored model felt weaker.
Then came 5.1. Anthropic’s launch post says its newest cybersecurity safeguards “block 60% fewer false positives than before,” and its own benchmark notes say cybersecurity tasks the safeguards intercepted were completed by Opus 4.8. That matches what I felt, closer but still not the original.
Why It Keeps Looking Like a Pattern

It’s starting to look like the model you’re using a few weeks in isn’t the one that posted the launch benchmarks.
On paper, Fable’s reason was a government order, and Astra’s were bugs and an experiment. Anthropic has also said, in a 2025 postmortem on earlier Claude complaints, “We never reduce model quality due to demand, time of day, or server load,” and that those problems came from infrastructure bugs.
I still think it comes down to capacity. Astra’s quality complaints landed the same week OpenAI called demand unprecedented and closed the door on new Pro subscribers. The Sol complaints came while OpenAI was experimenting with reasoning effort. On Anthropic’s side, Claude Code usage limits drop 17% on September 14.
How to Tell If Your AI Got Worse
Save a prompt that worked in launch week and run it again later, because that’s the test that caught Astra.
- Keep two or three real prompts and their answers from the first week, then rerun them with the same settings a few weeks later.
- Watch for fallback notices. Anthropic says it tells you when a Fable request gets blocked and sent to Opus 4.8.
- Check what the company has posted before you decide it’s in your head. OpenAI’s fix list and Anthropic’s safeguard notes both explained real changes.
Frequently Asked Questions
Was GPT-6 Astra nerfed?
OpenAI confirmed three problems hurting GPT-6 Astra’s quality: skills from older models firing too often, an opt-in context experiment that could cause early stops and affected an estimated 4,000 to 5,000 users, and badly configured engines that degraded a long tail of traffic. It fixed them and reset usage limits. OpenAI has not said it lowered quality on purpose.
Why did Claude Fable 5 get worse after it came back?
When Fable 5 returned on July 1 after a US export-control suspension, Anthropic added a stricter safety classifier that sends flagged requests to Claude Opus 4.8, and said it would flag benign coding and debugging requests more often. BridgeBench’s debugging score fell from 86.2 to 25.9, largely because rerouted tasks were scored as zero.
Is Claude Fable 5.1 better than Fable 5?
Anthropic says Fable 5.1’s cybersecurity safeguards block 60% fewer false positives than before, which means fewer ordinary requests get intercepted. Tasks the safeguards still catch are completed by Claude Opus 4.8, so it is closer to the original Fable 5 experience but not identical.
I’ve watched it happen twice now, with two different companies. That’s starting to look like a pattern to me.
Related reading: My first 24 hours with Claude Fable 5 | GPT-5.6 Sol vs Claude Fable 5 | Claude Code usage limits are dropping 17% | New to AI? Start here
WHO WROTE THIS
Moses Smith. I write Everyday AI for people who aren’t engineers. I go try the tools, then tell you honestly whether they were worth it. Sometimes the answer is no, and that’s kind of the point.
This blog is free and has no ads. If it saved you some time, you can buy me a coffee.









Leave a Reply