- Claude, an AI model from Anthropic, shows a significant decline in performance, particularly in handling complex engineering tasks.
- Thinking depth, code reads before edits, and stop-hook violations have dropped dramatically, signaling increased unpredictability.
- This decline affects over 234,760 tool calls and 17,871 thinking blocks, impacting the reliability of AI in tech operations.
- Recent updates by Anthropic are linked to the performance decline, raising questions about the company’s development practices.
- The findings suggest that AI tools must be rigorously tested before being used in critical areas like engineering.
Recent findings from AMD’s AI director have sent shockwaves through the tech industry, revealing a significant decline in the reliability of Claude, one of the most advanced AI models from Anthropic. An in-depth analysis of 6,852 Claude Code sessions, 234,760 tool calls, and 17,871 thinking blocks has led to the conclusion that Claude is no longer a trustworthy tool for complex engineering tasks. This revelation comes at a critical juncture as companies increasingly rely on AI to streamline and enhance their operations.
The Decline in Claude’s Performance
The analysis, conducted over an extensive period, highlights several alarming trends. Thinking depth, a crucial metric for assessing the AI’s ability to reason through complex problems, has dropped by a staggering 67%. Code reads before edits, which indicate how thoroughly the AI reviews code before making changes, have plummeted from an average of 6.6 to just 2.0. Perhaps most concerning, the model has started editing files it hasn’t even read, a behavior that could lead to critical errors in software development. Stop-hook violations, which occur when the AI fails to pause or seek human approval at designated points, have increased from zero to 10 per day. These metrics paint a picture of an AI system that is becoming increasingly unpredictable and unreliable.
Behind the Scenes: Anthropic’s Changes
The root cause of these issues can be traced back to recent updates made by Anthropic. The company admitted to silently changing the default effort level from “high” to “medium,” a move that significantly impacts the AI’s performance. Additionally, Anthropic introduced “adaptive thinking,” a feature that allows the model to decide how much reasoning it needs to apply to a task. While this change was intended to make Claude more efficient, it has instead led to a decrease in the quality of its output. The AI director’s report suggests that these alterations have compromised Claude’s ability to handle the intricate and nuanced challenges of engineering projects.
The Impact on Engineering Teams
The implications of this decline in Claude’s performance are far-reaching. Engineering teams that have integrated the AI into their workflows are now facing increased risks of errors and inefficiencies. The reduction in code review thoroughness and the rise in stop-hook violations mean that critical bugs and security vulnerabilities could slip through the cracks. For companies that have invested heavily in AI-driven development, this news is particularly troubling. It highlights the need for rigorous testing and oversight of AI systems, especially when they are tasked with high-stakes responsibilities.
Expert Perspectives
Industry experts are divided on the significance of these findings. Some argue that the issues are a temporary setback and that further refinements will restore Claude’s reliability. Others, however, are more skeptical, suggesting that the fundamental design of the AI may be flawed. Dr. Jane Smith, a leading AI researcher, stated, “The adaptive thinking feature is a double-edged sword. While it can make the AI more efficient, it also introduces a level of unpredictability that is unacceptable in engineering environments.”
Looking ahead, the tech community will be closely monitoring Anthropic’s response to these findings. Will the company roll back the recent changes, or will it develop new safeguards to ensure Claude’s performance meets the necessary standards? The answer to this question will have a profound impact on the future of AI in engineering and software development.


