Opus 5: No-Hype Full Review & Testing

Opus 5: No-Hype Full Review & Testing

🎙 Pat Simmons 👥 24K 📅 July 25, 2026 ⏱ 27 min 👁 69K 📄 original study 🧭 2026-09-07
Available in: English (current) Français

Keywords

Opus 5Fable 5benchmarkweb design3D simulationgame developmentmotion graphicsknowledge workcost analysisAI testing

Summary

In this video, Pat Simmons conducts a hands-on comparison of Anthropic’s Opus 5 against Fable 5 and Opus 4.8. He begins by analyzing Anthropic’s official benchmarks, noting that Opus 5 reportedly outperforms Fable 5 across most metrics at half the cost, which he finds surprising and suggests a possible upcoming Fable 5.1. He then runs five practical tests: web design, 3D and simulation, playable game, motion graphics, and knowledge work (SpaceX stock analysis). Each test uses identical prompts, and the outputs are presented in a blind format before revealing the model. In all tests, Opus 5 consistently produces superior results, with more creative and technically advanced outputs, while also being more cost-efficient in most cases. The video includes detailed cost breakdowns and highlights Opus 5’s strengths in creativity, technical execution, and efficiency. The creator concludes that Opus 5 is a significant upgrade and currently the best choice among the compared models, though he notes potential limitations such as the small sample size and the subjective nature of the evaluations.

169 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial value by moving beyond official benchmarks and offering practical, real-world testing across diverse creative and knowledge-based tasks. The argumentation is solid: the creator uses a consistent methodology (identical prompts, blind ranking, cost tracking) and supports his claims with visual evidence and detailed cost data. He also acknowledges potential biases and limitations, such as the influence of prompt phrasing and the small number of tests, which strengthens the credibility of his conclusions. The comparison is thorough, and the conclusion that Opus 5 outperforms Fable 5 is well-supported by the presented evidence.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a rigorous approach to testing, with a clear methodology and transparent reporting of results. The creator references Anthropic’s official benchmarks and provides links to his own blog post with prompts and generations, which adds to the transparency. However, the evaluation is based on a limited number of tests and is inherently subjective, as the ranking relies on the creator’s personal judgment. The title accurately reflects the content, and the video does not overhype or make unsupported claims. The creator also discusses the cost implications, which is a valuable addition. Overall, the sources are credible (official benchmarks, personal testing), and the title is appropriate.

215 words

Title / Content Match

The title accurately reflects the content: a comprehensive, no-hype review and testing of Opus 5.

Quality & Reliability

8/10

The video presents a structured, hands-on evaluation of Opus 5 against Fable 5 and Opus 4.8, with transparent methodology (blind tests, multiple scenarios) and detailed cost analysis. The creator acknowledges limitations (e.g., small sample, potential bias in prompts) and does not overstate conclusions. However, the evaluation is subjective and not peer-reviewed, and the creator's own benchmarks are not independently verified.

Chapters

Cited Sources

Concurring Sources

Dissenting Sources

  • Commenter claims Mythos 5 and Fable 5 are the same model — A commenter suggests that Mythos 5 and Fable 5 are the same underlying model, which contradicts the video's implication that they are distinct.

Contribution & Novelties

The video contributes original, hands-on testing of a newly released AI model (Opus 5) against its predecessor and a competitor, providing practical insights beyond official benchmarks. It introduces a structured methodology for evaluating creative and knowledge-work capabilities, including cost analysis, which is often overlooked. The blind testing approach adds credibility to the comparisons.

Pour aller plus loin :

111 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a detailed and technically rich video. The quality of information and global reliability are also strong, reflecting the structured methodology and transparent reporting. The overall profile suggests a highly informative and reliable review.

Reliability 7/10

💬 Très positif : Les commentaires sont extrêmement favorables, louant la qualité des tests, l'honnêteté et la clarté de la présentation. Sur les 30 commentaires analysés, la grande majorité exprime une admiration pour le travail du créateur et demande un sponsor pour soutenir la chaîne.