Grok 4 is so cracked

Grok 4 is so cracked

🎙 AI Search 👥 727K 📅 July 15, 2025 ⏱ 44 min 👁 183K 📄 review 🧭 2026-09-07
Available in: English (current) Français

Keywords

Grok 4AI capabilitiesprompt engineering3D simulationbenchmarks

Summary

The video is a comprehensive review and tutorial of Grok 4, xAI’s latest AI model. The creator demonstrates Grok 4’s capabilities through a series of tests, including generating interactive 3D simulations (asteroid impact, particle visualizers, black hole), creating maps with layers, performing financial analysis from PDFs, solving an IMO-level math problem, and solving visual puzzles from the ARC-AGI benchmark. The video emphasizes the importance of prompt engineering, showing that Grok 4 requires more specific instructions to avoid errors compared to other models. The creator also covers Grok 4’s specs, versions (Grok 4 and Grok 4 Heavy), and performance on benchmarks like ARC-AGI and other leaderboards. The video concludes with a discussion of Grok 4’s tendency to hallucinate and its overall potential, positioning it as a leading model for reasoning and coding tasks.

132 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial value by showcasing Grok 4’s practical applications in coding, science, and math, with detailed demonstrations. The argumentation is based on direct testing and comparison with other models like Gemini 2.5 Pro, highlighting Grok 4’s strengths and weaknesses. However, the evaluation is subjective and lacks rigorous scientific controls, as the creator’s prompt engineering skills significantly influence the outcomes. The video effectively argues that Grok 4’s performance is highly dependent on prompt quality, which is a key insight for users.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a moderate level of scientific rigor, with the creator providing clear examples and verifying some outputs (e.g., financial data). However, the evaluation is anecdotal and not peer-reviewed. The sources cited are primarily official links (grok.com, x.ai) and the creator’s own tools, which are relevant but not exhaustive. The title accurately reflects the content, which is a positive review of Grok 4’s capabilities. The video does not delve into potential biases or limitations beyond hallucinations, and the lack of a systematic methodology reduces its scientific credibility.

185 words

Title / Content Match

The title accurately reflects the content, which showcases Grok 4's impressive capabilities through various tests and demonstrations.

Quality & Reliability

7/10

The video provides a hands-on review of Grok 4 with multiple demonstrations, but lacks rigorous scientific methodology and relies on anecdotal evidence. The creator's expertise is evident in prompt engineering, but the evaluation is subjective and not peer-reviewed.

Chapters

Cited Sources

  • Grok — Official Grok platform for accessing the model.
  • xAI News: Grok 4 — Official announcement and details about Grok 4.
  • AI Search Tools — Creator's website for AI tools and jobs.
  • AI Search Newsletter — Creator's newsletter for AI updates.

Concurring Sources

  • xAI News: Grok 4 — Official announcement confirming Grok 4's release and features.

External References

Contribution & Novelties

The video offers a practical, hands-on evaluation of Grok 4, demonstrating its capabilities in generating complex 3D simulations and solving challenging problems. The emphasis on prompt engineering as a critical factor for performance is a valuable takeaway. The video also provides comparisons with other models, offering insights into Grok 4’s relative strengths.

Pour aller plus loin :

  • ARC-AGI Benchmark — The benchmark used to test visual reasoning, relevant to the video’s discussion of Grok 4’s performance.
  • International Mathematical Olympiad — The source of the math problem solved in the video, illustrating Grok 4’s mathematical reasoning.
  • Three.js — The JavaScript library used for 3D visualizations, central to many demonstrations in the video.

111 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, reflecting the video's detailed demonstrations. Quality of information and reliability are moderate, indicating a subjective review with some verified outputs. Overall, the video is informative but not fully rigorous.

Reliability 6/10

💬 Sur les 30 commentaires analysés, le climat est très positif, avec des réactions enthousiastes aux démonstrations, notamment les visualisations 3D et les capacités de raisonnement. Certains commentaires expriment des réserves sur le taux d'hallucinations et la nécessité de 'handholding' dans les prompts.