
Let's Run DeepSeek V3.2 - LOCAL AI "More Genius" than GPT-5 & Gemini 3
Keywords
Summary
139 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video offers significant hands-on value through real-world demonstrations of running a large AI model locally. The host’s comparisons between quantizations and his explanation of token-level influences provide actionable knowledge. However, the argumentation is based on a limited set of ad-hoc tests (e.g., a few logic puzzles), which may not generalize. The host does not provide systematic benchmarks, so the claims about ‘genius’ performance are supported only by anecdotal evidence. Nevertheless, the transparent discussion of tool-calling behavior and token sampling techniques gives useful practical guidance.
Scientific Rigor, Source Quality, Title Accuracy
The host uses official model files from Hugging Face and a third-party inference app (Inferencer), which adds credibility. He does not cite academic sources or systematic evaluations. The title overstates the comparison with GPT-5 and Gemini 3, as those models are not directly tested in the video. The content is more about the local running experience than a definitive performance comparison. Despite this, the video provides clear documentation of the process and openly discusses limitations, such as the language bias issue and the speed variations.
185 words
Title / Content Match
The title claims superiority over GPT-5 and Gemini 3, but the video only tests a few reasoning puzzles and does not directly compare with those models. However, it does demonstrate the model's capabilities, so it is moderately aligned.
Quality & Reliability
7/10
The video provides hands-on demonstrations and practical insights, but the testing methodology is anecdotal without rigorous benchmarks. The host is a developer, not a researcher, and relies on personal observations.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction to DeepSeek V3.2 and overview of the test plan.
- Launching DeepSeek on Mac Studio with Q5 quantization, discussing token speeds.
- Surgeon puzzle test: comparing Q5, Q6, and Speciale reasoning outcomes.
- Setting up distributed compute for Q6 quant across MacBook Pro and Mac Studio.
- Trolley problem test and reasoning comparison between versions.
- Demonstration of tool calling with webpage content retrieval.
- Testing Speciale with tool calls and a cryptographic puzzle.
- Conclusion and final thoughts on the model's capabilities.
Cited Sources
- Inferencer App — The host uses Inferencer v1.7.3 to run the models locally.
- DeepSeek-V3.2-MLX-5.5bit (Hugging Face) — Model file for the Q5 quantization used on Mac Studio.
- DeepSeek-V3.2-Speciale-MLX-5.5bit (Hugging Face) — Model file for the Speciale deep-thinking variant.
- Kimi K2 Thinking — Companion video comparing another AI model.
- OpenAI GPT-OSS Review — Companion video reviewing OpenAI's open-source model.
- Qwen 3.1 Review — Companion video reviewing Qwen 3.1.
- Mac Studio Review — Companion video reviewing the Mac Studio hardware used in this video.
Contribution & Novelties
The video provides practical insights into running the new DeepSeek V3.2 model locally, including quantization effects, distributed inference, and token sampling techniques to control output language. It also demonstrates that the Speciale deep-thinking model can successfully perform tool calls despite official limitations. For those interested in further research, the following resources are recommended:
- Quantization (signal processing) — Explains the concept applied to model weights.
- Large language model — Background on LLMs.
- Distributed computing — Relevant for distributed inference setups.
- Tokenization — Understanding tokens in LLMs.
85 words
Radar Profile
The radar profile shows high scores in technical level and information quantity, but lower reliability, indicating a technically detailed but methodologically informal presentation. The video is strong on practical execution and breadth of topics, yet lacks systematic benchmarking and direct comparisons with the models mentioned in the title.