
Let's Run Qwen 3.5 - Local AI HERO Model for OpenClaw, Writing, Coding & More
Keywords
Summary
127 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides concrete, reproducible data points: token speeds, memory usage, and completion quality across various tasks. The presenter grounds his claims through actual runs, comparing Qwen 3.5 against models like Kimi K2.5 and GLM, and highlights both strengths (creative writing, agent integration) and weaknesses (coding plagiarism, logic puzzles with thinking disabled). The argumentation is logical and transparent about hardware requirements, quantization trade-offs, and the need for enabling thinking for reliable tool calls. However, the absence of independent benchmark verification and the host’s admitted bias toward Qwen slightly weaken the objectivity.
Scientific Rigor, Source Quality, Title Accuracy
The video cites Hugging Face and Inferencer as primary references, and provides companion video links to other model evaluations. The presenter does not cite academic papers, relying instead on empirical testing, which is acceptable for a practical tutorial. The title accurately reflects the content, promising a local run of Qwen 3.5, which is fulfilled. The description includes affiliate links, but they do not affect the technical content. The host is transparent about assumptions and limitations, such as the use of Q9 quants and emulation for FP8. Overall, the information is presented with reasonable rigor, but the lack of external verification of claimed benchmarks restricts the score.
212 words
Title / Content Match
The title accurately reflects the content: running Qwen 3.5 locally, with emphasis on its performance for agent tasks, creative writing, and coding.
Quality & Reliability
7/10
The video provides a hands-on demonstration with multiple benchmarks, real-time token rates, and integration tests (OpenClaw, Kilo Code). The presenter is transparent about limitations and quantization choices, but the content includes affiliate links and the benchmark sources are not independently verified.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction and benchmarks of Qwen 3.5 against Gemini, Claude, and GPT
- Model specs: 397B params, 17B active, memory requirements
- Speed test: 25 tokens/sec with thinking disabled
- Batching demonstration: multiple inferences simultaneous
- Tool calling to ground information from Wikipedia
- OpenClaw integration: web search and Apple Notes
- Kilo Code setup and Flappy Birds generation (copied file)
- 3D Earth visualization and creative writing samples
- Minecraft attempt and logic puzzles (car wash, surgeon)
Cited Sources
- Qwen3.5-397B on Hugging Face — Model download and details
- Inferencer App — The inference software used for testing
- Kimi K2.5 Local Cluster — Companion video of another local model
- Kimi K2.5 with OpenClaw — Companion video of agent integration
- GLM-5 — Companion video of another model
- FLUX 2 Klein — Companion video of image gen
Concurring Sources
- Qwen 2.5 official blog — Previous Qwen releases that show performance trends.
External References
Contribution & Novelties
This video offers an original, real-time assessment of Qwen 3.5’s local performance, particularly its 397B-parameter open-weight architecture with only 17B active parameters, making it feasible for high-end consumer hardware. It highlights the model’s creative writing capabilities and its integration with agentic tools like OpenClaw, while also exposing practical quirks like prompt caching and quantization trade-offs. The host’s candid discussion of failures (e.g., coding plagiarism) provides useful insights for the community.
Pour aller plus loin :
- Qwen official — The official model family page with documentation and variations.
- Mixture of Experts — Background on the MoE architecture that Qwen 3.5 uses.
- Prompt Caching — Explanation of how caching reduces latency in LLM inference.
- Quantization in Large Language Models — Notes on quantizing models for local deployment (no URL if uncertain, but this is a known blog).
135 words
Radar Profile
The radar profile is balanced, with high scores in information quantity and practicality, and moderate scores in quality/rigor. The low level of technical depth compared to research papers is offset by the step-by-step tutorial nature.