Let's Run Devstral-2 123B & 24B - France's LOCAL Coding AI REVIEW

Let's Run Devstral-2 123B & 24B - France's LOCAL Coding AI REVIEW

🎙 xCreate 👥 26K 📅 December 10, 2025 ⏱ 18 min 👁 9K 📄 expert opinion 🧭 2026-09-09
Available in: English (current) Français

Keywords

Devstral-2123B24Bcoding modellocal LLMMLXquantizationbatchingperformance review

Summary

The video reviews Devstral-2, a new coding-focused AI model from Mistral, available in two sizes: a 123B dense model and a 24B small model. The presenter tests both on a Mac Studio using MLX quantized versions, measuring token generation speed, memory footprint, and coding capabilities. The 123B model runs at about 6.4 tokens per second on a high-end machine, while the small model reaches 32 tokens per second. Batching enables concurrent generations but reduces individual speed. The video compares Devstral-2 against DeepSeek V3.2 and Kimi K2 Thinking, noting Devstral-2’s score of 72.2% on SWE-Bench (vs DeepSeek’s 73%) and its advantage of requiring under 100 GB of RAM. Demonstrations include generating a 3D universe simulation, a spaceship game, a snooker game, a word processor, and a drawing app, often with functional but basic results. The presenter highlights the dense model’s strong code comprehension and the small model’s practicality for everyday tasks, while cautioning about license restrictions for commercial use and the physical limits of local inference.

165 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial practical value for testers and developers interested in running large coding models locally. It offers real-world numbers (token rates, memory usage) and demonstrates how to batch inferences, manage context windows, and leverage MLX optimizations. The argumentation is experience-based rather than theoretical, with the presenter openly acknowledging weaknesses (slow speed, limited outputs) and comparing multiple models on the same tasks. While the approach is not a rigorous benchmark, the direct hands-on testing and honest commentary give the argumentation a solid empirical foundation, especially for assessing usability on Apple Silicon hardware.

Scientific Rigor, Source Quality, Title Accuracy

The presenter cites Hugging Face model links, the Inferencer app, and companion video reviews of DeepSeek V3.2, Kimi K2, and GPT-OSS, providing transparent access to the exact artifacts tested. The discussion of licenses (Apache 2 for the 24B, custom terms for the 123B) shows attention to legal aspects. The title accurately reflects the content — the video indeed runs and reviews both Devstral-2 models. However, scientific rigor is limited by the uncontrolled environment (single machine, specific prompt examples) and reliance on self-reported benchmarks without independent verification. The presenter’s own admission of complexity and the absence of a formal methodology prevent a higher reliability score.

212 words

Title / Content Match

The title accurately reflects the content: the video runs both Devstral-2 models, tests their coding abilities, and provides a review of their local performance.

Quality & Reliability

6/10

The video offers a hands-on, honest assessment of the two Devstral-2 models on a Mac Studio, with real token rates, memory usage, and code examples. However, the evaluation is anecdotal and not a controlled scientific benchmark, and the creator openly acknowledges limitations such as slow speed and basic outputs. Affiliate links are present but do not distort the technical content.

Key Moments

Cited Sources

  • Devstral-2-123B-Instruct-2512-MLX-6.5bit — Hugging Face model page for the quantized 123B Devstral-2 model tested in the video.
  • Devstral-Small-2-24B-Instruct-2512-MLX-6.5bit — Hugging Face model page for the quantized 24B Devstral Small model tested in the video.
  • Inferencer App — The application used to run the models locally, with version 1.8 including modifications by the presenter.
  • DeepSeek V3.2 Review — Companion video comparing DeepSeek V3.2, which achieved 73% on SWE-Bench, slightly higher than Devstral-2.
  • Kimi K2 Thinking Review — Companion video referencing Kimi K2 Thinking, which scored lower than Devstral-2 in benchmarks.
  • OpenAI GPT-OSS Review — Companion video reviewing OpenAI's open-source model, mentioned for comparison of coding output quality.

Concurring Sources

  • Hugging Face Blog on MLX — Discusses MLX, the library used for optimization, aligning with the video's claims about performance and quantization.

Dissenting Sources

  • Mistral official documentation — The video's claim of 72.2% on SWE-Bench may differ from official benchmarks due to quantization or test setup; official documentation provides authoritative performance figures but is not referenced in the video.

External References

Contribution & Novelties

The video contributes a practical, first-hand evaluation of two new local coding AI models, providing concrete performance metrics and code outputs that are scarce in official announcements. It highlights the trade-offs between large dense models (high intelligence, low speed) and smaller models (speed vs. capability) in a real-world Mac Studio environment, along with tips for batching and memory management.

Pour aller plus loin :

  • MLX - Apple’s machine learning framework — Explains the framework used for running models efficiently on Apple Silicon, central to the video’s demonstrations.
  • Model quantization — The concept behind reducing model size and memory footprint (e.g., Q6, float16) without drastically losing quality, essential to the video’s testing methodology.
  • Mixture of Experts — Architecture used by competitors (e.g., DeepSeek) that trades speed for scale; contextualizes the dense model approach of Devstral-2.
  • SWE-Bench — The benchmark referenced for coding performance; accuracy and relevance in evaluating software engineering agents.

150 words

Radar Profile

The profile shows high 'quantite_information' and 'niveau_technique' but moderate 'qualite_information' and 'fiabilite_globale', indicating a rich but not rigorously controlled review. The presenter's hands-on approach and technical depth are strengths, while the anecdotal nature and lack of formal methodology temper the overall reliability.

Reliability 6/10