BEST Local AI for Tool Calls? GPT-OSS, DeepSeek, Kimi K2, MiniMax, GLM Compared

BEST Local AI for Tool Calls? GPT-OSS, DeepSeek, Kimi K2, MiniMax, GLM Compared

🎙 xCreate 👥 26K 📅 November 20, 2025 ⏱ 24 min 👁 6K 📄 expert opinion 🧭 2026-09-09
Available in: English (current) Français

Keywords

tool callinglocal AIGPT-OSSDeepSeekfunction calling

Summary

The video presents a comparative test of prominent open-weight local AI models—GPT-OSS (20B/120B), DeepSeek V3.1, Kimi K2, MiniMax M2, GLM 4.6, Qwen3 (1.7B/235B), and Apertus—focused on their ability to perform tool calls. The host uses the Inferencer app (v1.7) on a Mac to demonstrate real-time tool invocation, including weather retrieval, web search via DuckDuckGo, and webpage content extraction. The test highlights challenges such as rate limiting, hallucinated URLs, and model-specific parsing quirks. Notably, GPT-OSS 120B cleverly adapted by accessing Google News when search was disabled, while others struggled. The video also shows how to create custom tools and modify existing ones, and discusses the XML-based tool call format of MiniMax. The overall conclusion favors MiniMax and GPT-OSS for their smart handling, while smaller models like Qwen3-1.7B show limitations. The presentation is practical and informative, though not scientifically rigorous.

138 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable hands-on insights into the tool calling capabilities of local AI models, which are often underrepresented in benchmarks. The empirical tests, though informal, reveal real-world failure modes like rate limiting and hallucinated URLs, and demonstrate how models adapt. The argumentation is coherent, with clear comparisons and visual evidence. However, the lack of standardized metrics, single-run tests, and potential biases in hardware/software configuration limit the strength of conclusions. The host’s reasoning is transparent, and he acknowledges limitations, but the subjective evaluation of ‘smartness’ is not reproducible.

Scientific Rigor, Source Quality, Title Accuracy

The video references model cards on Hugging Face and the Inferencer app, providing traceability. The testing procedure is not controlled, but the host shows actual outputs and errors, enhancing credibility. The title is apt, as it accurately reflects the content. However, there is no external validation, and the affiliate links in the description introduce potential conflicts of interest. The absence of citations to peer-reviewed benchmarks reduces scientific rigor, but as an expert opinion piece, it offers practical value.

181 words

Title / Content Match

The title accurately describes the content, as the video systematically compares several local AI models on their tool calling capabilities, directly matching the viewer's expectation.

Quality & Reliability

6/10

The video is an informal, hands-on comparative test of local AI models for tool calling. While the demonstrations are real and reproducible, the methodology is not rigorously controlled, lacks statistical validation, and includes potential biases from affiliate marketing. However, the host transparently shows limitations and model behaviors, grounding the assessment in concrete outputs.

Key Moments

Cited Sources

  • GPT-OSS 20B model — One of the tested models for tool calling.
  • GPT-OSS 120B model — One of the tested models for tool calling.
  • DeepSeek V3.1 Terminus MLX — Tested model with prompt-based tool calling.
  • GLM 4.6 MLX — Tested model for tool calling.
  • Kimi K2 Thinking MLX — Tested model known for strong reasoning.
  • MiniMax M2 MLX — Tested model with XML-based tool call format.
  • Qwen3 1.7B — Small model tested for tool calling on resource-constrained devices.
  • Qwen3 235B A22B Instruct — Large Qwen model tested for tool calling.
  • Apertus 70B — Swiss model tested, showing basic tool calling.
  • Inferencer App — Application used to run local models and manage tool calls.

External References

Contribution & Novelties

The video offers a practical comparison of tool calling across multiple open-weight local AI models, revealing that while all can call basic functions, their robustness varies significantly under constraints like rate-limited searches. It demonstrates GPT-OSS 120B’s ability to creatively bypass failed search tools, and highlights the XML parsing issues in MiniMax. The host also shows how to easily create custom tools within Inferencer, enhancing extendibility. This hands-on approach provides insights that benchmarks often miss, particularly regarding real-world failure recovery.

Pour aller plus loin :

155 words

Radar Profile

The radar indicates a high quantity and quality of information, with a moderate level of technical depth and a lower global reliability due to the informal methodology. The video excels in providing practical demonstrations and diverse model comparisons, but the lack of controlled experiments and statistical rigor pulls down the reliability score.

Reliability 6/10