Overview
Reasoning models can process information in two stages:- Thinking stage: Internal reasoning process, returned in the
reasoning_contentfield of the response message - Response stage: The final answer, returned in the
contentfield
chat_template_kwargs parameter in your API requests. Reasoning output is billed as completion tokens. The number of tokens spent thinking is reported in usage.completion_tokens_details.reasoning_tokens, the same location OpenAI uses, and is a subset of completion_tokens. The top-level usage.reasoning_tokens field is kept as a deprecated alias.
GLM-5.1
GLM-5.1 supports reasoning through theenable_thinking parameter in chat_template_kwargs. Reasoning is enabled by default for GLM-5.1.
Default Behavior (Reasoning Enabled)
GLM-5.1 uses reasoning by default. You can explicitly enable it by setting"enable_thinking": true, or simply omit the parameter:
Disable Reasoning
To disable reasoning for GLM-5.1, you must explicitly set"enable_thinking": false in the chat_template_kwargs:
Qwen3.5-122B and Qwen3.6-35B
Qwen3.5-122B (Qwen/Qwen3.5-122B-A10B) and Qwen3.6-35B (Qwen/Qwen3.6-35B-A3B-FP8) have reasoning enabled by default. The model will include reasoning_content in responses automatically. To disable reasoning, set "enable_thinking": false in the chat_template_kwargs:
GPT-OSS-120B
GPT-OSS-120B (openai/gpt-oss-120b) always reasons before answering — its reasoning cannot be disabled. Instead, you control how much it thinks with the standard OpenAI reasoning_effort parameter (low, medium, or high; defaults to medium):
Kimi K2.5 and Kimi K2.6
Kimi K2.5 (moonshotai/kimi-k2.5) and Kimi K2.6 (moonshotai/kimi-k2.6) have reasoning enabled by default and include reasoning_content in responses automatically. To disable reasoning, set "thinking": false in the chat_template_kwargs (note the parameter name — thinking, not enable_thinking):
"thinking": false the response contains no reasoning trace and usage.completion_tokens_details.reasoning_tokens drops to 0.
Model-Specific Parameters
Different models use different parameter names for controlling reasoning:
Third-party models served through the gateway (OpenAI
o-series and GPT-5.x, Anthropic Claude, Gemini Pro, …) use their provider’s standard reasoning controls, such as reasoning_effort. The Kimi models are the exception: they do not respond to reasoning_effort — use chat_template_kwargs.thinking as described above.
Reading Reasoning Output
When reasoning is active, responses contain the reasoning trace alongside the answer:reasoning_content— the model’s internal thinking (streamed asdelta.reasoning_contentchunks whenstream: true)usage.completion_tokens_details.reasoning_tokens— how many of the completion tokens were spent thinking. It is present in non-streaming responses and in the final usage chunk of a stream when you setstream_options: {"include_usage": true}. The top-levelusage.reasoning_tokenscarries the same number and is deprecated.
When to Use Reasoning
Reasoning is particularly useful for:- Complex mathematical problems: Multi-step calculations and problem-solving
- Logical reasoning: Tasks requiring step-by-step analysis
- Code generation: Complex programming problems that need careful planning
- Scientific questions: Problems requiring structured thinking
- Simple queries: Straightforward questions with direct answers
- Fast responses: When latency is critical and the problem is simple
- Cost optimization: Reasoning can increase token usage and costs
Best Practices
- Test with and without reasoning: Compare results to determine if reasoning improves output quality for your use case
- Monitor token usage: Reasoning can increase the number of tokens used, affecting costs — watch
usage.reasoning_tokens - Use appropriate models: Not all models support reasoning - check model capabilities before enabling
- Stream responses: When using reasoning, streaming can provide better user experience for longer responses