Skip to main content
Some models in NEAR AI Cloud support advanced reasoning capabilities that allow them to “think” through problems before providing a final answer. This feature can improve the quality of responses for complex tasks that require step-by-step reasoning.

Overview

Reasoning models can process information in two stages:
  1. Thinking stage: Internal reasoning process, returned in the reasoning_content field of the response message
  2. Response stage: The final answer, returned in the content field
For most TEE-hosted models you can control whether the model uses reasoning by configuring the chat_template_kwargs parameter in your API requests. Reasoning output is billed as completion tokens. The number of tokens spent thinking is reported in usage.completion_tokens_details.reasoning_tokens, the same location OpenAI uses, and is a subset of completion_tokens. The top-level usage.reasoning_tokens field is kept as a deprecated alias.
Which models support reasoning?Check the supported_features array in the /v1/models response — reasoning-capable models list "reasoning".Known gap: Qwen/Qwen3.5-122B-A10B (cloud-api#703), moonshotai/kimi-k2.5, and moonshotai/kimi-k2.6 currently reason by default even though their metadata does not yet list the "reasoning" flag.

GLM-5.1

GLM-5.1 supports reasoning through the enable_thinking parameter in chat_template_kwargs. Reasoning is enabled by default for GLM-5.1.

Default Behavior (Reasoning Enabled)

GLM-5.1 uses reasoning by default. You can explicitly enable it by setting "enable_thinking": true, or simply omit the parameter:

Disable Reasoning

To disable reasoning for GLM-5.1, you must explicitly set "enable_thinking": false in the chat_template_kwargs:

Qwen3.5-122B and Qwen3.6-35B

Qwen3.5-122B (Qwen/Qwen3.5-122B-A10B) and Qwen3.6-35B (Qwen/Qwen3.6-35B-A3B-FP8) have reasoning enabled by default. The model will include reasoning_content in responses automatically. To disable reasoning, set "enable_thinking": false in the chat_template_kwargs:

GPT-OSS-120B

GPT-OSS-120B (openai/gpt-oss-120b) always reasons before answering — its reasoning cannot be disabled. Instead, you control how much it thinks with the standard OpenAI reasoning_effort parameter (low, medium, or high; defaults to medium):

Kimi K2.5 and Kimi K2.6

Kimi K2.5 (moonshotai/kimi-k2.5) and Kimi K2.6 (moonshotai/kimi-k2.6) have reasoning enabled by default and include reasoning_content in responses automatically. To disable reasoning, set "thinking": false in the chat_template_kwargs (note the parameter name — thinking, not enable_thinking):
With "thinking": false the response contains no reasoning trace and usage.completion_tokens_details.reasoning_tokens drops to 0.
reasoning_effort is not supported on KimiThe Kimi models ignore reasoning_effort (and OpenRouter-style reasoning request objects): requests that include them succeed, but the model reasons exactly as it would without them. Reasoning on Kimi K2.5 and K2.6 is on/off only — there are no effort levels or thinking budgets. Use chat_template_kwargs.thinking instead.

Model-Specific Parameters

Different models use different parameter names for controlling reasoning:
Always check the model documentation or use the model’s specific parameter name. Using the wrong parameter name will not enable reasoning.
Third-party models served through the gateway (OpenAI o-series and GPT-5.x, Anthropic Claude, Gemini Pro, …) use their provider’s standard reasoning controls, such as reasoning_effort. The Kimi models are the exception: they do not respond to reasoning_effort — use chat_template_kwargs.thinking as described above.

Reading Reasoning Output

When reasoning is active, responses contain the reasoning trace alongside the answer:
  • reasoning_content — the model’s internal thinking (streamed as delta.reasoning_content chunks when stream: true)
  • usage.completion_tokens_details.reasoning_tokens — how many of the completion tokens were spent thinking. It is present in non-streaming responses and in the final usage chunk of a stream when you set stream_options: {"include_usage": true}. The top-level usage.reasoning_tokens carries the same number and is deprecated.

When to Use Reasoning

Reasoning is particularly useful for:
  • Complex mathematical problems: Multi-step calculations and problem-solving
  • Logical reasoning: Tasks requiring step-by-step analysis
  • Code generation: Complex programming problems that need careful planning
  • Scientific questions: Problems requiring structured thinking
Reasoning may not be necessary for:
  • Simple queries: Straightforward questions with direct answers
  • Fast responses: When latency is critical and the problem is simple
  • Cost optimization: Reasoning can increase token usage and costs

Best Practices

  1. Test with and without reasoning: Compare results to determine if reasoning improves output quality for your use case
  2. Monitor token usage: Reasoning can increase the number of tokens used, affecting costs — watch usage.reasoning_tokens
  3. Use appropriate models: Not all models support reasoning - check model capabilities before enabling
  4. Stream responses: When using reasoning, streaming can provide better user experience for longer responses