Skip to main content
Fusion runs a private, server-side deliberation across multiple models and returns one OpenAI-compatible chat completion. Your client sends one request to NEAR AI Cloud; the inference proxy fans out to the configured panel models, optionally asks a judge model for structured guidance, then synthesizes the final answer with the original model. Fusion is available on /v1/chat/completions through the Cloud API gateway. Cloud API remains a pass-through: it bills the single request from the final aggregate usage, which includes panel, judge, and synthesis model tokens.

Request Shapes

NEAR AI Cloud accepts OpenRouter-style Fusion configuration in either a server tool or a plugin entry.
Use tool_choice: "required" when you want Fusion to run immediately. If you omit it, the outer model can decide whether to call Fusion.
Use the top-level request model for the model that should synthesize the final answer. The Fusion parameters.model or plugin model field selects the judge model. NEAR AI Cloud does not route the OpenRouter virtual model alias openrouter/fusion; the Cloud API gateway routes by real model name before the request reaches the inference proxy.

Parameters

Response Metadata

The response remains OpenAI-compatible and may include a top-level nearai_fusion object:
Raw panel answers are not returned by default. Panel failures are reported in metadata when at least one panel succeeds; if every panel fails, the request fails with a Fusion error. Fusion can use NEAR AI’s web_context_search tool inside panel and judge calls. Include {"type": "web_context_search"} in the same tools array and set max_tool_calls in the Fusion configuration to bound inner search calls.

See Also