What is Private Inference?
Private inference is a method of running AI models where both your input data and the model’s outputs remain completely hidden from everyone except the user and client even while the computation happens on remote servers you don’t control. Traditional cloud AI services require you to trust that providers won’t access your data. Private inference eliminates this need for trust by using hardware-based security that makes it technically impossible for anyone to see your data, even with physical access to the servers. NEAR AI Cloud’s private inference provides three core guarantees:Complete Privacy
Your prompts, model weights, and outputs are encrypted and isolated in hardware-secured environments. Infrastructure providers, model providers, and NEAR cannot access your data at any point in the process.
Cryptographic Verification
Attestation reports provide cryptographic evidence for a specific runtime and request path. When a response signature is available, you can also verify response integrity.
Production Performance
Hardware-accelerated TEEs with NVIDIA Confidential Computing deliver high-throughput inference with minimal latency overhead, making private inference practical for real-world applications.
How Private Inference Works
Trusted Execution Environment (TEE)
NEAR AI Cloud combines Intel TDX and NVIDIA TEE technologies to create isolated, secure environments for AI computation:- Intel TDX(Trust Domain Extensions) : Creates confidential virtual machines (CVMs) that isolate your AI workloads from the host system, preventing unauthorized access to data in memory.
- NVIDIA TEE : Provides GPU-level isolation for model inference, ensuring model weights and computations remain completely private during processing.
- Cryptographic Attestation : Each TEE environment generates cryptographic proofs of its integrity and configuration, enabling independent verification of the secure execution environment.
Client-Side Encryption
If you’re using a standard OpenAI SDK or curl, your prompts are automatically protected by TLS encryption—no additional setup required. For new integrations, use the NEAR AI Cloud Gateway. Its verification flow can establish the evidence required for a particular request path. Gateway mode routes throughcloud-api.near.ai, which runs in its own TEE before forwarding to the model TEE:
Key insight: TLS protects traffic in transit. When the Gateway TLS attestation flow passes for a particular connection, it binds that connection’s endpoint to verified attestation evidence.
Here’s why this works:
- Standard HTTPS = TLS encryption: When you make API calls using the OpenAI SDK or curl, you’re connecting via HTTPS. This means TLS encryption is applied automatically by your client before any data leaves your machine.
- TLS endpoint verification: Use Gateway TLS attestation when your application needs to verify the endpoint for a particular connection. The same TLS connection must supply both the peer certificate and attestation report.
- Evidence-driven trust: Use attestation and the workload policy appropriate for your application to decide which environment properties are required.
Direct Completions
Self-hosted models also have direct completions endpoints athttps://{slug}.completions.near.ai/v1, which skip the Gateway. They are experimental and have known verification limitations, so we don’t recommend them for new integrations or production verification. See Direct Completions.
The Inference Process
When you make a request to NEAR AI Cloud through the Gateway, your data flows through a secure pipeline designed to maintain privacy at every step. Via Gateway (cloud-api.near.ai):
- Request Initiation: You send chat completion requests via HTTPS to the LLM Gateway. TLS protects data in transit; use NEAR AI Cloud gateway TLS connection binding when your application needs to verify the endpoint for a particular connection.
- Secure Request Routing: The LLM Gateway routes your request to the appropriate Private LLM Node based on the requested model, availability, and load balancing requirements.
- Secure Inference: AI inference computations execute inside the Private LLM Node’s TEE, where all data and model weights are protected by hardware-enforced isolation.
- Attestation Generation Attestation reports provide evidence about the environment and its measured configuration.
- Response Signing: Supported flows can make a response signature available for integrity verification.
- Verifiable Response: Use verification to check the evidence required by your application. A response without an available signature is not response-verified.
Architecture Overview
NEAR AI Cloud operates through a distributed architecture consisting of an LLM Gateway and a network of Private LLM Nodes.Private LLM Nodes
Each Private LLM Node provides secure, isolated AI inference capabilities:- Standardized Hardware: 8x NVIDIA H200 GPUs per node, optimized for high-performance inference
- Intel TDX-enabled CPUs: Enable secure virtualization with hardware-enforced isolation
- Private-ML-SDK: Manages secure model execution, attestation generation, and cryptographic signing
- Health Monitoring: Automated liveness checks and monitoring ensure continuous availability
LLM Gateway
The LLM Gateway serves as the central orchestration layer:- Model Management: Registers and manages available models across the Private LLM Node network
- Request Routing: Intelligently routes requests to appropriate nodes based on model availability and load
- Attestation Verification: Validates and stores TEE attestation reports for audit and verification
- Access Control: Manages API keys, authentication, and usage tracking for billing and monitoring
Security Guarantees
Defense in Depth
NEAR AI Cloud’s private inference implements multiple layers of security to protect your data:- Hardware-Level Isolation : TEEs create isolated execution environments enforced at the hardware level, preventing unauthorized access to memory and computation even from privileged system administrators or cloud providers.
- Secure Communication : All communication between your applications and the LLM infrastructure uses TLS encryption. Use the route-specific verification guide to verify endpoint binding for a particular connection when that property is required.
- Cryptographic Attestation : Every TEE environment generates cryptographic proofs that verify the integrity of the execution environment, allowing you to independently confirm your computations occurred in a genuine, unmodified TEE.
- Result Authentication : When a response signature is available, it can be verified against the attested signer to check response integrity.
Threat Protection
NEAR AI Cloud’s architecture protects against common attack vectors:- Malicious Infrastructure Providers : Hardware-enforced TEE isolation prevents cloud infrastructure providers from accessing your prompts, model weights, or inference results, even with physical access to servers.
- Network-Based Attacks : End-to-end encryption protects your data during transmission, preventing man-in-the-middle attacks and network eavesdropping.
- Model Extraction Attempts : Model weights remain encrypted and isolated within the TEE, making extraction computationally infeasible even for attackers with privileged system access.
Next Steps
Verification
Understand how to verify and validate secure interactions with AI models
NEAR AI Cloud Gateway Verification
Verify gateway, model, TLS, and response evidence for requests sent through the NEAR AI Cloud gateway
E2EE Chat Completions
Add client-side encryption for defense-in-depth protection of your messages