Early access

Your own private AI server

Run chat, vision, speech, image and embedding models on your own hardware, behind the API your software already speaks. Nothing leaves your network.

  • OpenAI-compatible API
  • Windows & Linux
  • Works air-gapped
# Point any OpenAI SDK at your server
from openai import OpenAI

client = OpenAI(
    base_url="https://ai.example.internal:12443/v1",
    api_key="ofai_sk_…")

reply = client.chat.completions.create(
    model="qwen2.5:7b",
    messages=[{"role": "user",
               "content": "Summarise this contract."}])
print(reply.choices[0].message.content)
Data stays home

Prompts, documents and answers never leave your machines. Run it on a laptop, a GPU box or a cluster — or with no internet at all.

Drop-in compatible

OpenAI's API by default, plus Ollama's and Anthropic's. Existing tools, SDKs and apps work by changing one URL.

Any model, any GPU

GGUF and Ollama models, downloads from Hugging Face, NVIDIA (CUDA), any GPU through Vulkan, or plain CPU — one worker per GPU.

Ready for IT

API keys and quotas, single sign-on, a signed audit log, metrics and traces, Windows service and Linux packages.

Every kind of model, one server

Each capability runs in its own supervised worker process: a crash restarts the worker, not the server, and requests in flight are accounted for exactly.

Chat & agents

Streaming chat with tool calls, JSON and JSON-schema output, log-probabilities and several choices per request; the Responses API with stored conversations and background runs.

Vision

Send photos, receipts and screenshots with your prompt. Images are decoded on the server itself, and the server never fetches a URL on a caller's behalf.

Embeddings & reranking

Vectors for search and retrieval, and rerankers that put the right documents first — the building blocks of private RAG.

Speech to text

Transcription and translation with timestamps, streamed results, and live transcription over a WebSocket with voice-activity detection and live captions.

Text to speech

Natural voices streamed as they are spoken, in WAV, FLAC or raw PCM.

Image generation

Text to image, image to image and inpainting, with streamed previews while the picture forms, and optional screening of every image before it is returned.

Moderation

Guard models (Llama Guard 3, IBM Granite Guardian) score text — and a vision model screens images — in OpenAI's moderation format.

Files & batches

OpenAI's Batch API for large jobs: durable across restarts, every line answered once.

Scale out

Several GPUs in one machine, or other machines joined as remote workers over mutual TLS. Requests go where the model is already loaded.

Open standards first

OfflinAI Server speaks OpenAI's API as its default contract, so OpenAI's SDKs and the tools built on them work by pointing them at your server. It also answers Ollama's API and Anthropic's Messages API, exports traces over OpenTelemetry, and joins the caller's W3C trace.

  • /v1/chat/completions, /v1/responses, /v1/embeddings
  • /v1/audio/transcriptions, /v1/audio/speech
  • /v1/images/generations, /v1/images/edits
  • /v1/moderations, /v1/rerank, /v1/batches
  • Ollama /api/* and Anthropic /v1/messages
# Transcribe a meeting recording
curl https://ai.example.internal:12443/v1/audio/transcriptions \
  -H "Authorization: Bearer $OAS_KEY" \
  -F [email protected] -F model=whisper-1 \
  -F response_format=verbose_json

# Embed documents for search
curl https://ai.example.internal:12443/v1/embeddings \
  -H "Authorization: Bearer $OAS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nomic-embed-text","input":["first","second"]}'

Built for the people who run it

A web console in 16 languages, and every control your security team will ask about.

Keys, limits and budgets

API keys scoped to endpoints and models, with request and token limits, daily and monthly quotas, network restrictions and budgets per project. Exact usage from the engines' own counts.

Organisations and single sign-on

Organisations, projects and users with roles; sign-in with Microsoft Entra ID, Google, Okta or any OpenID provider.

Tamper-evident audit log

Every administrative action in a signed, hash-chained log you can verify, forwarded as it happens to your SIEM over syslog or HTTPS.

Content filters

Mask or block personal data — emails, phone numbers, card numbers, IBANs and your own patterns — in requests and replies, and have a guard model read them.

Encrypted by default

HTTPS for other machines with a built-in certificate authority, your own certificates or automatic ones (ACME), and client certificates for mutual TLS.

Observable

Prometheus metrics, OpenTelemetry request traces, daily log files and a live console view of what every worker is doing.

Air-gapped licensing

Activate online, or carry a licence request to a connected machine and back: the server never needs to call home.

Deploy your way

A Windows service or tray app, Debian and RPM packages with hardened systemd units, a 577 MB container image and a Helm chart.

Plans

Every plan runs any model you like. Plans differ in what the server does around them.

FeatureFreeProBusinessEnterprise
Modelsanyanyanyany
Serve other machines (HTTPS)–
Engine workers (GPUs, remote nodes)1unlimitedunlimitedunlimited
Active API keys225unlimitedunlimited
Per-request usage history7 days90 days1 yearas configured
Organisations / projects / users1 / 1 / 11 / 5 / 51 / unlimited / unlimitedunlimited
Batches (the Batch API)–
Single sign-on––
Policies, limits, budgets, content filters––
Audit forwarding (syslog, HTTPS)––
Request traces (OpenTelemetry)––
Residency tags, a fleet's audit chain–––
Automatic (ACME) and client certificates–––

Pricing is announced at launch. Request early access to discuss your deployment. Every plan is provided under the OfflinAI Server Licence.

Documentation

Every licensed server carries the full administrator and developer guide for its own version, in the web console under Guide. It covers:

  • Installing on Windows, Linux, containers and Kubernetes
  • Licensing, including air-gapped activation
  • Models, GPUs and remote workers
  • Every API with examples, and the differences from OpenAI's service
  • Keys, single sign-on, policies, content filters and the audit log
  • Operations, monitoring, backup, settings and troubleshooting

Evaluating for your organisation? Request early access and we'll share the guide with your team.

Try it in a container

The Free plan runs from the image on Docker Hub. On a Linux machine with Docker, it serves programs on that machine:

# Start the server (models and settings live in the volume)
docker run -d --name oas --network host \
  -v oas-data:/data offlinai/offlinai-server

# A one-time link to set the console's password
docker exec oas offlinai-server admin bootstrap \
  --url http://127.0.0.1:12436

# Any OpenAI client: base URL http://127.0.0.1:12436/v1

Image: offlinai/offlinai-server. Serving other machines over HTTPS (port 12443) needs a licence that includes it.

Questions

No. Once your models are downloaded the server runs entirely offline. Licences can be activated online, or offline for air-gapped sites: you carry a small request file to any connected machine and bring the signed licence back.

Any GGUF model, models from an existing Ollama store, and the server's own catalog of downloads from Hugging Face — chat and vision models, embedding models, rerankers, guard models, Whisper speech models, text-to-speech and Stable Diffusion-family image models. Every plan runs any model.

A 64-bit Windows or Linux machine. NVIDIA GPUs run through CUDA and other GPUs through Vulkan; without a GPU the server runs on the CPU. Memory decides which models fit: small models run on a laptop, large ones want a GPU with plenty of memory or several GPUs.

Yes — set the SDK's base URL to your server and use one of its API keys. The differences from OpenAI's own service are listed in the documentation: for example, speech comes back as WAV, FLAC or PCM rather than MP3, and images as base64 rather than links.

Not yet. OfflinAI Server runs on Windows and Linux, including in containers and on Kubernetes.

Request early access

OfflinAI Server is in early access. Tell us where you'd like to run it and we'll be in touch with a licence and the packages for your platform.

We'll only use your email to reply. No spam, ever.