Your own private AI server
Run chat, vision, speech, image and embedding models on your own hardware, behind the API your software already speaks. Nothing leaves your network.
- OpenAI-compatible API
- Windows & Linux
- Works air-gapped
# Point any OpenAI SDK at your server from openai import OpenAI client = OpenAI( base_url="https://ai.example.internal:12443/v1", api_key="ofai_sk_…") reply = client.chat.completions.create( model="qwen2.5:7b", messages=[{"role": "user", "content": "Summarise this contract."}]) print(reply.choices[0].message.content)
Data stays home
Prompts, documents and answers never leave your machines. Run it on a laptop, a GPU box or a cluster — or with no internet at all.
Drop-in compatible
OpenAI's API by default, plus Ollama's and Anthropic's. Existing tools, SDKs and apps work by changing one URL.
Any model, any GPU
GGUF and Ollama models, downloads from Hugging Face, NVIDIA (CUDA), any GPU through Vulkan, or plain CPU — one worker per GPU.
Ready for IT
API keys and quotas, single sign-on, a signed audit log, metrics and traces, Windows service and Linux packages.
Every kind of model, one server
Each capability runs in its own supervised worker process: a crash restarts the worker, not the server, and requests in flight are accounted for exactly.
Chat & agents
Streaming chat with tool calls, JSON and JSON-schema output, log-probabilities and several choices per request; the Responses API with stored conversations and background runs.
Vision
Send photos, receipts and screenshots with your prompt. Images are decoded on the server itself, and the server never fetches a URL on a caller's behalf.
Embeddings & reranking
Vectors for search and retrieval, and rerankers that put the right documents first — the building blocks of private RAG.
Speech to text
Transcription and translation with timestamps, streamed results, and live transcription over a WebSocket with voice-activity detection and live captions.
Text to speech
Natural voices streamed as they are spoken, in WAV, FLAC or raw PCM.
Image generation
Text to image, image to image and inpainting, with streamed previews while the picture forms, and optional screening of every image before it is returned.
Moderation
Guard models (Llama Guard 3, IBM Granite Guardian) score text — and a vision model screens images — in OpenAI's moderation format.
Files & batches
OpenAI's Batch API for large jobs: durable across restarts, every line answered once.
Scale out
Several GPUs in one machine, or other machines joined as remote workers over mutual TLS. Requests go where the model is already loaded.
Open standards first
OfflinAI Server speaks OpenAI's API as its default contract, so OpenAI's SDKs and the tools built on them work by pointing them at your server. It also answers Ollama's API and Anthropic's Messages API, exports traces over OpenTelemetry, and joins the caller's W3C trace.
/v1/chat/completions,/v1/responses,/v1/embeddings/v1/audio/transcriptions,/v1/audio/speech/v1/images/generations,/v1/images/edits/v1/moderations,/v1/rerank,/v1/batches- Ollama
/api/*and Anthropic/v1/messages
# Transcribe a meeting recording curl https://ai.example.internal:12443/v1/audio/transcriptions \ -H "Authorization: Bearer $OAS_KEY" \ -F [email protected] -F model=whisper-1 \ -F response_format=verbose_json # Embed documents for search curl https://ai.example.internal:12443/v1/embeddings \ -H "Authorization: Bearer $OAS_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"nomic-embed-text","input":["first","second"]}'
Built for the people who run it
A web console in 16 languages, and every control your security team will ask about.
Keys, limits and budgets
API keys scoped to endpoints and models, with request and token limits, daily and monthly quotas, network restrictions and budgets per project. Exact usage from the engines' own counts.
Organisations and single sign-on
Organisations, projects and users with roles; sign-in with Microsoft Entra ID, Google, Okta or any OpenID provider.
Tamper-evident audit log
Every administrative action in a signed, hash-chained log you can verify, forwarded as it happens to your SIEM over syslog or HTTPS.
Content filters
Mask or block personal data — emails, phone numbers, card numbers, IBANs and your own patterns — in requests and replies, and have a guard model read them.
Encrypted by default
HTTPS for other machines with a built-in certificate authority, your own certificates or automatic ones (ACME), and client certificates for mutual TLS.
Observable
Prometheus metrics, OpenTelemetry request traces, daily log files and a live console view of what every worker is doing.
Air-gapped licensing
Activate online, or carry a licence request to a connected machine and back: the server never needs to call home.
Deploy your way
A Windows service or tray app, Debian and RPM packages with hardened systemd units, a 577 MB container image and a Helm chart.
Plans
Every plan runs any model you like. Plans differ in what the server does around them.
| Feature | Free | Pro | Business | Enterprise |
|---|---|---|---|---|
| Models | any | any | any | any |
| Serve other machines (HTTPS) | – | |||
| Engine workers (GPUs, remote nodes) | 1 | unlimited | unlimited | unlimited |
| Active API keys | 2 | 25 | unlimited | unlimited |
| Per-request usage history | 7 days | 90 days | 1 year | as configured |
| Organisations / projects / users | 1 / 1 / 1 | 1 / 5 / 5 | 1 / unlimited / unlimited | unlimited |
| Batches (the Batch API) | – | |||
| Single sign-on | – | – | ||
| Policies, limits, budgets, content filters | – | – | ||
| Audit forwarding (syslog, HTTPS) | – | – | ||
| Request traces (OpenTelemetry) | – | – | ||
| Residency tags, a fleet's audit chain | – | – | – | |
| Automatic (ACME) and client certificates | – | – | – |
Pricing is announced at launch. Request early access to discuss your deployment. Every plan is provided under the OfflinAI Server Licence.
Documentation
Every licensed server carries the full administrator and developer guide for its own version, in the web console under Guide. It covers:
- Installing on Windows, Linux, containers and Kubernetes
- Licensing, including air-gapped activation
- Models, GPUs and remote workers
- Every API with examples, and the differences from OpenAI's service
- Keys, single sign-on, policies, content filters and the audit log
- Operations, monitoring, backup, settings and troubleshooting
Evaluating for your organisation? Request early access and we'll share the guide with your team.
Try it in a container
The Free plan runs from the image on Docker Hub. On a Linux machine with Docker, it serves programs on that machine:
# Start the server (models and settings live in the volume) docker run -d --name oas --network host \ -v oas-data:/data offlinai/offlinai-server # A one-time link to set the console's password docker exec oas offlinai-server admin bootstrap \ --url http://127.0.0.1:12436 # Any OpenAI client: base URL http://127.0.0.1:12436/v1
Image: offlinai/offlinai-server. Serving other machines over HTTPS (port 12443) needs a licence that includes it.
Questions
Request early access
OfflinAI Server is in early access. Tell us where you'd like to run it and we'll be in touch with a licence and the packages for your platform.