InferiaLLM is a self-hosted operating system for running LLMs in production. It sits between your applications and your AI infrastructure and provides the platform primitives that organizations need but nobody wants to build from scratch:
- Access control & RBAC — who can use which models, and how much
- Policy & safety — rate limits, quotas, token budgets, and content scanners (toxicity, PII) enforced per org and per deployment
- Inference routing — an OpenAI-compatible gateway with failover, load balancing, and backend selection
- Compute orchestration — provision and manage a fleet of GPU nodes across clouds, on-prem, and decentralized networks
- Cost controls — per-user quotas, token budgets, and rate limiting
- Audit logging — every request tracked, every policy decision recorded
These are operating-system responsibilities. InferiaLLM delivers them as a single, cohesive, single-port deployment.
Why this exists: LLMs, inference engines, and GPUs are all available — but they are not operable by an organization on their own. To run AI in production, teams end up stitching together a dozen tools. InferiaLLM consolidates that entire layer.
It is not a model, a runtime, or a training system. It governs how those systems are used.
InferiaLLM is the brain of a three-repository platform. This repo is the control + data plane; two sibling projects complete the picture:
| Repository | Language | Role |
|---|---|---|
| InferiaLLM (this repo) | Python · React | Control plane (auth, RBAC, policy, audit, orchestration) + data plane (OpenAI-compatible inference gateway) + admin dashboard. Ships as one single-port app. |
| inferia-worker | Go | GPU-node agent. Runs on every compute node (bare metal, self-hosted server, or cloud VM), connects back to the control plane, loads models on demand, and serves inference off its local GPUs. |
| InferiaAllSpark | Compose | One-box deployment of the full Inferia platform — shared InferiaAuth identity, InferiaGate, and InferiaLLM behind a single reverse proxy. |
flowchart LR
subgraph CP["InferiaLLM — control + data plane"]
direction TB
APP["Single-port unified app<br/>(APP_PORT, default 8000)"]
end
subgraph FLEET["Compute fleet"]
direction TB
W1["inferia-worker<br/>(GPU node)"]
W2["inferia-worker<br/>(GPU node)"]
W3["inferia-worker<br/>(cloud GPU)"]
end
DEV["Your apps<br/>(OpenAI-compatible client)"] -->|HTTPS| APP
ADMIN["Operators"] -->|Dashboard| APP
APP <-->|WebSocket channel<br/>+ heartbeat| W1
APP <-->|WebSocket channel| W2
APP <-->|WebSocket channel| W3
ALLSPARK["InferiaAllSpark<br/>(full-platform, one box)"] -.packages.-> CP
The whole surface runs as one ASGI app on a single port (APP_PORT, default 8000). Historically there were separate ports for the gateway, inference, and dashboard; these are now collapsed in-process behind mount prefixes:
| Mount | Sub-app | Responsibility |
|---|---|---|
/api |
API Gateway | Control plane — auth, RBAC, policy, audit, provider config, orchestration proxy |
/inf |
Inference Gateway | Data plane — OpenAI-compatible chat/completions/embeddings |
/v2/* |
OCI registry mirror | Model image/weights mirror (served at root; the OCI spec hard-codes <host>/v2) |
/ |
Dashboard SPA | React admin UI (static, with index.html fallback) |
flowchart TB
client["Client app / Dashboard"]
subgraph UNIFIED["Unified app — one port (APP_PORT)"]
direction TB
SPA["/ Dashboard SPA"]
API["/api API Gateway<br/>(control plane)"]
INF["/inf Inference Gateway<br/>(data plane)"]
V2["/v2 OCI model mirror"]
end
subgraph CONTROL["Control plane internals (api_gateway)"]
direction LR
AUTH["Auth & RBAC"]
POLICY["Policy engine<br/>rate · quota · budget · scanners"]
AUDIT["Audit log"]
MGMT["Provider & deployment mgmt"]
end
ORCH["Orchestration<br/>compute lifecycle · provisioning · scheduling"]
subgraph BACKENDS["Inference backends"]
direction LR
WK["inferia-worker nodes<br/>(vLLM · Ollama)"]
EXT["External APIs<br/>(OpenAI · Anthropic · …)"]
end
DB[("PostgreSQL")]
REDIS[("Redis")]
client -->|HTTPS| UNIFIED
API --> AUTH --> POLICY --> AUDIT
API --> MGMT
INF -->|"/internal policy + routing"| API
INF --> BACKENDS
API <-->|gRPC / HTTP| ORCH
ORCH --> WK
INF --> WK
INF --> EXT
CONTROL --- DB
CONTROL --- REDIS
ORCH --- DB
Data plane (/inf) handles inference traffic. The Inference Gateway normalizes every request to the OpenAI schema, evaluates policy via the control plane, then routes to a backend — a worker node or an external API — and streams the response back.
Control plane (/api) authenticates and authorizes, enforces policy and budgets, records audit, and manages provider credentials and deployments. The orchestration subsystem owns the compute fleet lifecycle.
Everything runs in the one unified app; the internal services (orchestration REST/gRPC, DePIN sidecar) sit on loopback and are remappable for host networking — see Services & ports.
InferiaLLM treats compute as a first-class, governed resource. A pool is a logical group of GPU nodes; each node runs the inferia-worker agent. Nodes are either self-hosted (you run the worker on your own hardware) or cloud-provisioned (InferiaLLM stands up the VM for you via Pulumi and cloud-init installs the worker automatically).
Once running, a worker exchanges a short-lived bootstrap token for a long-lived JWT, opens a persistent WebSocket channel to the control plane, and heartbeats live telemetry (CPU, memory, GPU utilization & VRAM, network, disk). The control plane sends LoadModel commands over that channel; the worker pulls the model (through the /v2 mirror) and begins serving OpenAI-compatible inference on its advertised URL.
sequenceDiagram
actor Op as Operator
participant CP as InferiaLLM control plane
participant Cloud as Pulumi / cloud-init
participant W as inferia-worker
Op->>CP: Create pool + add node
alt Cloud node
CP->>Cloud: Provision GPU VM (Pulumi)
Cloud->>W: cloud-init runs worker container
else Self-hosted node
Op->>W: docker compose up (with bootstrap token)
end
W->>CP: Register + exchange bootstrap token → JWT
W->>CP: Open WebSocket channel
loop Every few seconds
W->>CP: Heartbeat (GPU/CPU/mem/net/disk telemetry)
end
CP->>W: LoadModel(model)
W->>CP: Pull weights via /v2 mirror
W-->>CP: Ready (advertise URL)
Note over CP,W: Deployment RUNNING — routes inference to this node
On the control-plane host, scaffold a ready-to-run worker deployment:
inferiallm worker compose \
--pool-id <POOL_UUID> \
--node-name gpu-1 \
--advertise-url http://<worker-host>:8080 \
--out-dir ./inferia-worker-deployThis writes a .env (with a freshly minted bootstrap token) and a docker-compose.yml. Copy that directory to your GPU host (Docker + NVIDIA Container Toolkit required) and run:
cd inferia-worker-deploy && docker compose up -dThe node appears in the dashboard within seconds and is ready to receive model deployments. See the inferia-worker README for details.
| Provider | Type | How it connects |
|---|---|---|
| Self-hosted worker | Bare metal / your VMs | inferia-worker registers directly |
| AWS | Cloud GPU (EC2) | Provisioned via Pulumi; cloud-init runs the worker |
| GCP | Cloud GPU (Compute Engine) | Provisioned via Pulumi |
| Azure | Cloud GPU | Provisioned via Pulumi |
| Nosana | Decentralized GPU (DePIN) | Native sidecar integration |
| Akash | Decentralized cloud | SDL-based deployment |
| Kubernetes | On-prem / managed clusters | Direct orchestration |
Local inference engines: vLLM · Ollama · TEI · Infinity · Inferia Diffusion (image/video)
External API providers: OpenAI · Anthropic · Cohere · Gemini · Groq · Cerebras · OpenRouter
Every request flows through a governed pipeline before it reaches a model. Requests that fail auth, policy, or safety are rejected before any inference — no GPU time is wasted on unauthorized or unsafe traffic.
sequenceDiagram
participant App as Client app
participant INF as Inference Gateway (/inf)
participant API as API Gateway (/api)
participant POL as Policy engine
participant ORCH as Orchestration
participant BE as Backend (worker / external)
participant AUD as Audit log
App->>INF: POST /v1/chat/completions
INF->>API: Resolve context (auth + deployment)
API->>POL: Evaluate policy (RBAC, rate, quota, scanners)
alt Rejected
POL-->>App: 4xx (blocked — no compute used)
else Allowed
POL-->>INF: Pass + routing context
INF->>ORCH: Select backend
ORCH-->>INF: Target node / provider
INF->>BE: Forward request
BE-->>INF: Stream tokens
INF-->>App: Streamed response
INF->>AUD: Log tokens, latency, cost
end
Clone the repo and run the setup script — it does everything: generates a complete .env (all secrets included), builds and starts the whole stack (unified app + PostgreSQL + Redis), waits for health, and verifies routing.
git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM
./setup.shThat's it. Everything is served on one port:
Dashboard, API, and inference:
http://localhost:8000(or whatever you setAPP_PORTto) — Dashboard at/, control plane at/api, inference at/inf, model mirror at/v2.
Database migrations run automatically on boot. Log in with the superadmin credentials the script prints (or the ones you pass it).
Requirements: Docker (with the Compose plugin) and
python3.
Non-interactive & other flags
# Fully unattended (CI / remote box) — a password is generated if you don't pass one
./setup.sh --yes \
--public-url https://inferiallm.example.com \
--superadmin-email admin@example.com
./setup.sh --build # force an image rebuild
./setup.sh --no-up # only (re)generate .env, don't touch Docker
./setup.sh --down # stop the stack
./setup.sh --help # full list of flagsThe control plane runs in a container, so to deploy models onto a local
kind cluster it needs a route to the cluster's
network and a kubeconfig whose address resolves from inside a container — the
one on your host points at 127.0.0.1, which in there means the container.
kind create cluster # if you don't have one
./setup.sh --k8s # add --k8s-cluster NAME if you have severalWithout this, Kubernetes deployments fail with Service host/port is not set.
It also sets K8S_SERVICE_TYPE=NodePort and adds the cluster's node addresses
to UPSTREAM_ALLOWED_INTERNAL_HOSTS in .env. Without those, deployments
reach RUNNING but chat requests fail with Invalid upstream configuration.
Run it again if a node's address changes, for example after Docker restarts.
deploy/gpu-kind.sh takes a GPU machine from nothing to a cluster models can
be deployed on: the NVIDIA runtime for Docker, a kind cluster that can see the
GPU, the GPU operator and device plugin, KEDA, Prometheus, setup.sh --k8s,
and the vLLM image pre-loaded onto the node.
./deploy/gpu-kind.sh # everything
./deploy/gpu-kind.sh --no-engine # skip pre-loading the vLLM imageIt expects Linux x86_64 with an NVIDIA GPU, and the driver, Docker and NVIDIA container toolkit present already — an AWS Deep Learning Base AMI (Ubuntu) has all three. It stops with a clear message if any of them is missing. Every step checks before it acts, so it is safe to re-run.
This is one way to get a cluster, not a requirement: the control plane only
uses the standard Kubernetes API, so any cluster works. The automation here is
kind-specific though — on k3s, EKS or OpenShift, give the app container a
kubeconfig for that cluster yourself, and set the Prometheus scrape target in
deploy/k8s/prometheus.yaml to an address its pods can reach the control
plane on.
The dashboard listens on the machine, so reach it over a tunnel:
ssh -i <key.pem> -L 8000:localhost:8000 ubuntu@<machine>Prefer to wire it up by hand? setup.sh is just orchestration around a single docker-compose.yml that builds and runs the whole platform (unified app + PostgreSQL + Redis). You can drive it directly:
# 1. Clone
git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM
# 2. Configure
cp .env.example .env
# 3. Generate secrets and paste them into .env
openssl rand -hex 32 # JWT_SECRET_KEY
openssl rand -hex 32 # INTERNAL_API_KEY
openssl rand -hex 32 # LOG_ENCRYPTION_KEY
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())" # SECRET_ENCRYPTION_KEY
# Also set: POSTGRES_PASSWORD, PG_ADMIN_PASSWORD, DATABASE_URL password,
# SUPERADMIN_EMAIL, SUPERADMIN_PASSWORD
# 4. Launch
docker compose up -d --buildDatabase migrations run automatically on boot. Log in with the SUPERADMIN_EMAIL / SUPERADMIN_PASSWORD you set.
Deployment (single compose, env-driven ports & SSO)
# The single canonical stack: unified app + Postgres + Redis + ELK
docker compose up -d --buildThere is ONE compose file — the root docker-compose.yml. Everything else is env-driven in .env: ports (APP_PORT, and the internal HTTP_PORT / GRPC_PORT / DEPIN_SIDECAR_PORT for host networking), and external SSO (AUTH_PROVIDER + the EXTERNAL_AUTH_* / OAUTH_* block).
See deploy/README.md for reverse-proxy / TLS contracts.
pip install inferiallm
# Configure
curl -o .env https://raw.githubusercontent.com/InferiaAI/InferiaLLM/main/.env.example
nano .env # set DB, Redis, and secrets
inferiallm init # bootstrap database, roles, schemas
inferiallm start # start the unified appRequires Python 3.10–3.12 and reachable PostgreSQL + Redis.
git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM
python3 -m venv .venv && source .venv/bin/activate
pip install -e .
cp .env.example .env # edit: DB, Redis, secrets
make setup # install environment + dependencies (runs setup_project.sh)
inferiallm init --env dev
inferiallm startThe whole web surface runs on a single port, APP_PORT (default 8000) — the only port you publish. A few internal services listen on loopback; each is remappable via its own env var so the app can also run on host networking without colliding with other host services.
| Component | Default | Env var | Notes |
|---|---|---|---|
| Unified app | 8000 |
APP_PORT |
Dashboard /, gateway /api, inference /inf, mirror /v2 — the only published port |
| Orchestration REST | 8080 |
HTTP_PORT |
Compute lifecycle / worker channel (loopback; the gateway proxies to it) |
| Orchestration gRPC | 50051 |
GRPC_PORT |
Internal RPC (loopback) |
| DePIN sidecar | 3000 |
DEPIN_SIDECAR_PORT |
Nosana/Akash coordination (loopback Node service) |
| PostgreSQL | 5432 |
DATABASE_URL |
Primary datastore (external) |
| Redis | 6379 |
REDIS_PORT |
Rate limiting, pub/sub, streams (external) |
| inferia-worker | 8080 |
— | Inference port advertised by each GPU node (on the worker host) |
The in-process clients derive their target from the same env var, so changing e.g. HTTP_PORT moves both the orchestration listener and every caller together. An explicit ORCHESTRATION_URL / ORCHESTRATION_GRPC_ADDR / NOSANA_SIDECAR_URL / AKASH_SIDECAR_URL still wins (split / remote deployments).
InferiaLLM is configured via a .env file (start from .env.example). Key variables:
| Variable | Description | Default |
|---|---|---|
APP_PORT |
Single port for the unified app | 8000 |
ENVIRONMENT |
production or dev |
production |
AUTH_PROVIDER |
local or inferiaauth/oidc (external SSO) |
local |
ALLOWED_ORIGINS |
CORS allow-list (comma-separated) | — |
FORWARDED_ALLOW_IPS |
Trusted proxy IPs (for correct client IPs behind nginx/Caddy) | — |
| Variable | Description |
|---|---|
JWT_SECRET_KEY |
Signs access tokens. Min 32 chars (openssl rand -hex 32). |
INTERNAL_API_KEY |
Authenticates service-to-service calls. Min 32 chars. |
SECRET_ENCRYPTION_KEY |
Fernet key (URL-safe base64, 32 bytes) for encrypting provider credentials. Generate with python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())". |
LOG_ENCRYPTION_KEY |
32-byte hex key for encrypting sensitive log fields. |
SUPERADMIN_EMAIL / SUPERADMIN_PASSWORD |
Initial admin login. |
Fail closed: if
INTERNAL_API_KEYis unset, internal endpoints refuse service. Missing config is treated as an error, never a bypass.
| Variable | Description | Default |
|---|---|---|
DATABASE_URL |
PostgreSQL connection. Canonical form is bare postgresql:// (no +asyncpg). |
postgresql://inferia:…@postgres:5432/inferia |
REDIS_HOST / REDIS_PORT |
Redis connection | redis / 6379 |
| Variable | Description | Default |
|---|---|---|
INFERIA_WORKER_IMAGE |
Worker container image repository | ghcr.io/inferiaai/inferia-worker |
INFERIA_WORKER_IMAGE_TAG |
Worker image tag deployed to provisioned nodes | 0.2.7 |
INFERIA_MODEL_CACHE_DIR |
Host path for the model cache / mirror | /var/lib/inferia/models |
INFERIA_MODEL_MIRROR_BASE |
Public base URL of this control plane (for the /v2 + /hf mirror) |
— |
INFERIA_SSH_AUTHORIZED_KEYS_FILE |
Public keys baked into provisioned cloud workers | ./.ssh/authorized_keys |
inferiallm init # Bootstrap database, roles, schemas
inferiallm migrate # Apply pending migrations (auto-runs in Docker)
inferiallm start # Start the unified app (all services)
inferiallm start api-gateway # Split mode: control plane only
inferiallm start inference # Split mode: data plane only
inferiallm start orchestration # Split mode: orchestration only
# Provider credentials (DB-backed, encrypted)
inferiallm providers list
inferiallm providers add aws --name prod --type access_key_id --value AKIA...
inferiallm providers remove aws --name prod
# Compute nodes
inferiallm node add worker --name gpu-1 --pool-id <id> --advertise-url http://host:8080
inferiallm worker compose --pool-id <id> --node-name gpu-1 --advertise-url http://host:8080
inferiallm worker token --pool-id <id> --ttl-hours 1
inferiallm worker list --pool-id <id>| Layer | Technology |
|---|---|
| Language | Python 3.10–3.12 · Go (worker) |
| API | FastAPI (async), single-port ASGI mounts |
| Frontend | React 19 · Vite · TailwindCSS · Shadcn/UI · TanStack Query |
| Inter-service | gRPC + Protobuf · Redis Streams/Pub-Sub · WebSocket worker channel |
| Database | PostgreSQL 15 (async SQLAlchemy / asyncpg) |
| Cache / broker | Redis 7 |
| Auth | Stateless JWT; optional external SSO (InferiaAuth / OIDC) |
| Encryption | Fernet symmetric encryption for credentials at rest |
| Provisioning | Pulumi (AWS/GCP/Azure) + cloud-init |
| Observability | Prometheus-compatible metrics; JSON structured logs to stdout |
| Vector | pgvector / ChromaDB compatible |
To run the entire Inferia platform on a single box — one shared InferiaAuth identity authority consumed by InferiaGate and InferiaLLM, all behind your own reverse proxy — use InferiaAllSpark.
flowchart TB
user["Users"]
nginx["Your nginx<br/>(public hostnames + TLS)"]
subgraph BOX["One box — InferiaAllSpark"]
direction TB
auth["InferiaAuth<br/>:3001 — OIDC issuer + console"]
gate["InferiaGate<br/>:8080 — console (SSO)"]
llm["InferiaLLM<br/>:8081 — unified app<br/>(/ · /api · /inf · /v2)"]
end
user --> nginx
nginx -->|auth.example.com| auth
nginx -->|dashboard.example.com| gate
nginx -->|llm.example.com| llm
gate -. SSO .-> auth
llm -. SSO .-> auth
git clone https://github.com/InferiaAI/InferiaAllSpark.git && cd InferiaAllSpark
./up.sh # bring everything up + smoke-check
./up.sh --build # force-rebuild all images first
./down.sh -v # stop and wipe data volumesThe LLM host needs WebSocket upgrade, response buffering off, large request bodies, and long timeouts (for SSE token streaming and multi-GB model pulls over /api/hf and /v2). See deploy/README.md and the InferiaAllSpark nginx contract.
InferiaLLM/
├─ src/
│ ├─ unified_web/ # Single-port parent app (mounts /api, /inf, /v2, /)
│ ├─ api_gateway/ # Control plane: auth, rbac, policy, audit, gateway, management
│ ├─ inference/ # Data plane: OpenAI-compatible engines + provider adapters
│ ├─ orchestration/ # Compute lifecycle, node/pool API, worker channel, depin-sidecar
│ ├─ providers/ # Compute provider adapters (aws, gcp, azure, nosana, akash, k8s, worker, pulumi)
│ ├─ cli/ # `inferiallm` command-line entry point
│ ├─ common/ # Shared logging, errors, utilities
│ ├─ infra/ # SQL schemas + migrations
│ └─ dashboard/ # Built React SPA (source in apps/dashboard)
├─ apps/dashboard/ # Dashboard source (React 19 + Vite)
├─ deploy/ # Dockerfile, entrypoint, nginx
├─ docker-compose.yml # The single canonical deployment (project name: deploy)
└─ .env.example # Configuration template
Contributions are welcome. Each major component carries its own README with architecture context:
| Component | Responsibility | Documentation |
|---|---|---|
| API Gateway | Control plane: auth, policy, audit | README |
| Inference | OpenAI-compatible data plane | README |
| Orchestration | Compute lifecycle & worker fleet | README |
| RBAC | Identity and access boundaries | README |
| Gateway | Secure internal service routing | README |
| Policy | Quota, rate, and budget enforcement | README |
| Audit | Immutable execution and policy logs | README |
| Deployment | Docker (single root compose), reverse-proxy | deploy/README |
Open an issue to report bugs or request features. See CONTRIBUTING.md for the development workflow.
Own your intelligence.
inferia.ai · X (Twitter) · LinkedIn
InferiaLLM — Copyright © 2026 Inferia AI · Licensed under the Apache License, Version 2.0