Skip to content

Repository files navigation


What is InferiaLLM

InferiaLLM is a self-hosted operating system for running LLMs in production. It sits between your applications and your AI infrastructure and provides the platform primitives that organizations need but nobody wants to build from scratch:

  • Access control & RBAC — who can use which models, and how much
  • Policy & safety — rate limits, quotas, token budgets, and content scanners (toxicity, PII) enforced per org and per deployment
  • Inference routing — an OpenAI-compatible gateway with failover, load balancing, and backend selection
  • Compute orchestration — provision and manage a fleet of GPU nodes across clouds, on-prem, and decentralized networks
  • Cost controls — per-user quotas, token budgets, and rate limiting
  • Audit logging — every request tracked, every policy decision recorded

These are operating-system responsibilities. InferiaLLM delivers them as a single, cohesive, single-port deployment.

Why this exists: LLMs, inference engines, and GPUs are all available — but they are not operable by an organization on their own. To run AI in production, teams end up stitching together a dozen tools. InferiaLLM consolidates that entire layer.

It is not a model, a runtime, or a training system. It governs how those systems are used.


The Inferia platform

InferiaLLM is the brain of a three-repository platform. This repo is the control + data plane; two sibling projects complete the picture:

Repository Language Role
InferiaLLM (this repo) Python · React Control plane (auth, RBAC, policy, audit, orchestration) + data plane (OpenAI-compatible inference gateway) + admin dashboard. Ships as one single-port app.
inferia-worker Go GPU-node agent. Runs on every compute node (bare metal, self-hosted server, or cloud VM), connects back to the control plane, loads models on demand, and serves inference off its local GPUs.
InferiaAllSpark Compose One-box deployment of the full Inferia platform — shared InferiaAuth identity, InferiaGate, and InferiaLLM behind a single reverse proxy.
flowchart LR
    subgraph CP["InferiaLLM — control + data plane"]
        direction TB
        APP["Single-port unified app<br/>(APP_PORT, default 8000)"]
    end

    subgraph FLEET["Compute fleet"]
        direction TB
        W1["inferia-worker<br/>(GPU node)"]
        W2["inferia-worker<br/>(GPU node)"]
        W3["inferia-worker<br/>(cloud GPU)"]
    end

    DEV["Your apps<br/>(OpenAI-compatible client)"] -->|HTTPS| APP
    ADMIN["Operators"] -->|Dashboard| APP
    APP <-->|WebSocket channel<br/>+ heartbeat| W1
    APP <-->|WebSocket channel| W2
    APP <-->|WebSocket channel| W3

    ALLSPARK["InferiaAllSpark<br/>(full-platform, one box)"] -.packages.-> CP
Loading

Architecture

The whole surface runs as one ASGI app on a single port (APP_PORT, default 8000). Historically there were separate ports for the gateway, inference, and dashboard; these are now collapsed in-process behind mount prefixes:

Mount Sub-app Responsibility
/api API Gateway Control plane — auth, RBAC, policy, audit, provider config, orchestration proxy
/inf Inference Gateway Data plane — OpenAI-compatible chat/completions/embeddings
/v2/* OCI registry mirror Model image/weights mirror (served at root; the OCI spec hard-codes <host>/v2)
/ Dashboard SPA React admin UI (static, with index.html fallback)
flowchart TB
    client["Client app / Dashboard"]

    subgraph UNIFIED["Unified app — one port (APP_PORT)"]
        direction TB
        SPA["/  Dashboard SPA"]
        API["/api  API Gateway<br/>(control plane)"]
        INF["/inf  Inference Gateway<br/>(data plane)"]
        V2["/v2  OCI model mirror"]
    end

    subgraph CONTROL["Control plane internals (api_gateway)"]
        direction LR
        AUTH["Auth & RBAC"]
        POLICY["Policy engine<br/>rate · quota · budget · scanners"]
        AUDIT["Audit log"]
        MGMT["Provider & deployment mgmt"]
    end

    ORCH["Orchestration<br/>compute lifecycle · provisioning · scheduling"]

    subgraph BACKENDS["Inference backends"]
        direction LR
        WK["inferia-worker nodes<br/>(vLLM · Ollama)"]
        EXT["External APIs<br/>(OpenAI · Anthropic · …)"]
    end

    DB[("PostgreSQL")]
    REDIS[("Redis")]

    client -->|HTTPS| UNIFIED
    API --> AUTH --> POLICY --> AUDIT
    API --> MGMT
    INF -->|"/internal policy + routing"| API
    INF --> BACKENDS
    API <-->|gRPC / HTTP| ORCH
    ORCH --> WK
    INF --> WK
    INF --> EXT
    CONTROL --- DB
    CONTROL --- REDIS
    ORCH --- DB
Loading

Data plane (/inf) handles inference traffic. The Inference Gateway normalizes every request to the OpenAI schema, evaluates policy via the control plane, then routes to a backend — a worker node or an external API — and streams the response back.

Control plane (/api) authenticates and authorizes, enforces policy and budgets, records audit, and manages provider credentials and deployments. The orchestration subsystem owns the compute fleet lifecycle.

Everything runs in the one unified app; the internal services (orchestration REST/gRPC, DePIN sidecar) sit on loopback and are remappable for host networking — see Services & ports.


Compute Fleet

InferiaLLM treats compute as a first-class, governed resource. A pool is a logical group of GPU nodes; each node runs the inferia-worker agent. Nodes are either self-hosted (you run the worker on your own hardware) or cloud-provisioned (InferiaLLM stands up the VM for you via Pulumi and cloud-init installs the worker automatically).

Once running, a worker exchanges a short-lived bootstrap token for a long-lived JWT, opens a persistent WebSocket channel to the control plane, and heartbeats live telemetry (CPU, memory, GPU utilization & VRAM, network, disk). The control plane sends LoadModel commands over that channel; the worker pulls the model (through the /v2 mirror) and begins serving OpenAI-compatible inference on its advertised URL.

sequenceDiagram
    actor Op as Operator
    participant CP as InferiaLLM control plane
    participant Cloud as Pulumi / cloud-init
    participant W as inferia-worker

    Op->>CP: Create pool + add node
    alt Cloud node
        CP->>Cloud: Provision GPU VM (Pulumi)
        Cloud->>W: cloud-init runs worker container
    else Self-hosted node
        Op->>W: docker compose up (with bootstrap token)
    end
    W->>CP: Register + exchange bootstrap token → JWT
    W->>CP: Open WebSocket channel
    loop Every few seconds
        W->>CP: Heartbeat (GPU/CPU/mem/net/disk telemetry)
    end
    CP->>W: LoadModel(model)
    W->>CP: Pull weights via /v2 mirror
    W-->>CP: Ready (advertise URL)
    Note over CP,W: Deployment RUNNING — routes inference to this node
Loading

Adding a self-hosted GPU node

On the control-plane host, scaffold a ready-to-run worker deployment:

inferiallm worker compose \
  --pool-id <POOL_UUID> \
  --node-name gpu-1 \
  --advertise-url http://<worker-host>:8080 \
  --out-dir ./inferia-worker-deploy

This writes a .env (with a freshly minted bootstrap token) and a docker-compose.yml. Copy that directory to your GPU host (Docker + NVIDIA Container Toolkit required) and run:

cd inferia-worker-deploy && docker compose up -d

The node appears in the dashboard within seconds and is ready to receive model deployments. See the inferia-worker README for details.

Providers

Provider Type How it connects
Self-hosted worker Bare metal / your VMs inferia-worker registers directly
AWS Cloud GPU (EC2) Provisioned via Pulumi; cloud-init runs the worker
GCP Cloud GPU (Compute Engine) Provisioned via Pulumi
Azure Cloud GPU Provisioned via Pulumi
Nosana Decentralized GPU (DePIN) Native sidecar integration
Akash Decentralized cloud SDL-based deployment
Kubernetes On-prem / managed clusters Direct orchestration

Local inference engines: vLLM · Ollama · TEI · Infinity · Inferia Diffusion (image/video)

External API providers: OpenAI · Anthropic · Cohere · Gemini · Groq · Cerebras · OpenRouter


Request Lifecycle

Every request flows through a governed pipeline before it reaches a model. Requests that fail auth, policy, or safety are rejected before any inference — no GPU time is wasted on unauthorized or unsafe traffic.

sequenceDiagram
    participant App as Client app
    participant INF as Inference Gateway (/inf)
    participant API as API Gateway (/api)
    participant POL as Policy engine
    participant ORCH as Orchestration
    participant BE as Backend (worker / external)
    participant AUD as Audit log

    App->>INF: POST /v1/chat/completions
    INF->>API: Resolve context (auth + deployment)
    API->>POL: Evaluate policy (RBAC, rate, quota, scanners)
    alt Rejected
        POL-->>App: 4xx (blocked — no compute used)
    else Allowed
        POL-->>INF: Pass + routing context
        INF->>ORCH: Select backend
        ORCH-->>INF: Target node / provider
        INF->>BE: Forward request
        BE-->>INF: Stream tokens
        INF-->>App: Streamed response
        INF->>AUD: Log tokens, latency, cost
    end
Loading

Quick Start

Clone the repo and run the setup script — it does everything: generates a complete .env (all secrets included), builds and starts the whole stack (unified app + PostgreSQL + Redis), waits for health, and verifies routing.

git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM
./setup.sh

That's it. Everything is served on one port:

Dashboard, API, and inference: http://localhost:8000 (or whatever you set APP_PORT to) — Dashboard at /, control plane at /api, inference at /inf, model mirror at /v2.

Database migrations run automatically on boot. Log in with the superadmin credentials the script prints (or the ones you pass it).

Requirements: Docker (with the Compose plugin) and python3.

Non-interactive & other flags
# Fully unattended (CI / remote box) — a password is generated if you don't pass one
./setup.sh --yes \
  --public-url https://inferiallm.example.com \
  --superadmin-email admin@example.com

./setup.sh --build      # force an image rebuild
./setup.sh --no-up      # only (re)generate .env, don't touch Docker
./setup.sh --down       # stop the stack
./setup.sh --help       # full list of flags

Deploying models to Kubernetes

The control plane runs in a container, so to deploy models onto a local kind cluster it needs a route to the cluster's network and a kubeconfig whose address resolves from inside a container — the one on your host points at 127.0.0.1, which in there means the container.

kind create cluster          # if you don't have one
./setup.sh --k8s             # add --k8s-cluster NAME if you have several

Without this, Kubernetes deployments fail with Service host/port is not set.

It also sets K8S_SERVICE_TYPE=NodePort and adds the cluster's node addresses to UPSTREAM_ALLOWED_INTERNAL_HOSTS in .env. Without those, deployments reach RUNNING but chat requests fail with Invalid upstream configuration. Run it again if a node's address changes, for example after Docker restarts.

On a GPU machine

deploy/gpu-kind.sh takes a GPU machine from nothing to a cluster models can be deployed on: the NVIDIA runtime for Docker, a kind cluster that can see the GPU, the GPU operator and device plugin, KEDA, Prometheus, setup.sh --k8s, and the vLLM image pre-loaded onto the node.

./deploy/gpu-kind.sh              # everything
./deploy/gpu-kind.sh --no-engine  # skip pre-loading the vLLM image

It expects Linux x86_64 with an NVIDIA GPU, and the driver, Docker and NVIDIA container toolkit present already — an AWS Deep Learning Base AMI (Ubuntu) has all three. It stops with a clear message if any of them is missing. Every step checks before it acts, so it is safe to re-run.

This is one way to get a cluster, not a requirement: the control plane only uses the standard Kubernetes API, so any cluster works. The automation here is kind-specific though — on k3s, EKS or OpenShift, give the app container a kubeconfig for that cluster yourself, and set the Prometheus scrape target in deploy/k8s/prometheus.yaml to an address its pods can reach the control plane on.

The dashboard listens on the machine, so reach it over a tunnel:

ssh -i <key.pem> -L 8000:localhost:8000 ubuntu@<machine>

Self-Hosting with Docker Compose

Prefer to wire it up by hand? setup.sh is just orchestration around a single docker-compose.yml that builds and runs the whole platform (unified app + PostgreSQL + Redis). You can drive it directly:

# 1. Clone
git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM

# 2. Configure
cp .env.example .env

# 3. Generate secrets and paste them into .env
openssl rand -hex 32                 # JWT_SECRET_KEY
openssl rand -hex 32                 # INTERNAL_API_KEY
openssl rand -hex 32                 # LOG_ENCRYPTION_KEY
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"   # SECRET_ENCRYPTION_KEY
# Also set: POSTGRES_PASSWORD, PG_ADMIN_PASSWORD, DATABASE_URL password,
#           SUPERADMIN_EMAIL, SUPERADMIN_PASSWORD

# 4. Launch
docker compose up -d --build

Database migrations run automatically on boot. Log in with the SUPERADMIN_EMAIL / SUPERADMIN_PASSWORD you set.

Deployment (single compose, env-driven ports & SSO)
# The single canonical stack: unified app + Postgres + Redis + ELK
docker compose up -d --build

There is ONE compose file — the root docker-compose.yml. Everything else is env-driven in .env: ports (APP_PORT, and the internal HTTP_PORT / GRPC_PORT / DEPIN_SIDECAR_PORT for host networking), and external SSO (AUTH_PROVIDER + the EXTERNAL_AUTH_* / OAUTH_* block).

See deploy/README.md for reverse-proxy / TLS contracts.

Install from PyPI

pip install inferiallm

# Configure
curl -o .env https://raw.githubusercontent.com/InferiaAI/InferiaLLM/main/.env.example
nano .env                       # set DB, Redis, and secrets

inferiallm init                 # bootstrap database, roles, schemas
inferiallm start                # start the unified app

Requires Python 3.10–3.12 and reachable PostgreSQL + Redis.

Build from source

git clone https://github.com/InferiaAI/InferiaLLM.git && cd InferiaLLM

python3 -m venv .venv && source .venv/bin/activate
pip install -e .

cp .env.example .env            # edit: DB, Redis, secrets
make setup                      # install environment + dependencies (runs setup_project.sh)
inferiallm init --env dev
inferiallm start

Services & ports

The whole web surface runs on a single port, APP_PORT (default 8000) — the only port you publish. A few internal services listen on loopback; each is remappable via its own env var so the app can also run on host networking without colliding with other host services.

Component Default Env var Notes
Unified app 8000 APP_PORT Dashboard /, gateway /api, inference /inf, mirror /v2 — the only published port
Orchestration REST 8080 HTTP_PORT Compute lifecycle / worker channel (loopback; the gateway proxies to it)
Orchestration gRPC 50051 GRPC_PORT Internal RPC (loopback)
DePIN sidecar 3000 DEPIN_SIDECAR_PORT Nosana/Akash coordination (loopback Node service)
PostgreSQL 5432 DATABASE_URL Primary datastore (external)
Redis 6379 REDIS_PORT Rate limiting, pub/sub, streams (external)
inferia-worker 8080 — Inference port advertised by each GPU node (on the worker host)

The in-process clients derive their target from the same env var, so changing e.g. HTTP_PORT moves both the orchestration listener and every caller together. An explicit ORCHESTRATION_URL / ORCHESTRATION_GRPC_ADDR / NOSANA_SIDECAR_URL / AKASH_SIDECAR_URL still wins (split / remote deployments).


Configuration

InferiaLLM is configured via a .env file (start from .env.example). Key variables:

Core & web

Variable Description Default
APP_PORT Single port for the unified app 8000
ENVIRONMENT production or dev production
AUTH_PROVIDER local or inferiaauth/oidc (external SSO) local
ALLOWED_ORIGINS CORS allow-list (comma-separated) —
FORWARDED_ALLOW_IPS Trusted proxy IPs (for correct client IPs behind nginx/Caddy) —

Security

Variable Description
JWT_SECRET_KEY Signs access tokens. Min 32 chars (openssl rand -hex 32).
INTERNAL_API_KEY Authenticates service-to-service calls. Min 32 chars.
SECRET_ENCRYPTION_KEY Fernet key (URL-safe base64, 32 bytes) for encrypting provider credentials. Generate with python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())".
LOG_ENCRYPTION_KEY 32-byte hex key for encrypting sensitive log fields.
SUPERADMIN_EMAIL / SUPERADMIN_PASSWORD Initial admin login.

Fail closed: if INTERNAL_API_KEY is unset, internal endpoints refuse service. Missing config is treated as an error, never a bypass.

Datastores

Variable Description Default
DATABASE_URL PostgreSQL connection. Canonical form is bare postgresql:// (no +asyncpg). postgresql://inferia:…@postgres:5432/inferia
REDIS_HOST / REDIS_PORT Redis connection redis / 6379

Compute & models

Variable Description Default
INFERIA_WORKER_IMAGE Worker container image repository ghcr.io/inferiaai/inferia-worker
INFERIA_WORKER_IMAGE_TAG Worker image tag deployed to provisioned nodes 0.2.7
INFERIA_MODEL_CACHE_DIR Host path for the model cache / mirror /var/lib/inferia/models
INFERIA_MODEL_MIRROR_BASE Public base URL of this control plane (for the /v2 + /hf mirror) —
INFERIA_SSH_AUTHORIZED_KEYS_FILE Public keys baked into provisioned cloud workers ./.ssh/authorized_keys

CLI

inferiallm init                        # Bootstrap database, roles, schemas
inferiallm migrate                     # Apply pending migrations (auto-runs in Docker)
inferiallm start                       # Start the unified app (all services)
inferiallm start api-gateway           # Split mode: control plane only
inferiallm start inference             # Split mode: data plane only
inferiallm start orchestration         # Split mode: orchestration only

# Provider credentials (DB-backed, encrypted)
inferiallm providers list
inferiallm providers add aws  --name prod --type access_key_id --value AKIA...
inferiallm providers remove aws --name prod

# Compute nodes
inferiallm node add worker --name gpu-1 --pool-id <id> --advertise-url http://host:8080
inferiallm worker compose --pool-id <id> --node-name gpu-1 --advertise-url http://host:8080
inferiallm worker token   --pool-id <id> --ttl-hours 1
inferiallm worker list    --pool-id <id>

Tech Stack

Layer Technology
Language Python 3.10–3.12 · Go (worker)
API FastAPI (async), single-port ASGI mounts
Frontend React 19 · Vite · TailwindCSS · Shadcn/UI · TanStack Query
Inter-service gRPC + Protobuf · Redis Streams/Pub-Sub · WebSocket worker channel
Database PostgreSQL 15 (async SQLAlchemy / asyncpg)
Cache / broker Redis 7
Auth Stateless JWT; optional external SSO (InferiaAuth / OIDC)
Encryption Fernet symmetric encryption for credentials at rest
Provisioning Pulumi (AWS/GCP/Azure) + cloud-init
Observability Prometheus-compatible metrics; JSON structured logs to stdout
Vector pgvector / ChromaDB compatible

Deploying the full platform — InferiaAllSpark

To run the entire Inferia platform on a single box — one shared InferiaAuth identity authority consumed by InferiaGate and InferiaLLM, all behind your own reverse proxy — use InferiaAllSpark.

flowchart TB
    user["Users"]
    nginx["Your nginx<br/>(public hostnames + TLS)"]

    subgraph BOX["One box — InferiaAllSpark"]
        direction TB
        auth["InferiaAuth<br/>:3001 — OIDC issuer + console"]
        gate["InferiaGate<br/>:8080 — console (SSO)"]
        llm["InferiaLLM<br/>:8081 — unified app<br/>(/ · /api · /inf · /v2)"]
    end

    user --> nginx
    nginx -->|auth.example.com| auth
    nginx -->|dashboard.example.com| gate
    nginx -->|llm.example.com| llm
    gate -. SSO .-> auth
    llm  -. SSO .-> auth
Loading
git clone https://github.com/InferiaAI/InferiaAllSpark.git && cd InferiaAllSpark
./up.sh                # bring everything up + smoke-check
./up.sh --build        # force-rebuild all images first
./down.sh -v           # stop and wipe data volumes

The LLM host needs WebSocket upgrade, response buffering off, large request bodies, and long timeouts (for SSE token streaming and multi-GB model pulls over /api/hf and /v2). See deploy/README.md and the InferiaAllSpark nginx contract.


Repository layout

InferiaLLM/
├─ src/
│  ├─ unified_web/      # Single-port parent app (mounts /api, /inf, /v2, /)
│  ├─ api_gateway/      # Control plane: auth, rbac, policy, audit, gateway, management
│  ├─ inference/        # Data plane: OpenAI-compatible engines + provider adapters
│  ├─ orchestration/    # Compute lifecycle, node/pool API, worker channel, depin-sidecar
│  ├─ providers/        # Compute provider adapters (aws, gcp, azure, nosana, akash, k8s, worker, pulumi)
│  ├─ cli/              # `inferiallm` command-line entry point
│  ├─ common/           # Shared logging, errors, utilities
│  ├─ infra/            # SQL schemas + migrations
│  └─ dashboard/        # Built React SPA (source in apps/dashboard)
├─ apps/dashboard/      # Dashboard source (React 19 + Vite)
├─ deploy/              # Dockerfile, entrypoint, nginx
├─ docker-compose.yml   # The single canonical deployment (project name: deploy)
└─ .env.example         # Configuration template

Contributing

Contributions are welcome. Each major component carries its own README with architecture context:

Component Responsibility Documentation
API Gateway Control plane: auth, policy, audit README
Inference OpenAI-compatible data plane README
Orchestration Compute lifecycle & worker fleet README
RBAC Identity and access boundaries README
Gateway Secure internal service routing README
Policy Quota, rate, and budget enforcement README
Audit Immutable execution and policy logs README
Deployment Docker (single root compose), reverse-proxy deploy/README

Open an issue to report bugs or request features. See CONTRIBUTING.md for the development workflow.


Own your intelligence.

inferia.ai · X (Twitter) · LinkedIn

InferiaLLM — Copyright © 2026 Inferia AI · Licensed under the Apache License, Version 2.0

Releases

Packages

Contributors

Languages