Inference control-plane landscape
Research date: 2026-09-08
inferctl source baseline: main, commit 86ef4c2
External-source review date: 2026-09-07, unless a row gives another date
Conclusion
Inference runtimes execute models. Request gateways receive and control live inference requests. inferctl is different: it is an out-of-band local inference control plane. It inspects configured backends, gathers evidence, plans named task routes, and tests readiness without sending an inference prompt. It does not proxy inference traffic, execute prompts, retry requests, load models, or change backend state.
This is a living index. It records the comparison research that is complete and the research that is still in progress. It does not make a performance, replacement, or universal-backend claim.
The three tool classes
| Class | Main job | Inference request path | Example role with inferctl |
|---|---|---|---|
| Inference runtime | Loads and executes one or more models. | Owns model execution. | inferctl can inspect the configured runtime before a client sends a request. |
| Request gateway | Receives, proxies, routes, retries, or observes live requests. | Is in the request path. | A caller can use inferctl for local route evidence before it calls the gateway. |
| Out-of-band control plane | Inspects state and prepares a route without carrying traffic. | Is outside the request path. | inferctl selects a configured route and returns readiness evidence. |
The categories can overlap in one product. For example, a server can provide health APIs and still be a runtime or gateway when it receives inference requests. The deciding question here is whether the tool carries the live request.
Comparison index
| Candidate | Class | Overlap with inferctl | Main difference | Comparison status |
|---|---|---|---|---|
| OpenClaw Infer | Provider inference command | Model selection and request preparation | Sends provider requests; inferctl does not. | Complete |
| InferFlux | Inference server and gateway | Backend health, model state, and routing | Serves live requests; inferctl can inspect it as a backend. | Complete |
| LocalAI | Local runtime with operations features | Discovery, configuration, resource, and worker evidence | Serves requests and can change runtime state; inferctl is read-only across configured backends. | Complete |
| llama-swap | Local lifecycle proxy | Compatible-server health, model routing, and profiles | Proxies requests and manages server lifecycle; inferctl does neither. | Complete |
| LiteLLM Proxy | Multi-provider request gateway | Provider selection, routing, fallback, policy, and visibility | Controls live requests; inferctl makes a route decision before a request. | Complete |
| llama.cpp server router | Local server and model router | Model-name routing and load state | Owns a server process and its model subprocesses; inferctl inspects independent backends. | Complete |
| Otari | Inference gateway with control-plane features | Routing, credentials, budgets, usage, and policy | Includes a gateway request path; inferctl does not carry traffic. | Complete |
| Bifrost | Multi-provider request gateway | Routing, failover, load balancing, and governance | Provides data-plane gateway control rather than local installed and loaded model evidence. | Complete |
| NVIDIA Dynamo | Distributed inference platform | Worker health, routing, canaries, and capacity control | Operates distributed serving; inferctl has local, out-of-band scope. | Complete |
| llm-d Router and KServe | Kubernetes routing and serving stack | Model pools, deployment state, and request routing | Operates Kubernetes model servers and request paths; inferctl does not. | Complete |
| Ray Serve LLM | Distributed LLM serving system | Deployment, health, multi-model serving, and routing | Deploys and serves traffic; inferctl does not deploy or serve models. | Complete |
| SGLang Model Gateway | Model gateway | Model routing, health checks, retry, and circuit control | Handles request-time faults for SGLang deployments; inferctl checks readiness before submission. | Complete |
| Ollama API | Local inference runtime | Loaded-model list, model metadata, and runtime state | A backend that inferctl can inspect, not a separate control plane. | Integration context |
| vLLM | Inference runtime | Health, model inventory, loading, and runtime metrics | A backend and possible integration target, not a direct substitute. | Integration context |
The classes, overlaps, and differences in the table are a first-pass research summary. The sources in the candidate names were reviewed on 2026-09-07. A completed article pins its own project revision or source-inspection date.
Product position
inferctl is a local control-plane CLI. It can inspect configured inference backends, report backend and model evidence where a backend exposes it, select a named route, and run a bounded readiness check without a model prompt. Its machine-readable commands support scripts and agents.
inferctl does not send inference or model-lifecycle requests. The client that uses its result owns authentication, streaming, retries, timeouts, request execution, and lifecycle actions. A successful readiness result is evidence for a selected route. It is not proof of model quality or a successful model response. See the pinned preparation contract and agent guide.
Composition
inferctl can compose with both runtimes and gateways. A typical workflow is:
- Configure the runtime or gateway as an inferctl backend.
- Run
inferctl route <task> --jsonorinferctl preflight <task> --json. - Let the caller stop, choose a permitted fallback, or use the returned handoff.
- Send the inference request directly to the selected runtime or gateway.
The runtime or gateway remains responsible for request execution and any request-time routing, retries, model lifecycle, policy, or observability. inferctl remains outside that request path.
Completed research
- inferctl compared with OpenClaw Infer
- OpenClaw preparation pipeline assessment
- inferctl compared with InferFlux
- inferctl compared with LocalAI
- inferctl compared with llama-swap
- inferctl compared with LiteLLM Proxy
- inferctl compared with llama.cpp server router
- inferctl compared with Otari
- inferctl compared with Bifrost
- inferctl compared with NVIDIA Dynamo
- inferctl compared with llm-d Router and KServe
- inferctl compared with Ray Serve LLM
- inferctl compared with SGLang Model Gateway
All candidates in this comparison index have a completed source-based article. New candidates will be added only after source review.
Sources and update policy
The table uses the primary project documentation and repositories linked in each candidate name. Check the named project version, commit, release, or documentation date again before publication or when this page changes. Update this index after each completed comparison.