<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Build the Inference Stack Platform on Modelplane Docs</title><link>/platform/</link><description>Recent content in Build the Inference Stack Platform on Modelplane Docs</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Mon, 01 Jan 0001 00:00:00 +0000</lastBuildDate><atom:link href="/platform/index.xml" rel="self" type="application/rss+xml"/><item><title>Define Hardware Classes</title><link>/platform/inference-class/</link><pubDate/><guid>/platform/inference-class/</guid><description>&lt;p&gt;&lt;strong&gt;API:&lt;/strong&gt; &lt;a href="/reference/inferenceclasses/"&gt;&lt;code&gt;modelplane.ai/v1alpha1&lt;/code&gt; · InferenceClass&lt;/a&gt;
&lt;/p&gt;
&lt;!-- vale write-good.Passive = NO --&gt;
&lt;p&gt;An &lt;code&gt;InferenceClass&lt;/code&gt; is a tested recipe for a GPU node pool. It bundles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Devices&lt;/strong&gt;: the node&amp;rsquo;s hardware as a list of Dynamic Resource Allocation (DRA)
style devices, each with a driver, count, typed attributes, and capacity. The
scheduler matches a member&amp;rsquo;s &lt;code&gt;nodeSelector&lt;/code&gt; against these devices, and GPUs
bind to pods through DRA.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Provisioning&lt;/strong&gt; (optional): how to create a node pool of this class on a
specific cloud. Classes without provisioning are for existing clusters where
the pool already exists.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Different clouds and GPU types imply different classes. A GKE L4 pool is
&lt;code&gt;gke-l4-1x-g2&lt;/code&gt;. A bare-metal H100 pool is &lt;code&gt;h100-8x-byo&lt;/code&gt; (no provisioning).&lt;/p&gt;</description></item><item><title>Register a Cluster</title><link>/platform/inference-cluster/</link><pubDate/><guid>/platform/inference-cluster/</guid><description>&lt;p&gt;&lt;strong&gt;API:&lt;/strong&gt; &lt;a href="/reference/inferenceclusters/"&gt;&lt;code&gt;modelplane.ai/v1alpha1&lt;/code&gt; · InferenceCluster&lt;/a&gt;
&lt;/p&gt;
&lt;!-- vale write-good.Passive = NO --&gt;
&lt;p&gt;An &lt;code&gt;InferenceCluster&lt;/code&gt; represents a Kubernetes cluster configured for model
serving. Platform teams create these to provide GPU capacity.&lt;/p&gt;
&lt;p&gt;Each cluster has:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;cluster source&lt;/strong&gt;: &lt;code&gt;GKE&lt;/code&gt;, &lt;code&gt;EKS&lt;/code&gt;, &lt;code&gt;AKS&lt;/code&gt;, &lt;code&gt;Nebius&lt;/code&gt;, &lt;code&gt;Vultr&lt;/code&gt; or &lt;code&gt;Civo&lt;/code&gt; (Modelplane
provisions the full cluster) or &lt;code&gt;Existing&lt;/code&gt; (bring a cluster you manage yourself). See
&lt;a href="/platform/providers/"&gt;Supported Providers&lt;/a&gt;
for the clouds and
neoclouds Modelplane runs on.&lt;/li&gt;
&lt;li&gt;One or more &lt;strong&gt;node pools&lt;/strong&gt;, each referencing an &lt;code&gt;InferenceClass&lt;/code&gt; for its
hardware capabilities and provisioning recipe.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Labels&lt;/strong&gt; for organizational metadata: tier, region, provider. These are the
matching surface for &lt;code&gt;ModelDeployment.clusterSelector&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Modelplane installs a serving stack on every cluster it manages, including
existing clusters, which it assumes are solely for its use.&lt;/p&gt;</description></item><item><title>Set Up the Gateway</title><link>/platform/inference-gateway/</link><pubDate/><guid>/platform/inference-gateway/</guid><description>&lt;p&gt;&lt;strong&gt;API:&lt;/strong&gt; &lt;a href="/reference/inferencegateways/"&gt;&lt;code&gt;modelplane.ai/v1alpha1&lt;/code&gt; · InferenceGateway&lt;/a&gt;
&lt;/p&gt;
&lt;!-- vale write-good.Passive = NO --&gt;
&lt;p&gt;The &lt;code&gt;InferenceGateway&lt;/code&gt; is the front door for inference requests: the address a
caller sees. It speaks the OpenAI and Anthropic APIs and routes each request on
to a cluster serving the model it asked for.&lt;/p&gt;
&lt;p&gt;It runs on an &lt;code&gt;InferenceCluster&lt;/code&gt;, named by &lt;code&gt;spec.clusterName&lt;/code&gt;. The cluster it
names can serve models too, or run the gateway alone.&lt;/p&gt;
&lt;p&gt;Create as many as you need, one per cluster. A second gateway naming a cluster
that already has one reports &lt;code&gt;ClusterAlreadyHasGateway&lt;/code&gt; and doesn&amp;rsquo;t become
ready. A gateway is where a request enters your fleet, so run one per place
requests should enter from. &lt;code&gt;spec.serviceSelector&lt;/code&gt; decides
which &lt;code&gt;ModelService&lt;/code&gt;s each one serves. Left unset, a gateway serves every
service. Scoping a gateway to a region is how you express residency: label a
service for the EU and it reaches only EU gateways, and from there only the
endpoints it selects.&lt;/p&gt;</description></item><item><title>Drain a Cluster</title><link>/platform/drain-cluster/</link><pubDate/><guid>/platform/drain-cluster/</guid><description>&lt;p&gt;&lt;strong&gt;API:&lt;/strong&gt; &lt;a href="/reference/inferenceclusters/"&gt;&lt;code&gt;modelplane.ai/v1alpha1&lt;/code&gt; · InferenceCluster&lt;/a&gt;
· &lt;a href="/reference/modeldeployments/"&gt;&lt;code&gt;ModelDeployment&lt;/code&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;p&gt;Taking a cluster out of service, for maintenance or decommissioning, means
telling Modelplane to stop scheduling there and, when you can&amp;rsquo;t wait for work to
finish, to move what&amp;rsquo;s already running. You do this by tainting the
&lt;code&gt;InferenceCluster&lt;/code&gt;. It&amp;rsquo;s the fleet-level counterpart to &lt;code&gt;kubectl drain&lt;/code&gt; on a
node. Modelplane reschedules the affected replicas onto other clusters whose
hardware fits, the same way it placed them to begin with.&lt;/p&gt;</description></item><item><title>Monitor the Fleet</title><link>/platform/telemetry/</link><pubDate/><guid>/platform/telemetry/</guid><description>&lt;!-- vale write-good.Passive = NO --&gt;
&lt;p&gt;Modelplane runs an OpenTelemetry collector on every inference cluster. It
collects from every component Modelplane installs. This includes the inference
server engine, inference gateway and Envoy proxy, router, and the GPU exporter
your cloud provides. It renames each component&amp;rsquo;s series to a single
&lt;code&gt;modelplane_*&lt;/code&gt; vocabulary and exports them to wherever you say - any backend the
collector has an exporter for, not only OTLP.&lt;/p&gt;</description></item><item><title>Supported Providers</title><link>/platform/providers/</link><pubDate/><guid>/platform/providers/</guid><description>&lt;p&gt;Modelplane is built on &lt;a href="https://crossplane.io"&gt;Crossplane&lt;/a&gt;
and shares its
infrastructure providers, so the set of clouds and neoclouds it reaches grows
alongside Crossplane itself. This page shows where Modelplane runs today and
where it&amp;rsquo;s headed.&lt;/p&gt;
&lt;p&gt;A provider can show up here in three ways:&lt;/p&gt;
&lt;div class="admonition note"&gt;
&lt;div class="admonition-title"&gt;
&lt;svg class="bi flex-shrink-0" role="img" aria-label="note:"&gt;&lt;use
xlink:href="#info"/&gt;&lt;/svg&gt;
&lt;span class="ps-1"&gt;Note&lt;/span&gt;
&lt;/div&gt;
&lt;div class="admonition-content"&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Provisioning supported.&lt;/strong&gt; Modelplane creates and manages the whole cluster
from an &lt;code&gt;InferenceCluster&lt;/code&gt;, selected through &lt;code&gt;provisioning.provider&lt;/code&gt;. GKE, EKS, AKS, Nebius mk8s, Vultr VKE, Civo K3s work this way today.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bring your own supported.&lt;/strong&gt; Register a cluster you already run with
&lt;code&gt;source: Existing&lt;/code&gt;. This works on any provider whose Kubernetes meets
Modelplane&amp;rsquo;s requirements (Dynamic Resource Allocation and a recent Kubernetes
version), so you can run on the providers below now, ahead of native
provisioning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Crossplane provider exists.&lt;/strong&gt; A Crossplane provider is published for the
cloud. That provider is the path by which native provisioning lands, so it
marks where Modelplane can grow next.&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="clouds-and-neoclouds"&gt;Clouds and neoclouds &lt;a class="anchor-link" id="clouds-and-neoclouds" href="#clouds-and-neoclouds" aria-label="Link to this section: Clouds and neoclouds"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Listed alphabetically, spanning hyperscalers and GPU-specialist neoclouds. Each
runs a managed Kubernetes service with GPU node pools, so the bring-your-own path
covers them all today. Where a Crossplane provider exists, it&amp;rsquo;s the path to
native provisioning.&lt;/p&gt;</description></item></channel></rss>