LLMOps& AI Platforms
Model Serving & Inference on your GPU infrastructure
Extend AI Infrastructure into model APIs, from DGX Spark and DGX Station GB300 to SuperPOD, with Kubernetes / Slurm, RDMA and shared GPU resource management.
Model Serving
& Inference on GPU
Extend your installed GPUs with a software stack for model APIs, team resource allocation and measured acceptance before handover
Connect with AI Infrastructure ↗An API, model version and configuration for application teams
Reviewable access, quotas and tenant boundaries
Benchmark report, dashboards, runbook and rollback procedure
From a desktop to an inference cluster
Start with models and usage patterns, then size GPUs, networking and replicas
Which models can you start serving on each DGX?
Starting candidates based on memory and precision. Confirm the checkpoint, engine and license before deployment.
DGX Spark
1 × GB10 Grace Blackwell
- Memory
- 128 GB LPDDR5x unified
- Bandwidth
- 273 GB/s
- Peak compute
- 1 PFLOP · FP4 sparse
Prototype RAG and private assistants · 1 system
Starting models / precisionQwen3-32B · FP4 / gpt-oss-120b · MXFP4
Start with NVIDIA Spark playbooks and a GB10-compatible runtime. Reserve unified memory for the OS and KV cache.
Model card / deployment guide ↗DGX Station GB300
1 × Blackwell Ultra + Grace CPU
- Memory
- 252 GB HBM3e + 496 GB CPU
- Bandwidth
- 7.1 TB/s · GPU HBM
- Peak compute
- 15 / 20 PFLOPS · FP4 dense / sparse
Team development and model evaluation · 1 system
Starting models / precisionQwen3-32B · BF16 / Llama 3.3 70B · BF16 / FP8
Single-GPU evaluation candidates: 70B BF16 weights use roughly 140 GB before KV cache. CPU memory has different bandwidth from HBM.
Model card / deployment guide ↗DGX H200
8 × H200 · Hopper
- Memory
- 141 GB / GPU · 1,128 GB total
- Bandwidth
- 4.8 TB/s / GPU
Pilot a chatbot or coding assistant · 1 node
Starting models / precisionLlama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8
Evaluation candidates: 70B across 2–8 GPUs; start R1 FP8 across eight GPUs with tensor / expert parallelism, then measure KV cache and concurrency.
Model card / deployment guide ↗DGX B200
8 × B200 · Blackwell
- Memory
- 180 GB / GPU · 1,440 GB total
- Bandwidth
- 8 TB/s / GPU · 64 TB/s total
- Peak compute
- 72 / 144 PFLOPS · FP4 dense / sparse
Evaluate large models or separate serving replicas · 1 node
Starting models / precisionLlama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8
Evaluate R1 FP8 across eight GPUs or separate 70B replicas. Choose GPUs per replica from the latency target.
Model card / deployment guide ↗DGX B300
8 × B300 · Blackwell Ultra
- Memory
- 288 GB / GPU · 2,304 GB total
- Bandwidth
- 8 TB/s / GPU · 64 TB/s total
More memory for model weights and KV cache · 1 node
Starting models / precisionLlama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8 / BF16 (converted)
Start with R1 FP8. Converting to BF16 with a compatible engine uses roughly 1.34 TB of weights across eight GPUs before KV cache.
Model card / deployment guide ↗GB300 NVL72
72 × Blackwell Ultra · 36 Grace CPUs
- Memory
- 20 TB GPU HBM · rack total
- Bandwidth
- 576 TB/s · rack aggregate
- Peak compute
- 1,080 / 1,440 PFLOPS · FP4 dense / sparse
Inference pool for multiple models and teams · 1 rack
Starting models / precisionDeepSeek-R1 671B · FP8 / BF16 (converted) · multi-replica
Partition GPUs into serving groups for replicas or disaggregated prefill/decode. Benchmark one group before scaling across the rack.
Model card / deployment guide ↗DGX Vera Rubin NVL72
72 × Rubin · 36 Vera CPUs
- Memory
- 20.7 TB HBM4 · rack total
- Bandwidth
- 1,400 TB/s · rack aggregate
- Peak compute
- 3,600 PFLOPS · NVFP4 inference
Plan reasoning and long-context inference · 1 rack
Starting models / precisionDeepSeek-R1 671B · FP8 · evaluation candidate
Preliminary NVIDIA specifications; the referenced table does not specify sparsity for NVFP4 inference. The model is an evaluation candidate: confirm Rubin-compatible runtimes and kernels before deployment.
Model card / deployment guide ↗Peak compute uses the stated precision and sparsity; it is not measured model tokens/s. Aggregate GPU memory requires parallelism and room for weights, KV cache and runtime. These examples are not Zotect benchmark results.
Specifications reviewed · H200 / B200 / B300 bandwidth ↗ · Spark model support ↗
Scale beyond one system as workloads grow
Add replicas for multiple teams or distribute a model across nodes. These are benchmark starting points; concurrency is a test load to validate with real workloads.
DGX / HGX H200 · B200 · B300
8–256 nodes / 64–2,048 GPUsReplicas & service continuityShared inference with replicas and spare capacity for maintenance and releases
If a replica fits one node: example 8 nodes = 4 serving + 2 spare + 2 batch; a two-node replica needs a different reserve plan
H200 · B200 · B300 cluster
8–256 nodes / 64–2,048 GPUsDistributed inferenceEvaluate 1–3T-parameter frontier MoE models or disaggregated prefill/decode for long-context agents
This is total cluster capacity across replicas; start each serving-group benchmark at supported FP4/FP8, 32–64K context and 8–32 requests; budget all expert weights, KV cache, communication buffers and spare replicas
GB300 NVL72 → SuperPOD
1–48 racks / 72–3,456 GPUs18 compute trays per rack · 4 GPUs per trayLarge inference pools for multiple models and tenants, with online and batch capacity separated
Plan in racks, not eight-GPU nodes: 1–48 racks = 18–864 four-GPU compute nodes. SuperPOD naming requires the NVIDIA reference architecture
DGX Vera Rubin NVL72
1–48 racks / 72–3,456 GPUs72 Rubin GPUs + 36 Vera CPUs per rackPlan a Rubin inference pool for reasoning, long-context workloads and multiple tenants
The rack range is a planning example. NVIDIA specifications are preliminary; confirm configuration, software support, power and cooling before final design
How we turn examples into a deployment size
Measure weights + KV cache + runtime overhead, then benchmark representative prompts, output lengths and concurrency. Add replicas for throughput and complete-replica failure capacity. Use multi-node parallelism when memory or computation requires it; tokens/s does not scale linearly with node count
Budget all MoE expert weights, KV cache and runtime. Measure p95 TTFT, p95 inter-token latency and aggregate tokens/s using representative context, output and concurrency before setting capacity and HA.
Connect GPUs, networking and scheduling
Choose orchestration for the workload; every deployment does not need every tool
Kubernetes + NVIDIA Operators
Install NVIDIA GPU Operator for drivers, container runtime, device plugins and telemetry against the support matrix, with NVIDIA Network Operator for network drivers, RDMA device plugins and secondary networks such as Multus / SR-IOV as the topology requires
Documentation ↗RDMA & distributed serving
Validate GPU–NIC locality, InfiniBand or RoCE fabric, GPUDirect RDMA and inter-node NCCL before inference. Choose tensor, pipeline or expert parallelism for the model. RDMA requires compatible hardware, drivers and fabric, not just a plugin
Documentation ↗Slurm for scheduled workloads
Use Slurm partitions, GRES, QoS and accounting for batch inference, evaluation and fine-tuning. Use Kubernetes for persistent APIs, or separate resource pools when both schedulers are deployed so they do not allocate the same GPUs
Documentation ↗Spark / Station are development and local-inference entry points. Cluster deployments must align GPU architecture, Arm64/x86, OS, drivers, NICs and Kubernetes/Slurm versions. Confirm NIM/operator support for each platform before proposing the stack
Shared GPUs. Clear team boundaries.
Administrators allocate through UI/API; users see and consume resources within their team permissions
NVIDIA Base Command Manager
Provision OS images, configure nodes and monitor cluster health centrally, integrating Kubernetes or Slurm management for the selected design
Platform capabilities ↗NVIDIA Run:ai · GPU / CPU / RAM
Administrators manage departments, projects, node pools, quotas and borrowing policies through UI/API. Users submit within their permissions; teams receive GPU, CPU and CPU-memory budgets with scheduling and resource visibility
Platform capabilities ↗Storage / disk & tenancy
Combine PVCs / storage classes, requests.storage and ephemeral-storage quotas with Kubernetes ResourceQuota and storage-backend policy. Separate namespaces, RBAC, secrets and network policies per tenant; quotas alone are not security isolation
Platform capabilities ↗Monitoring & capacity trends
Connect DCGM Exporter, Prometheus and Grafana for GPU utilization and memory, CPU/RAM, storage, fabric errors and per-team quota usage. Retain history for trends alongside TTFT, latency, tokens/s and queue depth
Platform capabilities ↗Choose a serving stack your team can operate
Combine NVIDIA enterprise and open-source components around compatibility and support needs
NVIDIA AI Enterprise + NIMNVIDIA enterprise
Inference microservices with an enterprise-support path. Match containers and GPUs to the support matrix; verify production entitlements separately from model licenses
Tool documentation ↗NVIDIA Dynamo + TensorRT-LLMNVIDIA open source
Distributed serving, prefill/decode and KV-aware routing with Dynamo; select TensorRT-LLM where the model and backend are supported
Tool documentation ↗vLLM / SGLangOpen source
API serving and continuous batching; choose the runtime from model architecture, quantization and measured performance
Tool documentation ↗KV cache + LMCacheOpen source
KV cache stores attention state; LMCache supports cache reuse, offload and transfer with compatible backends. Validate memory budgets, hit rates and isolation
Tool documentation ↗LiteLLMOpen source + enterprise options
API gateway for routing, keys, rate limits and token budgets, depending on edition; not a GPU scheduler, and token budgets are not GPU quotas
Tool documentation ↗DCGM Exporter · Prometheus · GrafanaOpen tooling
Resource and endpoint telemetry, dashboards and alerts, with retention suited to capacity planning
Tool documentation ↗Start with one serving engine. Add a gateway, KV-cache layer or distributed serving when measurements justify it. Verify NVIDIA AI Enterprise entitlements and commercial features for the delivered versions; open weights do not imply identical open-source licensing
Models to evaluate with your workloads
Selected high-ranking LLM Stats examples, linked to publisher weights. These are candidates, not configurations already certified by Zotect
LLM Stats · Checked · Overall ranks across all models at review date
Kimi K3
Candidate for coding, reasoning and multimodal-agent evaluation
moonshotai/Kimi-K3GLM-5.3
Candidate for coding and multi-step agent evaluation
zai-org/GLM-5.3Qwen3.8 Max / open checkpoint
Score is for Qwen3.8 Max. The related open checkpoint is not the identical hosted API; re-evaluate quality and features
Qwen/Qwen3.8-2.4T-A95BDeepSeek-V4-Pro-0813
Candidate for reasoning, coding and tool-using agents
deepseek-ai/DeepSeek-V4-Pro-0813Benchmark rankings are not adoption statistics and do not establish Thai-language or domain quality. Recheck quality, licensing, engine support and checkpoint memory; benchmark real context lengths rather than treating model-card maximum context as guaranteed serving capacity
How do open weights compare with frontier AI?
Compare by task to shortlist models for evaluation on your GPUs.
Snapshot checked September 23, 2026. Frontier references are the versions in the source tables, not a live ranking. Open-weight describes access under each model’s license; open-weight models can also be frontier models.
Kimi K3 (max)compared withGPT-5.6 Sol (max)
Reported by Moonshot AI ↗GPQA Diamond
Terminal-Bench 2.1
Scores are close on these two benchmarks. Terminal results use different agent harnesses; repeat evaluation in a shared workflow.
GLM-5.3compared withGPT-5.6 Sol
Reported by Z.AI ↗Terminal-Bench 2.1
DeepSWE v1.1
Close on terminal tasks, with a wider gap on DeepSWE. Coding selection depends on the task and agent harness.
DeepSeek-V4-Pro-0813compared withClaude Fable 5 (w/ fallback)
Reported by DeepSeek ↗Terminal-Bench 2.1
DeepSWE
Terminal scores are close; DeepSWE differs by 7.3 points. The source labels the Claude results as including fallback.
Higher is better on these benchmarks. Bars share a 0–100 scale; do not average across benchmarks. Publisher results may use different reasoning budgets, tools and harnesses. Similar scores do not establish statistical equivalence, universal interchangeability, Thai-language quality, or speed and cost on your GPUs.
Already have GPUs? Start with a serving baseline
Share GPU types and node counts, target models, context, concurrency and latency targets to scope deployment and benchmarking
Deployment and Release
Control versions, promotion, and rollback for models and configuration.
- Model version
- Promotion
- Rollback
RAG Production Integration
Connect retrieval sources and applications with defined inspection points.
- Data sources
- Retrieval
- Application
Performance and Cost
Capture throughput, latency, and resource use for operational decisions.
- Throughput
- Latency
- Resource usage
Observability and Governance
Record operating signals, changes, and access.
- Operating signals
- Change records
- Access
Managed AI Platform Operations
Provide proactive platform care after technical-baseline acceptance.
- Technical baseline
- Platform care
- Service agreement
What to track when systems change
Define owners, release records, rollback paths and operating signals for each service.
Release Control
Record versions and approvals for models, prompts, and configuration.
Observability
View application, model, and infrastructure signals in one context.
Security
Control access to data, model endpoints, and operational tools.
Start with the state of your system
Each stage is a separate scope, starting with a shared review of the technical baseline.
Discover
Assess serving, data paths, and operations, then prioritize gaps.
Build
Implement the platform, integrations, observability, and release path.
Support
Provide reactive help with incidents, configuration reviews, and model changes.
Operate
Provide proactive platform care under a service agreement.
Start with your models, data and platform
Review serving, data paths and operations to define the scope together.
Explore AI Infrastructure