Choosing an Inference Path on AWS: From Managed APIs to Self-Hosted Accelerators
A decision guide for choosing an inference path on AWS — Amazon Bedrock, Amazon SageMaker, or self-hosted on ECS/EKS - and the accelerator beneath it (CPU, GPU, Inferentia, Trainium).
Teams commonly run model inference on AWS in one of three ways: an Amazon Bedrock API call, an Amazon SageMaker endpoint, or a self-managed serving stack (for example, vLLM) on Amazon EKS. Each is appropriate for different requirements, and the difference in outcome usually comes less from the technology selected than from how deliberately it was matched to the workload.
A frequent anti-pattern is defaulting to GPU-backed self-hosting for every workload. The result is often low accelerator utilization, over-provisioned endpoints, and a serving stack that carries more operational burden than the team can sustain, all of which increase cost without improving customer outcomes. A useful starting point is to select the most managed option that satisfies your requirements and adopt greater operational ownership only when a concrete requirement justifies it, recognizing that each step toward self-hosting increases both control and operational responsibility.
This guide front-loads that decision: it opens with a framework for choosing a path, then profiles each hosting option, examines the compute layer beneath (CPU, GPU, Inferentia, Trainium), and covers the cross-cutting concerns (latency, cost modeling, operations) that apply regardless of path. A self-hosted deep dive, which readers on managed services can skip, comes last.
The decision framework
Make the drivers of the decision explicit before comparing services. Five axes cover the majority of cases; the recommendations throughout this guide map back to them.
| Axis | Question to answer | Tends to favor |
|---|---|---|
| Control and customization | Do you require a specific model, custom weights, or a bespoke serving stack? | Self-hosting |
| Operational burden | How much infrastructure is the team resourced to operate? Is there GPU and Kubernetes expertise? | Managed services |
| Cost at scale | Is traffic spiky and low-volume, or steady and high-throughput? | Steady high volume can favor self-hosting on unit cost |
| Latency objective | What is the p95 target? | Strict objectives favor dedicated capacity and the appropriate accelerator |
| Data residency | Must data remain in your account, VPC, or Region? | Self-managed in-VPC serving |
Choosing a path
Proceed through the questions in order and stop at the first option that satisfies your requirements.

- Can a foundation model behind an API meet the requirement? → Amazon Bedrock. Fastest to production, no model ownership required, well suited to spiky or low-to-moderate volume.
- Do you own the model but want AWS to operate the infrastructure? → Amazon SageMaker. Managed hosting for your own or open-weight models across real-time, serverless, asynchronous, and batch paths.
- Do you require a specific serving stack, hardware control, scale economics, or full control of the data path, and have the expertise to operate it? → Self-hosted on Amazon ECS or EKS.
Signals to change paths
Movement between paths goes both directions. Teams move down the stack for control and up for operational relief; neither direction is inherently better, and the right position is the one that meets current requirements. Move only on a specific signal, not intuition, and validate it with a cost-per-token comparison at your projected utilization (see Cost modeling), since managed options often remain cheaper once operational cost is included.
Signals to move down the stack (toward more control):
- Bedrock → SageMaker: you need to serve your own fine-tuned or open-weight model, require a deployment shape Bedrock does not offer (for example, asynchronous or batch transform on your own model), or want per-model control of instance type and scaling.
- SageMaker → self-hosted: you need a specific serving engine, custom kernels, or a parallelism or GPU-sharing strategy the managed service does not expose; steady high-volume traffic makes owning the fleet cheaper per token; or you need full control of the data path and scheduler.
Signals to move up the stack (toward less operational burden):
- Self-hosted → SageMaker: utilization stays low for weeks, on-call burden exceeds team capacity, or the serving-engine customization you self-hosted for is now available in the managed service.
- SageMaker → Bedrock: your model is now available through Bedrock (including via Custom Model Import), or traffic has dropped enough that per-token pricing is cheaper than a persistent endpoint.
Hosting options
AWS inference hosting ranges from fully managed to fully self-operated. Select the most managed option that meets the workload's requirements.

1. Amazon Bedrock: fully managed foundation models
A single API to a catalog of foundation models (Anthropic, Meta, Mistral, Amazon, and open-weight families), with no infrastructure to operate. Several capabilities affect cost and latency:
- On-Demand (per-token): No commitment; pay per input/output token. Best for spiky traffic and the fastest path to production.
- Provisioned Throughput: Reserves dedicated model capacity (model units) at committed hourly pricing, for steady, high-volume, or latency-sensitive workloads and for custom or imported models. It becomes favorable once steady utilization makes its committed hourly cost lower than the equivalent per-token spend, so keep spiky traffic on On-Demand.
- Batch inference: Asynchronous processing at approximately 50% of On-Demand pricing, for non-real-time work such as bulk summarization or embeddings.
- Custom Model Import and fine-tuning: Bring fine-tuned weights (for supported architectures) or fine-tune hosted models, served behind the same API.
- Latency-optimized inference: An opt-in mode for supported models that reduces response latency for interactive use cases.
- Prompt caching: Reuses cached prompt prefixes so they are not reprocessed each call, a cost and latency lever for retrieval-augmented and agentic applications with stable context (system prompts, tool definitions, retrieved context).
How you use it. Call the
InvokeModel or Converse API with your model ID; there is no infrastructure to provision. Start on On-Demand, move steady traffic to Provisioned Throughput once it crosses the break-even point, and route non-real-time work to Batch. Enable prompt caching for large stable prefixes and latency-optimized inference where time-to-first-token binds. Bedrock Guardrails (content and safety policies), Knowledge Bases (managed retrieval-augmented generation), and Agents (tool-calling orchestration) are available as managed building blocks, so you can add governance and RAG without operating that infrastructure yourself.Consider Bedrock when you want a foundation model behind an API and do not need to own the serving stack; traffic is spiky or low-to-moderate volume, so per-token pricing (no charge for idle capacity) is preferable to reserved capacity; or you want managed guardrails, knowledge bases, and agent orchestration rather than building them.
Trade-off. Bedrock provides the least infrastructure-level control: you do not select the instance or tune the serving engine, and custom models are limited to supported import paths. For many teams building generative AI features, it is an appropriate starting point and often a suitable long-term choice. Bedrock is a Regional service; for multi-Region resilience, call it in more than one Region behind Amazon Route 53 or a global routing layer.
2. Amazon SageMaker: managed hosting for your own models
You bring the model (your own or open-weight) and AWS operates the infrastructure. Four deployment options cover most requirements:
- Real-time endpoints: persistent, low-latency online inference.
- Serverless endpoints: fully managed, auto-scaling, with scale-to-zero for intermittent traffic.
- Asynchronous endpoints: large payloads and long-running inference, queued.
- Batch transform: high-throughput offline scoring without a persistent endpoint.
How you use it. Package the model with an inference container (a prebuilt Deep Learning Container, a framework image such as Hugging Face TGI or DJL/LMI, or your own), register it as a SageMaker model, and create an endpoint of the appropriate type on a CPU, GPU, or Neuron instance with an autoscaling policy. For multiple models, define each as an inference component with its own accelerator, memory, and scaling policy; SageMaker packs them onto shared instances and scales each independently (each model scales its own copy count while SageMaker adds or removes instances to fit). AWS reports this reduces deployment cost by about 50% on average relative to one model per endpoint. One packing limit is worth noting: an inference component's
NumberOfAcceleratorDevicesRequired has a minimum of 1, so an accelerator-backed component claims at least one whole GPU or Inferentia device (there is no fractional accelerator sharing), whereas CPU-only components pack more densely because CPU cores are allocatable from 0.25 and memory in MB. Where many models must share one accelerator, a common pattern is a single inference component that claims one device and routes among the models inside its own container, trading per-model scaling and metrics for higher density.Consider SageMaker when you own the model and want managed hosting rather than raw infrastructure; you need real-time and batch or asynchronous paths within one service; or you want to pack multiple models onto shared instances with per-model autoscaling, without operating Kubernetes.
Trade-off. SageMaker offers more control than Bedrock and correspondingly more operational involvement. You select instance types and scaling policies, while AWS manages the underlying fleet.
3. Self-hosted on Amazon ECS or Amazon EKS: full control
You run the serving stack (for example, vLLM, TorchServe, or Triton) in containers on Amazon ECS or Amazon EKS, on instances you select (GPU, Graviton CPU, or Inferentia/Trainium), with your own autoscaling (Karpenter, KEDA) and accelerator-sharing strategy. This path exposes the serving engine, parallelism, batching, and scheduling that the managed services abstract away. It also provides fine-grained accelerator scheduling that neither Bedrock nor SageMaker offers.
Consider self-hosting when you require a specific serving engine, custom kernels, or a parallelism strategy that Bedrock and SageMaker do not expose; traffic is steady and high-volume, so operating the fleet improves unit economics through direct control of accelerator utilization; or you have strict data-path or audit requirements that call for full control over the serving stack, and the platform expertise to operate it.
Trade-off. Self-hosting provides the most control and can achieve the best unit cost at scale, in exchange for owning drivers, device plugins, autoscalers, capacity planning, and on-call responsibilities. The Self-hosted deep dive, later in this guide, walks that operational surface. Evaluate managed options first and adopt self-hosting when a specific requirement justifies it.
At a glance
| Bedrock | SageMaker | Self-hosted (ECS/EKS) | |
|---|---|---|---|
| You provide | API calls | The model | The model and the serving stack |
| AWS operates | Everything | The fleet | Nothing (you do) |
| Cost model | Per-token, or provisioned/batch | Per-instance-hour, packed per model | Per-instance-hour, you optimize |
| Operational burden | Lowest | Moderate | Highest |
| Best-fit signal | Spiky/low volume, no model ownership | Own model, want managed hosting | Custom stack, steady scale, or data-path control |
The compute layer: CPU, GPU, Inferentia, and Trainium
The accelerator decision sits beneath whichever hosting option you select, and it accounts for a significant share of cost. A common error is provisioning a GPU by default when the workload does not require one; CPUs and accelerators are complementary rather than competing.
| Factor | Favors CPU | Favors GPU / Trainium |
|---|---|---|
| Model size | SLMs 1–8B (quantized), embeddings, classifiers | 8B+ for low-latency online serving; 70B+ in general |
| Latency objective | p95 of 100–500 ms acceptable | p95 below 50 ms required |
| Concurrency | Below ~100 req/s per endpoint | Above ~100 req/s sustained |
| Workload type | Orchestration, retrieval, ETL, batch scoring | Online inference, fine-tuning, training |
| Capacity | Immediate availability, no reservations | Often requires reserved capacity |
| Cost sensitivity | Best cost per token for eligible workloads | Amortizes at high utilization |
| Team expertise | Standard Kubernetes operations | Requires GPU operations expertise |
As guidelines by model size: SLMs (1–8B, quantized) default to CPU, moving to GPU when p95 must be below 50 ms or sustained concurrency exceeds ~100 req/s; medium models (8–30B) use CPU for batch or offline work and GPU for online serving; and large models (70B+) use GPU or Trainium for interactive workloads, with CPU viable only for offline or batch use with heavy quantization. Treat these as starting points and validate against your model, data, and latency budget.
| Option | Description | Best for | Watch for |
|---|---|---|---|
| CPU (incl. AWS Graviton) | General-purpose arm64/x86 compute | SLMs, embeddings, classical ML, and the CPU-bound work around inference (retrieval, orchestration, guardrails) | Not for larger models (~8B+) at low latency or high concurrency |
| GPU (NVIDIA) | Broadly compatible accelerator (e.g., G5; higher-end families for the largest models) | Low-latency online inference; training and fine-tuning | Scarce capacity may need reservations; only cost-effective at high utilization |
| AWS Inferentia (Inf2) | Purpose-built inference silicon (NeuronCores) | High-throughput LLM and diffusion inference at scale | Requires ahead-of-time Neuron compilation |
| AWS Trainium (Trn) | Purpose-built training/fine-tuning silicon (NeuronCores) | Cost-effective training and fine-tuning of large (100B+) models; distributed training | Uses the Neuron SDK; confirm your model architecture is supported |
Inferentia2 delivers up to 4× higher throughput and 10× lower latency than the prior generation (Inferentia1), and AWS reports that Inf2 and Trn instances can offer up to 50% lower cost to deploy models such as Llama 3 relative to comparable EC2 instances. A common pattern is to train on GPUs (or Trainium) and serve on Inferentia; the AWS Neuron SDK keeps the model portable across them. On Inferentia or Trainium you compile the model ahead of time with the Neuron SDK for a fixed batch size, sequence length, and core count, then serve with a Neuron-enabled vLLM or TGI container.
Terminology note: Inf2 and Trn instances do not expose GPUs. They provide NeuronCores (aninf2.48xlargehas 24), scheduled in Kubernetes through theaws.amazon.com/neuronresource rather thannvidia.com/gpu.
Cross-cutting concerns
The following apply to every hosting path.
Latency: reasoning beyond throughput
Throughput (RPS, tokens per second) indicates how much a fleet can serve; latency indicates whether the experience meets user expectations. The two are in tension: increasing batch size raises throughput while individual requests wait longer. Reason about latency explicitly rather than inferring it from throughput, and track two metrics:
- Time to first token (TTFT): the interval before the first token appears, which corresponds to perceived responsiveness in a streaming interface. It is dominated by the prefill phase (processing the full prompt), which is compute-bound: long prompts increase prefill cost and therefore TTFT.
- End-to-end latency and time per output token (TPOT): the total time to complete the response, dominated by the decode phase (generating tokens sequentially), which is memory-bandwidth-bound: each step reads the full KV cache and weights. This is why memory-optimized accelerators and a compact KV cache matter, and why decode dominates end-to-end latency for long outputs.
The split has practical consequences: a long-input, short-output workload (for example, 16K in / 512 out) is prefill-heavy, so TTFT is the primary risk; a short-prompt, long-answer workload is decode-heavy, so TPOT and memory bandwidth dominate. Levers for meeting a latency objective include chunked prefill (interleaves a long prompt's prefill with decode steps so it does not stall other requests), batch sizing (capping batch width protects p95 under load), tensor parallelism (sharding across accelerators lowers per-token latency, to a point), speculative decoding (a draft model proposes tokens the target verifies in parallel; supported on GPU and Inferentia2), and quantization (smaller weights reduce memory traffic per decode step). On Bedrock, prompt caching and latency-optimized inference are the managed equivalents.
Define objectives as a pair. A target such as "p95 TTFT below 300 ms and p95 end-to-end below 3 s at 50 RPS" is specific and testable; a single throughput figure is not, and relying on it can produce a fleet that benchmarks well but feels unresponsive.
Cost modeling: compare cost per token, not hourly price
Accelerator families cannot be compared on hourly price alone. The reliable approach is a structured benchmark against your own model and traffic, converted into a cost per token at your projected utilization.
Fix the workload and objective before measuring (model, representative input and output token lengths, and a target p95 latency), and hold them identical across candidate instance families so the comparison is fair. On each candidate, tune tensor parallelism to meet the latency target on a single model copy, then scale with data parallelism until the hardware is saturated (see Parallelism in the deep dive). This establishes the best throughput each instance sustains within your latency objective. Then convert to cost per token:
1
2
tokens per hour = RPS × 3600 × (input + output tokens per request)
cost per 1M tokens = (hourly instance price ÷ tokens per hour) × 1,000,000
An instance with a higher hourly rate can have a lower cost per token if its throughput is proportionally higher, which is the basis for purpose-built accelerators' price-performance advantage at scale. Crucially, self-hosted cost is fixed hourly regardless of load, so the effective cost per token is the saturated figure divided by utilization: an endpoint at 25% utilization costs about 4× its saturated rate, and at 10% about 10×. Keeping the accelerator busy therefore has more impact on unit cost than the accelerator choice itself.
As an illustration of the method (not an AWS-published benchmark): a hypothetical endpoint sustaining 40 requests per second of 512-token responses on a $5.67/hr instance processes about 73.7M tokens/hour, or roughly $0.077 per 1M tokens at saturation; at 25% utilization that same endpoint costs about $0.31 per 1M tokens. Substitute your own measured throughput and current prices. This is why the decision resolves as:
- Steady, high utilization (above roughly 60–70%): self-hosting or SageMaker on the appropriate accelerator can win on unit cost; tune parallelism to saturate the hardware.
- Spiky or low volume: Bedrock On-Demand, where you pay per token and the idle-time penalty does not apply.
- Intermediate: SageMaker with autoscaling and inference components, or Bedrock Provisioned Throughput sized to the baseline with On-Demand for peaks.
Compute cost per 1M tokens for each candidate at your projected utilization rather than at 100%: an instance that appears expensive at saturation is often the lower-cost option in practice, and the reverse also holds.
How you purchase capacity is a separate lever from how efficiently you use it, and it should match the traffic shape. On-Demand Capacity Reservations and EC2 Capacity Blocks for ML reserve scarce GPU capacity for defined windows; Savings Plans and Reserved Instances discount steady long-term use; and Spot Instances can reduce cost by up to 90% for interruption-tolerant training and batch inference (Trainium and Inferentia are available on Spot as well). Steady baseline load suits Savings Plans or Provisioned Throughput; predictable bursts suit Capacity Blocks or Spot; and unpredictable interactive traffic favors On-Demand or Bedrock per-token pricing.
Self-hosted deep dive
This section covers the operational surface of running inference on Amazon EKS: a reference vLLM deployment, followed by the serving-efficiency, GPU-sharing, and autoscaling techniques.
Reference implementation: serving an LLM on Amazon EKS with vLLM
A minimal deployment: vLLM on a GPU node, fronted by a service, with Karpenter provisioning the hardware.
1. Provision GPU nodes on demand with Karpenter (
NodePool):1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: gpu
spec:
template:
spec:
requirements:
- key: node.kubernetes.io/instance-type
operator: In
values: ["g5.12xlarge"] # or inf2.48xlarge for Neuron
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand", "spot"] # Spot for interruption-tolerant serving
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: gpu
limits:
nvidia.com/gpu: 100
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized # reclaim idle capacity
2. Deploy vLLM as an OpenAI-compatible server (
Deployment):1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
apiVersion: apps/v1
kind: Deployment
metadata:
name: llama-vllm
spec:
replicas: 1 # scale out for data parallelism
selector:
matchLabels: { app: llama-vllm }
template:
metadata:
labels: { app: llama-vllm }
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- "--model=meta-llama/Llama-3.1-8B-Instruct"
- "--tensor-parallel-size=4" # TP=4: shard across 4 A10G GPUs
- "--max-model-len=16384"
- "--gpu-memory-utilization=0.90"
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: "4" # claim all 4 GPUs on the node
nodeSelector:
karpenter.sh/nodepool: gpu
3. Expose and invoke the service. vLLM implements the OpenAI API, so clients are straightforward:
1
2
3
4
5
6
7
kubectl expose deployment llama-vllm --port=80 --target-port=8000 --type=LoadBalancer
curl http://<lb-address>/v1/completions \
-H "Content-Type: application/json" \
-d '{"model": "meta-llama/Llama-3.1-8B-Instruct",
"prompt": "Explain tensor parallelism in one sentence.",
"max_tokens": 128}'
To run the same stack on Inferentia2, substitute a Neuron-enabled vLLM or TGI container, target
inf2.48xlarge, request aws.amazon.com/neuron devices rather than nvidia.com/gpu, and pre-compile the model with the Neuron SDK. vLLM abstracts the accelerator, so the serving interface stays consistent across GPU and Neuron.Improving accelerator utilization
Many deployments use only a fraction of the accelerator they pay for. A 7B model in FP16 needs about 14 GB of VRAM, but an H100 provides 80 GB, so a single model on its own GPU leaves roughly 66 GB idle. Industry analyses commonly report GPU utilization below 40% at moderate traffic.
There are two separate causes of waste, each with its own fix.
Cause 1: memory is underused within one model's serving. Fix: a modern serving engine. The largest memory consumer in LLM serving is the KV cache (the key/value vectors kept for every token, layer, and head). It grows with batch size and sequence length; for Llama-3.1-70B at 131K context it reaches roughly 43 GB per request, larger than the weights themselves.
Two techniques in engines like vLLM, TGI, and TensorRT-LLM address this:
- PagedAttention allocates the KV cache in small fixed-size blocks on demand, the way an operating system manages RAM. This minimizes wasted memory and lets requests with identical prefixes share it.
- Continuous batching adds and removes requests at each decoding step, so the accelerator never idles waiting for a batch to finish.
Together they deliver 2–4× higher throughput at the same latency versus naive static batching, and adopting such an engine is often the single highest-leverage change. Tune three settings:
--gpu-memory-utilization (VRAM share for weights and cache), --max-model-len (caps cache growth per request), and --max-num-seqs (batch width), leaving headroom to avoid out-of-memory errors under bursty long-context traffic.Cause 2: a whole GPU is oversized for one model. Fix: GPU sharing. For a single large LLM this rarely applies: that workload wants the whole device, scaled across nodes with data parallelism. Sharing pays off for the other workloads on a platform: many small models, classical ML, embeddings, and dev/test endpoints that would each otherwise strand most of an expensive GPU. NVIDIA offers three mechanisms, trading isolation for density:
| Mechanism | Isolation | When to use | Primary consideration |
|---|---|---|---|
| MIG: partitions a GPU into up to 7 hardware instances | Hardware-enforced (memory and fault) | Production multi-tenant with defined SLOs; regulated workloads | A100/H100 and later only; static profiles; cannot burst idle capacity between slices |
| MPS: processes share one CUDA context, kernels overlap | Soft (shared fault domain) | Many small concurrent models (approximately 2×+ density) | A misbehaving process can affect neighbors |
| Time-slicing: round-robins GPU time | None | Development and test; bursty or low-concurrency | No memory isolation; contention above ~80% utilization |
In short: MIG when you need isolation, MPS for many small concurrent models, and time-slicing for dev/test only.
On managed hosting, SageMaker achieves the same packing through inference components (described under Amazon SageMaker), with multi-model endpoints as a lighter option for a long tail of rarely used models on one framework.
Fine-grained scheduling: Dynamic Resource Allocation (DRA)
This level of scheduling control exists only when you run the scheduler yourself; Bedrock and SageMaker manage placement and do not expose it.
The traditional Kubernetes model (request
nvidia.com/gpu: 1 and receive one opaque GPU) cannot express a modern accelerator's memory, topology, or partitioning modes. Dynamic Resource Allocation (DRA) replaces it with a claim-based model: a workload asks for a class of device (for example, "a GPU with at least 40 GB and NVLink") and the scheduler assigns one that fits. The core APIs reached general availability in Kubernetes 1.34, and DRA is expected to become the standard way multi-tenant accelerator scheduling is expressed, with MIG, MPS, and time-slicing requested through a claim rather than configured per node.AWS availability (as of 2026): On Amazon EKS, the NVIDIA DRA driver is not yet supported with Karpenter or EKS Auto Mode, so use the NVIDIA device plugin for the architecture in this guide.
Parallelism and autoscaling
Two parallelism controls set how a model uses accelerators, and the order matters:
- Tensor parallelism (TP) splits one model's weights across several accelerators in a single instance. Use it when the model does not fit on one device, or to lower latency. It adds cross-device communication, so more is not always better.
- Data parallelism (DP) runs independent copies of the model to add throughput.
Tune TP first to meet the latency target on one copy, then scale out with DP. Over-sharding with TP wastes interconnect bandwidth; under-sharding runs out of memory.
The complete production picture, including managed alternatives, looks like this:

Decision recap
The most cost-effective inference is the workload that does not occupy an idle accelerator. Direct workloads to managed APIs and CPUs where appropriate, reserve accelerators for workloads that require them, tune tensor and data parallelism to keep them saturated, and avoid idle capacity. When one decision axis is the binding constraint, it typically resolves as follows:
| Binding constraint | Indicated direction |
|---|---|
| Control and customization | Self-hosted on ECS/EKS, the path that exposes the internals |
| Operational burden | Bedrock first, then SageMaker; avoid self-hosting |
| Cost at scale (steady, high volume) | Self-hosting or SageMaker on the right accelerator, saturated; validate with cost per token at your utilization |
| Latency objective | The appropriate accelerator with dedicated capacity; tune TP, chunked prefill, speculative decoding; Bedrock latency-optimized inference if managed |
| Data residency | Self-managed, or a managed service deployed in-VPC, rather than an external API |
To get started, prototype on Bedrock or a SageMaker endpoint before committing to self-hosting, instrument utilization from the outset with Amazon CloudWatch and Cost Explorer, and validate the choice by benchmarking one representative workload against your own p95 objective.
Caveats
Two caveats keep this grounded. First, results are directional: cost and throughput outcomes depend on your model, sequence length, batch size, latency target, and Region, and prices and quotas change over time, so the benchmarking and cost-modeling method is the transferable asset, not any specific figure. Second, managed versus self-hosted is a total-cost-of-ownership decision, not only an hourly-rate one: self-hosting can lower compute cost while raising engineering time, on-call burden, and opportunity cost. This guide also addresses the common path rather than every option (it does not cover EC2 self-managed serving without containers, SageMaker HyperPod, Bedrock Marketplace models, or edge inference).
References
- AWS Inferentia2 builds on Inferentia1 by delivering 4x throughput and 10x lower latency
- AWS Inferentia and Trainium deliver lowest cost to deploy Llama 3 in Amazon SageMaker JumpStart
- Amazon SageMaker adds new inference capabilities to reduce foundation model deployment costs and latency (inference components)
- CPU Inference and Orchestration (Amazon EKS Best Practices Guide)
- Navigating GPU Challenges: Cost Optimizing AI Workloads on AWS
Further reading
- How to run AI model inference with GPUs on Amazon EKS Auto Mode
- Guidance for Scalable Model Inference and Agentic AI on Amazon EKS
About the Author
Samar Singhal is a Sr. Solutions Architect at AWS specializing in Advanced AI Systems. With deep passion for Generative AI, Containers, and Kubernetes, Samar focuses on developing next-generation solutions that combine artificial intelligence with cloud-native technologies to solve complex enterprise challenges.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article