AWS Builder Center

Batch LLM Inference at Scale on AWS

A step-by-step methodology to maximize performance for batch LLM inference at scale on AWS, backed by real Llama-3.1-8B benchmarks on H200s, plus a practical guide to choosing between Amazon Bedrock, SageMaker, AWS Batch, and self-managed EKS for offline workloads.

Authors: Hossam Basudan (Sr. GTM SSA, GenAI/ML, MENAT) and Oussama Kandakji (Sr. GTM SSA, AI/ML GenAI, France)

Overview

Large Language Model (LLM) inference workloads fall into two categories: real-time (low-latency, interactive) and batch (high-throughput, offline). While real-time inference dominates conversational AI use cases, batch inference is critical for scenarios where throughput and cost-efficiency matter more than latency like document summarization, content generation pipelines, data enrichment, evaluation harnesses, and embedding generation
This guide covers architectural patterns, service options, and best practices for running batch LLM inference at scale on AWS.

When to Use Batch Inference

Use CaseExampleWhy Batch?
Document processingSummarize 1M legal contractsHigh volume, no real-time requirement
Data enrichmentClassify/tag product catalog entriesCost-sensitive, parallelizable
Evaluation & benchmarkingRun evals across model versionsReproducibility, throughput
Embedding generationGenerate vector embeddings for RAGBulk pre-computation
Content generationProduce marketing copy variantsOffline, review-before-publish
Synthetic dataGenerate training data at scaleVolume-driven, quality-filtered

Methodology: Maximizing Performance on a Given Instance

Getting peak batch throughput out of a fixed instance is a repeatable, measurable process that applies regardless of GPU, model, or serving engine. The goal is to find and then saturate whichever resource limits your workload. LLM inference has two phases with opposite profiles: prefill (processing the input prompt) is usually compute-bound, while decode (generating output tokens) is usually memory-bandwidth-bound. Which phase dominates depends on your input-to-output token ratio. Long inputs with short outputs are prefill-heavy, while short inputs with long outputs are decode-heavy. These two phases might also interfere with each other when we use batching because an inference step in a batch can be a prefill step for input or a decode step for an output. The diagram below shows the anatomy of inference requests for LLMs in the case of continuous batching  in vLLM.
The steps below identify the binding constraint first, then systematically close the gap against the corresponding ceiling.

Step 1: Shape the Workload

Input and output length are major throughput levers, and they are properties of the request rather than the hardware. Fix them early: sequence length sets the per-sequence memory cost that the KV cache and concurrency steps depend on, and it feeds the prefill/decode classification in Step 2.
Start by compressing inputs, since shorter prompts leave more memory for concurrent sequences; trim boilerplate and chunk long documents. Cap outputs at the minimum the task needs, because shorter outputs raise the request rate. Where you can, batch similar-length requests together (mixing lengths wastes batch slots), and cache shared prefixes so repeated system prompts across the batch reuse computation when the engine supports it.

Step 2: Establish the Hardware Ceiling

Before tuning anything, compute the roofline for your model and accelerator and classify your workload, so you know what "good" looks like.
Pull the accelerator's memory bandwidth and peak compute from its spec sheet, and use your workload's input-to-output token ratio (fixed in Step 1) to gauge how much time goes to prefill versus decode. Then run high-throughput benchmarks at increasing concurrency to find the token throughput ceiling of the model on the chosen accelerator.

Step 3: Reduce the Model's Memory Footprint

Shrinking the weights is often the highest-leverage single change. A smaller footprint frees memory for KV cache and increases effective bandwidth per token.
Apply quantization (for example FP8 or NVFP4) appropriate to your quality tolerance, and prefer pre-quantized checkpoints where available so scale factors are computed offline instead of at runtime. Always validate output quality afterward, since the right precision is the lowest one that still meets your accuracy bar.

Step 4: Choose the Right Parallelism Layout

On a multi-accelerator node, how you split the model matters as much as whether you split it. Together with the footprint from Step 3, this sets the free memory available per model instance, the supply side of the KV cache budget.
For models that fit in a single device's memory, run independent replicas (one per device) rather than sharding. This avoids cross-device communication overhead and scales near-linearly, with each replica processing its own partition of the input to maximize aggregate node throughput. Reserve tensor or pipeline parallelism for models too large to fit on a single device.

Step 5: Maximize KV Cache Memory

KV cache capacity determines how many sequences can run concurrently, which directly sets achievable batch size.
Allocate as much device memory to the KV cache as stability allows. Smaller weights (Step 3) and the right parallelism layout (Step 4) leave more room for it, raising the concurrent-sequence limit, so confirm the engine isn't reserving excess memory for unused features or oversized buffers. You can also activate KV cache offloading to move unused KV blocks to cheaper storage and memory layers instead of deleting them outright, which maximizes prefix cache hits and improves prefill performance.

Step 6: Apply prefill and decode optimizations

To optimize prefill and decode and reduce interference between them, you can apply a few optimizations.
Speculative decoding  accelerates token generation, whether through speculator heads (EAGLE 3), a smaller draft model, or model-less techniques like suffix-based decoding. Chunked prefill improves inter-token latency by breaking prefill into chunks that interfere less with adjacent decode steps, which helps most when prefill steps (large inputs) are long relative to decode steps. And beyond KV cache offloading, enabling prefix caching maximizes prefill cache hits and further reduces prefill latency.

Step 7: Measure, Compare, Iterate

Tuning is empirical. Re-benchmark after each change and compare against the Step 2 ceiling.
Measure throughput (tokens/s and requests/s) at your target input and output lengths, and profile the accelerator to confirm time is spent in core compute kernels rather than overhead. If efficiency stalls well below your target, revisit concurrency (Step 6) and KV cache pressure (Step 5) before blaming the kernels.
Bottom line: You cannot beat the roofline. You can only approach whichever ceiling binds your workload. Reduce the memory footprint, fill the freed memory with KV cache, and push concurrency until the binding resource (memory bandwidth for decode-heavy work, compute for prefill-heavy work) is the limit.

Real-World Throughput Benchmarks

The following benchmarks apply the methodology above to a concrete workload: Meta Llama-3.1-8B-Instruct on a single p5en node (8× NVIDIA H200 141GB) running vLLM v0.23.0. Each step below corresponds to the methodology section and shows the decisions made and their measured impact.

Step 1: Workload Definition

The target workload is a high-volume batch classification/summarization task: roughly 1,000 input tokens on average (customer prompts with context) and around 400 output tokens (structured responses), a 2.5:1 input-to-output ratio. That makes it decode-heavy, so memory bandwidth is the binding constraint at high concurrency.

Step 2: Hardware Ceiling (Roofline  Analysis)

Before tuning, we compute the theoretical ceiling for H200 + Llama-8B:
H200 SpecValue
HBM3e bandwidth4,800 GB/s
FP8 compute1,979 TFLOPS
HBM capacity141 GB
Memory budget per GPU (FP8 pre-quantized Llama-8B):
ComponentSize
Model weights (FP8)8.0 GB
Available for KV cache~131 GB
KV cache per token (BF16)128 KB
Max tokens in KV cache~1,023k
Max concurrent sequences (1.4k tokens each)~731
Is decode compute-bound or memory-bandwidth-bound?
The roofline model classifies a workload by comparing its arithmetic intensity (the compute performed per byte read from memory, in FLOP/byte) against the accelerator's ridge point, defined as Peak Compute ÷ HBM Bandwidth. Below the ridge point a workload is memory-bandwidth-bound (throughput is limited by bandwidth, not compute); above it, compute-bound.
Ridge point (H200). Although weights are stored in FP8, the KV cache and activations remain in BF16, so the effective compute bound is the BF16 dense rate (~989 TFLOPS): 989 TFLOPS ÷ 4,800 GB/s ≈ 206 FLOP/byte.
Arithmetic intensity (this workload). Each decode step performs 2 × 8 B parameters × 731 sequences ≈ 11,700 GFLOP of compute, and reads the full model weights (8 GB) plus the KV cache for all active sequences (731 × 1,400 tokens × 128 KB per token ≈ 131 GB), ≈ 139 GB in total. That gives 11,700 GFLOP ÷ 139 GB ≈ 84 FLOP/byte. Arithmetic intensity grows with batch size but is bounded, so even at the largest batch the device can hold (731) it stays at 84 FLOP/byte.
Because 84 FLOP/byte is well below the 206 FLOP/byte ridge point at every batch size in the feasible range, decode is memory-bandwidth-bound. Adding more compute won't raise throughput; only more memory bandwidth will.
Roofline throughput by batch size:
Having established that decode is memory-bandwidth-bound, we estimate its throughput ceiling from the memory-bandwidth model. Each decode step reads all model weights plus the KV cache for every active sequence, producing one token per sequence:
Roofline tok/s = (Batch Size × HBM Bandwidth) ÷ (Model Weights + Batch Size × Avg Seq Length × KV per Token)
Example (batch size = 1): With 4,800 GB/s bandwidth and 8 GB of FP8 weights, the KV cache contribution is negligible at batch 1, so: (1 × 4,800) ÷ 8 ≈ 587 tok/s. At batch size 731 the KV cache dominates: (731 × 4,800) ÷ (8 + 731 × 1,400 × 0.000128 GB) ≈ 25,240 tok/s.
Batch SizeRoofline output tok/s (per GPU)
1587
6415,779
25622,807
51224,637
73125,240
The ceiling for this workload is 25,240 tok/s per GPU, reached at batch size 731. The table stops at 731 because that is the memory-capacity limit, not an arbitrary cutoff: at ~1,400 tokens per sequence, 731 concurrent sequences consume the ~131 GB of HBM left for the KV cache after weights (the ~731 max concurrent sequences derived in the memory budget above). Larger batches don't fit, so additional requests must queue rather than run concurrently, which makes 731 the memory-bound concurrency ceiling for this sequence length. Throughput also flattens as it nears this point (512 to 731 adds only ~2.4%) because the KV cache increasingly dominates the denominator, so 731 is both the memory ceiling and the point of diminishing returns.

Step 3: Reduce Memory Footprint (Quantization)

Fixed parameters: avg input = 1,000 tokens, avg output = 400 tokens
ConfigurationPeak req/s (8 GPUs)Per-GPU output tok/svs. BF16
BF16 (baseline)255.412,770baseline
FP8 on-the-fly quantization287.414,370+12.5%
Pre-quantized FP8 (ModelOpt)330.216,510+29.3%
Pre-quantized FP8 checkpoints deliver the best throughput, 29% above BF16 and 15% above on-the-fly FP8. The pre-computed scale factors eliminate runtime calibration overhead. FP8 also halves the weight footprint (16 GB → 8 GB), freeing memory for KV cache in Step 5.

Step 4: Parallelism Layout

Llama-8B in FP8 fits in 8 GB, well within a single H200's 141 GB. We run 8 independent replicas (one per GPU, tensor parallelism = 1) to avoid cross-device communication overhead and achieve near-linear scaling across all GPUs.

Step 5: KV Cache Capacity

With 8 GB of weights, each GPU has ~131 GB available for KV cache. At 128 KB per token (BF16 KV), this supports ~1,023k tokens or ~731 concurrent sequences of 1.4k tokens each. vLLM's --max-num-seqs 731 and --max-num-batched-tokens 32768 were tuned to fill this budget without OOM.

Step 6: Apply prefill and decode optimizations

Here are some configurations to apply prefill and decode optimizations:
ParameterImpact
--enable-chunked-prefillReduces inter-token latency
--enable-prefix-cachingAccelerates/eliminates prefill according to cache hit
Speculative decoding (see configs below)Accelerates decode steps depending on acceptance rate
Speculative decoding (Suffix decoding):
1
2
3
4
--speculative-config {
"method": "suffix",
"num_speculative_tokens": 8
}
Speculative decoding (Eagle3):
1
2
3
4
5
--speculative-config {
"method": "eagle3",
"model": "RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3",
"num_speculative_tokens": 2
}

Step 7: Measure and Iterate

Result: 16,510 tok/s per GPU = 65% of roofline (25,240 tok/s)
(GPU kernel profiling):
Operation% GPU TimeKernel
FP8 GEMMs (MatMul)59.6%cutlass_3x_gemm_sm90_fp8
Flash Attention v329.1%FlashAttnFwdSm90
Fused elementwise (SiLU, scale, cast)2.9%triton_poi_fused_*
KV cache update2.7%reshape_and_cache_flash_kernel
Sampling0.5%TopPSamplingFromProbKernel
Other5.2%Prefill GEMMs, normalization
89% of GPU time is spent in just two operations, both using state-of-the-art SM90 Cutlass/Flash Attention kernels. There is no faster implementation available. The gap between measured and roofline is accounted for by prefill overhead, scheduling, and memory management.

Sensitivity Analysis

With the optimized configuration established (pre-quantized FP8, 8 replicas, max KV cache), we sweep the workload parameters from Step 1 to quantify their impact:

Input Length Sensitivity

Fixed parameters: pre-quantized FP8, avg output = 400 tokens
Avg Input TokensPeak req/s (1 node)Estimated Time for 10M Requests (1 node)
5004506.5 hours
1,0003308.5 hours
2,00015018.5 hours
3,0005055 hours
Reducing from 3k to 1k tokens yields 6× improvement. KV cache memory scales linearly with sequence length, so longer inputs leave fewer slots for concurrent sequences.

Output Length Sensitivity

Fixed parameters: pre-quantized FP8, avg input = 1,000 tokens
Output TokensPeak req/s (1 node)Output tok/s (1 node)
128880112,640
200620124,000
400330132,083
Halving output from 400 to 200 tokens nearly doubles request throughput. Note that total output tok/s remains relatively stable. Shorter outputs simply complete faster, freeing batch slots sooner.

Sizing Your Deployment

Use these reference numbers (Llama-3.1-8B, FP8, p5en nodes) to estimate infrastructure needs:

Theoretical Maximum Throughput (1× p5en.48xlarge = 8 GPUs)

MetricPer GPU1 Node (8 GPU)
Roofline (theoretical max)25,240 tok/s201,920 tok/s
Achievable at 65% efficiency16,510 tok/s132,080 tok/s
Scale linearly: Throughput scales nearly linearly with node count for batch workloads (each GPU runs an independent replica). Add nodes to reduce wall-clock time proportionally.

Architecture Patterns

Pattern 1: Amazon Bedrock Batch Inference

Amazon Bedrock provides a fully managed batch inference API that requires zero infrastructure management.
How it works:
  1. Prepare a JSONL input file with your prompts in S3
  2. Submit a batch inference job via the CreateModelInvocationJob API
  3. Bedrock processes all records asynchronously
  4. Results are written to your specified S3 output location
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
import boto3

bedrock = boto3.client("bedrock")

response = bedrock.create_model_invocation_job(
jobName="contract-summarization-batch",
modelId="anthropic.claude-sonnet-4-20250514",
roleArn="arn:aws:iam::123456789012:role/BedrockBatchRole",
inputDataConfig={
"s3InputDataConfig": {
"s3Uri": "s3://my-bucket/batch-input/contracts.jsonl",
"s3InputFormat": "JSONL"
}
},
outputDataConfig={
"s3OutputDataConfig": {
"s3Uri": "s3://my-bucket/batch-output/"
}
}
)
Benefits: it requires no infrastructure to manage, handles scaling and throughput optimization automatically, and can deliver up to 50% cost savings versus on-demand inference.

Pattern 2: SageMaker Batch Transform

SageMaker Batch Transform provides managed batch inference with control over instance types and scaling.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
import sagemaker
from sagemaker.transformer import Transformer

transformer = Transformer(
model_name="my-llm-endpoint-model",
instance_count=10,
instance_type="ml.g5.48xlarge",
output_path="s3://my-bucket/batch-output/",
max_payload=6, # MB
strategy="MultiRecord",
assemble_with="Line"
)

transformer.transform(
data="s3://my-bucket/batch-input/",
content_type="application/jsonlines",
split_type="Line",
join_source="Input"
)
When to choose SageMaker Batch Transform: you need specific instance types (such as p5e, p4d, or Inf2), or you have complex pre- and post-processing logic.

Pattern 3: Self-Managed on EKS/EC2

For maximum control, deploy inference servers on Amazon EKS or EC2 and drive them with a batch client that reads the input from S3, sends requests to the inference endpoint, and writes the results back to S3.
Architecture:
Key components: The inference engine can be vLLM, TGI (Text Generation Inference), or TensorRT-LLM, run as a Kubernetes Deployment behind a Service that exposes an OpenAI-compatible endpoint. A batch driver reads the S3 input, partitions it, and streams requests to that endpoint concurrently, writing the results back to S3. Karpenter provisions GPU nodes while a Horizontal Pod Autoscaler scales replicas on GPU utilization or throughput. For reliability, the driver checkpoints completed records and retries failures, so a re-run resumes from where it stopped.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
# Karpenter NodePool for batch inference
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: batch-inference-gpu
spec:
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: gpu-nodes
requirements:
- key: node.kubernetes.io/instance-type
operator: In
values: ["p5e.48xlarge", "p4d.24xlarge", "g5.48xlarge"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
taints:
- key: nvidia.com/gpu
effect: NoSchedule
When to choose self-managed: you need multi-model serving or model routing, or custom batching logic such as priority scheduling and SLA tiers.

Pattern 4: AWS Batch for GPU Inference Jobs

AWS Batch is a fully managed job scheduler purpose-built for batch workloads. It handles job queuing, compute provisioning (including GPU instances), and automatic scaling, without the operational overhead of managing Kubernetes clusters.
How it works:
  1. Define a job definition with your inference container (vLLM, TGI, or custom)
  2. Submit array jobs, one per input partition (e.g., 1,000 JSONL chunks)
  3. AWS Batch provisions GPU instances, schedules jobs, and handles retries
  4. Results written to S3 upon completion
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
{
"jobDefinitionName": "llm-batch-inference",
"type": "container",
"containerProperties": {
"image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/vllm-batch:latest",
"resourceRequirements": [
{ "type": "GPU", "value": "8" },
{ "type": "VCPU", "value": "192" },
{ "type": "MEMORY", "value": "1400000" }
],
"command": [
"python", "run_batch.py",
"--input", "Ref::s3_input",
"--output", "Ref::s3_output",
"--model", "meta-llama/Llama-3.1-8B-Instruct"
],
"environment": [
{ "name": "VLLM_QUANTIZATION", "value": "fp8" }
]
},
"platformCapabilities": ["EC2"]
}
Submit an array job (1,000 partitions):
1
2
3
4
5
6
7
8
9
aws batch submit-job \
--job-name "llm-inference-10M-records" \
--job-queue "gpu-batch-queue" \
--job-definition "llm-batch-inference" \
--array-properties size=1000 \
--parameters '{
"s3_input": "s3://my-bucket/batch-input/",
"s3_output": "s3://my-bucket/batch-output/"
}'
Compute environment with Spot + On-Demand fallback:
1
2
3
4
5
6
7
8
9
10
11
12
{
"computeEnvironmentName": "gpu-inference-env",
"type": "MANAGED",
"computeResources": {
"type": "SPOT",
"allocationStrategy": "SPOT_CAPACITY_OPTIMIZED",
"instanceTypes": ["p5en.48xlarge", "p5e.48xlarge", "g5.48xlarge"],
"maxvCpus": 3072,
"minvCpus": 0,
"spotIamFleetRole": "arn:aws:iam::123456789012:role/AWSBatchSpotFleetRole"
}
}
When to choose AWS Batch: you want managed scheduling without Kubernetes overhead, or your work is embarrassingly parallel (array jobs with independent partitions), or you need a mix of Spot and On-Demand with automatic fallback. It also fits when you have job dependencies and multi-stage pipelines (job A feeding job B), need automatic retry with configurable backoff, or want fair-share scheduling across teams and workloads.

Choosing the Right Approach

Summary

Batch LLM inference at scale on AWS offers several paths depending on your requirements. Start with Bedrock Batch for simplicity and managed scaling with foundation models, or SageMaker Batch Transform for a managed experience with a simple SDK. Reach for AWS Batch when you want managed GPU job scheduling with built-in Spot and retry, and build on EKS when you need more control and custom batching logic.
Regardless of the approach, apply cost optimization (spot, quantization, right-sizing), implement checkpointing for reliability, and monitor throughput metrics so your batch pipelines deliver results efficiently at scale.

Related Resources

Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article