
7 Ways to Maximize Video Encoding Performance on AWS Graviton
AWS Graviton processors are purpose-built for video encoding workloads - but only a tuned build fully realizes that potential. This article walks through 7 practical optimizations, from compiler flags to SIMD acceleration, that unlock Graviton's full performance for FFmpeg and OpenCV pipelines.
7 Ways to Maximize FFmpeg Video Encoding Performance on AWS Graviton
AWS Graviton processors have become one of the most compelling options for video encoding in the cloud. Graviton4 achieves up to 73% more frames per second on H.265 encoding compared to Graviton2, according to AWS benchmarks , and consistently ranks at the top of price-performance comparisons across EC2 instance families.
But selecting the right instance is only half the job. Graviton's full performance is unlocked at the build layer — through the right compiler settings, codec configuration, and hardware acceleration. Without these, teams migrating from x86 can see significant regressions. With them, Graviton delivers the price-performance it's built for.
This article walks through what those optimizations are, why they matter, and exactly how to apply them.
Why Graviton Is a Natural Fit for Video Encoding
Video encoding is parallel and compute-heavy by nature — the same mathematical operations repeated millions of times per frame. That's exactly the kind of workload Graviton's architecture is designed for.
The key is SIMD — a capability that lets the processor handle many values in a single operation instead of one at a time. Every Graviton generation includes progressively more powerful SIMD support. Graviton2 supports NEON, which processes 16 pixel-channels per instruction. Graviton3 adds SVE with 256-bit vector registers — double the width of NEON — enabling significantly more data to be processed per clock cycle. Graviton4 adds SVE2, which runs at 128-bit vector width but introduces specialized media instructions — like
histseg — that accomplish in a single operation what previously required five separate steps. The advantage of SVE2 is algorithmic efficiency, not raw width, and it's why H.265 encoding sees such a pronounced improvement on Graviton4.AWS and the open source community have been contributing encoding optimizations that take advantage of these hardware capabilities. The H.265 codec library (x265) alone improved 16–34% between its 3.6 and 4.1 releases — and up to 2.3–2.5x compared to the older versions still shipped in most Linux distributions today.
The hardware is ready. Getting your build to use it is the whole game.
What Full Graviton Optimization Looks Like
A fully optimized Graviton FFmpeg build addresses three layers: the compiler, FFmpeg and its codec libraries, and — if your pipeline includes ML-based processing — the Python stack. Each layer contributes meaningfully to total throughput.
1. Tell the Compiler Which Processor It's Building For
This is the single highest-impact change you can make. By default, a generic ARM64 build targets the lowest common denominator — a baseline that covers phones, embedded devices, and cloud servers alike. It doesn't know it's running on Graviton, so it can't take advantage of Graviton's specific capabilities.
Setting the right CPU target flag tells the compiler exactly which Graviton generation it's building for, unlocking hardware-specific scheduling and instruction selection that a generic build simply cannot do.
For teams building container images on x86 CI hosts and deploying to Graviton, this is the most common source of misconfiguration — the build target needs to match the Graviton generation, not just "ARM64."
Compiler choice also matters. For H.265 encoding specifically, using Clang instead of GCC delivers approximately 11% better throughput on Graviton4. The latest codec libraries are written in a way that Clang handles more effectively on ARM.
2. Use the AWS Official FFmpeg Builder
AWS maintains a purpose-built FFmpeg builder for Graviton, available in the aws-graviton-getting-started repository. This is the fastest path to a production-ready, fully optimized binary.
The builder handles all the complexity: the right settings for each codec (H.264, H.265, AV1), the latest codec library versions that include Graviton-specific improvements, and the correct hardware acceleration settings. It outputs a package ready to install directly in your Dockerfile.
FFmpeg binaries from standard Linux distribution packages do not include Graviton-specific optimizations — they use generic ARM64 paths that leave the most impactful hardware capabilities unused.
3. Swap in ARM-Optimized Math Libraries for Python
If your pipeline does any Python-based processing alongside encoding — resizing frames, preparing data for an ML model, or any numerical work with NumPy — those operations are likely running on a math library that was not built for Graviton.
Arm Performance Libraries is a free, Graviton-optimized replacement that speeds up exactly these operations with no code changes required. You swap the library, reinstall NumPy, and that's it. Since it's updated regularly, always grab the latest version from the Arm Developer portal when building your container.
4. Control Thread Usage When Running PyTorch Alongside FFmpeg
On Graviton instances with high core counts, thread management matters in ways that don't surface on smaller x86 instances. When PyTorch inference runs in parallel with FFmpeg encoding — a common pattern in ML-enhanced video pipelines — PyTorch's default behavior can consume more CPU threads than intended, competing with the encoder and degrading both workloads at once.
The fix is two lines of code at initialization: one to limit PyTorch to a single thread, and one to skip unnecessary computation during inference. For pipelines where this contention is present, the improvement can be dramatic.
5. Upgrade Your Graviton Generation
Each Graviton generation brings meaningful encoding improvements through more capable hardware. The jump is particularly significant for H.265:
| Instance | Generation | vs. Graviton2 | SIMD capability |
|---|---|---|---|
| c6g | Neoverse N1 (Graviton2) | Baseline | NEON — 128-bit vectors |
| c7g | Neoverse V1 (Graviton3) | Improvement over Graviton2 | NEON + SVE — 256-bit vectors |
| c8g | Neoverse V2 (Graviton4) | +12–15% vs. Graviton3 / +73% vs. Graviton2 on H.265 | NEON + SVE2 — 128-bit vectors with specialized media instructions |
Source: AWS Video Encoding on Graviton in 2025 — frames per second, fully-loaded parallel encoding runs.
Migrating between generations is a launch template or node pool change — no application code to modify. Just update the instance family (c7g for Graviton3, c8g for Graviton4) and the CPU target flag in your build to match.
Tip — combine with EC2 Spot Instances: Video encoding is typically asynchronous and can tolerate interruptions — a natural fit for Spot. Combining Graviton's price-performance with Spot discounts of up to 90% off On-Demand pricing makes for a very cost-efficient encoding pipeline. Use a mix of c7g and c8g instance types in your Spot configuration to maximize availability.
6. Use the Right OS Base Image
The AWS FFmpeg builder supports both Ubuntu 22.04 and Amazon Linux 2023 (AL2023). AL2023 ships with a newer kernel that includes improvements to how it schedules work on ARM processors, which benefits encoding workloads. If you're not already on AL2023, it's worth running a canary to measure the impact on your specific pipeline.
7. Verify You're Running Native ARM64
Before benchmarking any of the above, confirm that your Python environment and installed packages are truly running on ARM64 — not x86 packages running under emulation. This is a surprisingly common issue in containerized pipelines, particularly when package caches are built on x86 hosts and reused without validation.
1
2
3
4
5
6
7
# Confirm Python is native ARM64
python3 -c "import platform; print(platform.machine())"
# Expected: aarch64
# Confirm no x86 packages are installed
pip list -v | grep x86_64
# Expected: no resultsIf any x86 packages appear, reinstall them from ARM64 wheels before proceeding.
The Build Flags Reference
Every optimization above maps to specific flags. Here they are, with what each one does and how to set it.
Compiler flags — set before any build command
| Flag | Why it matters | How to set it |
|---|---|---|
-mcpu | Targets the exact Graviton processor, enabling hardware-specific optimizations the compiler otherwise cannot apply. The single highest-impact flag — without it, you get a generic ARM64 build that misses most Graviton improvements. | export CFLAGS="-mcpu=neoverse-v1" for Graviton3, -mcpu=neoverse-v2 for Graviton4. Set in both CFLAGS and CXXFLAGS. |
-moutline-atomics | Ensures thread-safe operations work correctly across different kernel versions — important in containerized environments. | Add alongside -mcpu in CFLAGS and CXXFLAGS. |
-O3 | Tells the compiler to apply more aggressive optimizations. Most distribution builds use a more conservative setting that leaves throughput on the table. | Include in CFLAGS: -mcpu=neoverse-v1 -moutline-atomics -O3 |
AWS FFmpeg builder flags
| Flag | Why it matters | How to set it |
|---|---|---|
--target-platform | Activates Graviton-specific hardware acceleration for all codecs. Without this, the output binary uses generic ARM64 paths that miss the most impactful optimizations. | --target-platform graviton3 for c7g instances, graviton4 for c8g. |
--target-distro | Produces a package in the right format for your base image, ready to install directly in your Dockerfile. | --target-distro jammy for Ubuntu 22.04, al2023 for Amazon Linux 2023. |
--compiler | Clang produces approximately 11% better H.265 throughput than GCC on Graviton4 due to how it handles the latest codec libraries. | --compiler clang-latest |
FFmpeg configuration flags (verify with
ffmpeg -version)| Flag | Why it matters | How to set it |
|---|---|---|
--enable-neon | Activates ARM hardware acceleration across FFmpeg's filters and codecs. Absent in standard distribution builds. | Set automatically by the AWS FFmpeg builder. Confirm it appears in ffmpeg -version. |
--enable-libx265 | Enables H.265 encoding. The codec version matters — the AWS builder uses the latest version, which includes significant Graviton improvements. Distribution packages often ship older versions. | Set by the AWS FFmpeg builder. Confirm with ffmpeg -version. |
--enable-libsvtav1 | Enables AV1 encoding. AWS has contributed Graviton-specific optimizations to this codec. | Set by the AWS FFmpeg builder when AV1 is configured. |
--enable-libx264 | Enables H.264 encoding, which also benefits from Graviton-specific compiler tuning. | Set by the AWS FFmpeg builder when H.264 is configured. |
Python stack
| Setting | Why it matters | How to set it |
|---|---|---|
| Arm Performance Libraries | The default math library behind NumPy is not optimized for Graviton. This free replacement speeds up numerical operations in frame processing — no code changes needed, just swap the library and reinstall NumPy. | export BLAS=/opt/arm/armpl/lib/libarmpl.so then reinstall NumPy: pip install numpy --no-binary numpy --force-reinstall. Check the Arm Developer portal for the latest version. |
torch.set_num_threads(1) | Prevents PyTorch from consuming more CPU threads than intended when running alongside FFmpeg, avoiding contention that degrades both workloads. | Call at pipeline initialization, before any inference. |
torch.no_grad() | Skips unnecessary computation during inference, reducing overhead for pipelines that only run predictions, not training. | Wrap all inference calls: with torch.no_grad(): result = model(input_tensor) |
Validating Your Configuration
Two quick checks before benchmarking:
Is FFmpeg Graviton-tuned?
1
2
ffmpeg -version
# Look for: --enable-neon in the configuration lineIs NumPy using Arm Performance Libraries?
1
2
python3 -c "import numpy as np; np.show_config()"
# Look for: armplThen benchmark end-to-end:
1
time ffmpeg -i input_4k.mp4 -c:v libx265 -preset medium -t 60 -y /dev/nullRecord a baseline before any changes, apply optimizations one at a time, and re-run after each. Deploy to a 10% traffic canary and monitor for at least 30 minutes before promoting to production.
What This Looks Like in Practice
When teams apply these optimizations together, the results align with what AWS's own benchmarks show. In AWS's 2025 price-performance analysis across EC2 instance families for FFmpeg workloads, Graviton-powered instances took three of the top five spots — including Graviton2, which remains competitive for H.264 workloads despite being the oldest generation. Graviton4 ranked first overall.
The biggest gains come from the compiler and codec configuration — everything else builds on top. Applied together on a pipeline that started from a generic ARM64 build, the cumulative effect transforms Graviton from a good instance choice into an exceptional one.
Key Takeaways
Graviton's hardware is built for encoding. Each generation adds more capable SIMD processing: NEON at 128-bit on Graviton2, SVE at 256-bit on Graviton3, and SVE2 on Graviton4 — which trades raw width for specialized media instructions that make H.265 encoding dramatically more efficient. AWS actively contributes optimizations to FFmpeg, x265, and SVT-AV1 to take advantage of each generation.
The AWS FFmpeg Builder is the fastest path to a correctly configured binary. It handles codec configuration, compiler selection, and library versions in one step — no manual flag hunting required.
The 73% H.265 improvement is real and well-sourced. It reflects frames-per-second throughput on Graviton4 vs. Graviton2, from AWS's April 2025 benchmark analysis , using fully-loaded parallel encoding runs.
Thread management matters when mixing encoding and inference. A couple of initialization settings prevent PyTorch from competing with FFmpeg for CPU resources on high-core-count instances.
Graviton on Spot is a natural fit for video encoding. Asynchronous, interruptible workloads get the full benefit of both Graviton's price-performance and Spot discounts.
Share Your Results
Tried these optimizations on your pipeline? We'd love to hear what you found — share your benchmarks and questions on AWS re:Post or open an issue on the AWS Graviton Getting Started GitHub repository. Your real-world results help the community and feed back into the open source work happening on x265, SVT-AV1, and FFmpeg.
References
| Resource | Link |
|---|---|
| AWS Graviton FFmpeg Builder | View on GitHub |
| Video Encoding on Graviton in 2025 | AWS Open Source Blog |
| Optimized Video Encoding with FFmpeg on AWS Graviton (2022) | AWS Open Source Blog |
| AWS Graviton C/C++ Optimization Guide | aws.github.io |
| AWS Graviton SIMD & Vectorization Guide | aws.github.io |
| Arm Performance Libraries (latest) | developer.arm.com |
| Amazon EC2 Spot Instances | aws.amazon.com |
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article