
Rift Rewind - Clean Up Match Data with SageMaker
Use the visual ETL feature of SageMaker to transform the match data from the Riot API!
Rift Rewind Developer Challenge #2 - Day 2
NOTE: Due to comments notifying us that you are removed from free tier due to IAM Identity Center Requirements, you only need to read this walk through to complete the task, comment done once you have read through the blog and you will be marked complete!
Welcome back to Challenge #2 - Day 2! Yesterday, you learned about fetching data from the Riot API using AWS Lambda and storing it in Amazon S3. Today, we’ll dive into quickly and easily processing that data using Amazon SageMaker Unified Studio! Think of it as unlocking a legendary item that makes orchestrating data tasks as easy as last-hitting minions! 🎮
📊 Why Process Data?
Data processing is basically the mid-game phase between getting the raw API data and making awesome insights!Some examples of when you might want to use data processing:
- Cleaning data (removing those pesky missing values like dodged games)
- Transforming formats (converting timestamps to actual dates, like knowing exactly when Baron was slain)
- Aggregating statistics (calculating your average CS per minute or KDA across multiple matches)
- Filtering (focusing only on your main role, like a true OTP Teemo)
- Joining datasets (combining champion stats with match history, like pairing Yuumi with her perfect ADC)
🎯 What You Will Accomplish Today
✅ Set up a SageMaker domain✅ Create a SageMaker Unified Studio project✅ Set up a workflow environment✅ Create a data processing job✅ Query your Riot API data from Day 1’s challenge
🔮 What You Will Need
- An AWS account 💻
- An email address 📧
- Data in your S3 bucket from yesterday’s challenge 🪣
- 10 mins for setup + 30 mins for setup to load + 20 mins for challenge ⏰
🌿 What’s Amazon SageMaker?
Amazon SageMaker is the ultimate all-in-one training arena for your machine learning adventures – think of it as the Summoner's Rift for data scientists! Just as League gives you everything you need to battle with champions, SageMaker provides all the tools to build, train, and deploy your ML models without the frustrating grind!Ready to queue up for your first ML victory? SageMaker's waiting in the lobby! GLHF! 😄
🧪 What's SageMaker Unified Studio?
Amazon SageMaker Unified Studio is your all-in-one data and AI development hub where you can access all your organization's data and tackle it with the perfect tools for your mission. It's like having a complete inventory of weapons and spells for any data challenge you encounter! 🗡️Today, we’ll be learning about the visual workflows builder that lets you orchestrate your data tasks without writing code manually. With visual workflows, you can:
- Create workflows using an intuitive drag-and-drop interface (easier than Yuumi's kit)
- Connect notebooks, queries, and data processing jobs graphically (chain them like a perfect combo)
- Automatically convert your visual designs to Python DAG definitions for Airflow (automatic CS improvement!)
Now that you know what SageMaker Unified Studio is, let’s dive in and start working on it!
Setting Up SageMaker Unified Studio (30-40 mins)
Important Note: To work with SageMaker Unified Studio, there are a couple steps we need to follow to set up our workflow environment. It’s pretty straightforward to set up, but AWS can take 20-30 minutes to create your workflow environment (the final step). You should knock out these three steps, then grab a bite to eat, or play a LoL match while you wait! ⏳
🌐 Setup Step 1: Create a SageMaker Unified Studio domain (~5 min)
The first thing we’ll need to do to set up SageMaker is to create a Unified Studio domain.Domains are like private League of Legends servers where your team can develop machine learning strategies together—they provide secure, isolated environments where data scientists can collaborate on models using shared resources.Similar to how League's Summoner's Rift has different lanes and jungle areas for specialized champions, a SageMaker Unified Studio domain organizes your machine learning workspace with different user profiles, applications, and computing resources all managed under one administrative umbrella.
- Open up your AWS Console and search Amazon SageMaker.
- On the SageMaker homepage, click Create a Unified Studio domain.
- Keep the Quick setup selected. For the VPC, choose Create VPC.
- This will take you to AWS CloudFormation. CloudFormation is like using a blueprint for your entire gaming setup—instead of manually installing League of Legends, configuring your graphics settings, and setting up Discord, you create one template that automatically builds everything exactly how you want it. Scroll to the bottom of the page and click Create stack.
- Go back to your SageMaker domain setup. Scroll down and expand the Quick setup settings. Scroll down within the Quick setup settings, and you’ll see a section where you can add the SageMakerUnifiedStudioVPC that CloudFormation just created for you! Make sure to also add the 3 private subnets.
Note: You may need to click the blue refresh button if the VPC isn’t appearing yet.
6. On the Create IAM Identity Center user page, create an SSO user (account with IAM Identity Center) or select an existing SSO user to log in to the Amazon SageMaker Unified Studio. IAM roles that create the Amazon SageMaker unified domains cannot log in to the Amazon SageMaker Unified Studio. The SSO selected here is used as the administrator in the Amazon SageMaker Unified Studio. Finally, click Create domain.
- After a few minutes, you should receive an email to the address you provided, granting you access to the domain. Click on the link provided in the email and follow the steps to register as an SSO user. Make sure to note down the email and password you choose!
Congratulations, you’ve created your SageMaker domain!
🗂️ Setup Step 2: Create a SageMaker Unified Studio project (~5 min)
- Next up, we’ll need to make a SageMaker Unified Studio project. SageMaker Studio projects are like your League of Legends premade team compositions, but for data science! Just as you'd gather your favorite champions with complementary abilities to dominate the Rift, Studio projects let you bundle all your machine learning resources—notebooks, datasets, and models—into one organized team that works together perfectly!
- First, you’ll need to login to your SageMaker Unified Studio domain using the email and password you chose in the prior step.
- After you’ve logged in, you should see the SageMaker Unified Studio domain homepage. Select Create project.
- Give your project a name. Enter something like
rift-rewind-project-[your-name]. In Description, enter “Riot API data processing”. For Project Profile, select All capabilities. - Select Continue, then scroll to the bottom of the second page and press Continue again. Finally, on the third page, review all your configurations and click Create Project.
Note: Project creation can take several minutes to complete.
Once the project is finished being created, you’ll be put on the Project overview screen.
⚙️ Setup Step 3: Setup a Workflow Environment (<1 min setup + 20-30 mins wait time)
The final step to setting up SageMaker Unified Studio is to create a Workflow Environment. After this, we’ll be able to start processing our data! ✨
- In the menu on the left-hand side, click Compute. This takes you to the Compute page.
- On the Compute page, select the Workflow environments tab on the right side. Then choose Create.
- In the Create workflow environment window, choose Create workflow environment.
Note: Workflow environment creation can take several minutes (20-30) to complete.
Congratulations, you’ve successfully created your SageMaker Unified Studio environment!
Processing Your Data Using Visual ETL (25 mins)
🪪 Step 1: Assign IAM Permissions to Access Your S3 Bucket (~5 mins)
Remember how we created an IAM role yesterday for our Lambda function, allowing it to access your S3 bucket? This let the function write the match data it got from the Riot API into your S3 bucket. We’ll need to do the same thing today, only this time, we’ll create an IAM role allowing SageMaker to access it!
Creating the Policy
- Go to IAM - Search “IAM” in the AWS console
- Click Policies > Create Policy
- In the JSON editor, paste the following policy:
{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["s3:ListBucket"],"Resource": "arn:aws:s3:::[YOUR BUCKET NAME]"},{"Effect": "Allow","Action": ["s3:GetObject"],"Resource": "arn:aws:s3:::[YOUR BUCKET NAME]/*"}]}5. Name your policy something descriptive, like
SageMakerS3BucketAccess-rift-rewind- Click Create Policy
Attach Policy to Role
- Click Roles in the IAM menu on the left side
- Search for
datazone_usr_role_64d46teve978dl_dkx17oztq4zq21 - Select the role and click Add permissions > Attach policies
- Find and attach the policy you just created!
🪄 Step 2: Create a Data Processing Job (~15 mins)
Now, we'll create a data processing job to filter out just the matches we want to see - like using Oracle Lens to find only the enemies you want to target!
Setting up Data Source
- Navigate to Build > Visual ETL flow in the top menu
2. Select the project you just created (rift-rewind-project-[your-name])
3. Click Create flow
4. Select Data managed using full-table access (compatibility) and click Continue - Under Add nodes, ensure Data sources is selected. Choose Amazon S3 as the data source. You should now see the visual editor.
- Getting data from Amazon S3: Next up, we need to get our data from the Amazon S3 bucket we created yesterday. In a new tab (keep this one open because we’ll be coming back to it), navigate to Amazon S3 in the console and open your
rift-rewind-match-data-[your-name]bucket.- You should have a folder inside called
match-historywhich was automatically created by the Lambda function we ran. - Inside
match-history, you’ll find a folder with user data for a specific user (for example,Hide on bush#KR1). Click on that folder to expand it. - Inside the user data folder, there should be two subfolders:
fullandstats.fullcontains the full match data with all the different variables, whilestatsjust includes player stats. Click on stats. - Select one of the the JSON objects in the bucket and click Copy S3 URI. This is the resource identifier, or the unique ID of that individual match! We’ll give this URI to SageMaker Unified Studio so it has access to the player data.
- You should have a folder inside called
Note: If Copy S3 URI is greyed out, make sure you’ve navigated all the way to the bottom of the folders. Ensure the path on top looks something like: Amazon S3 > Buckets > rift-rewind-match-data-[YOUR NAME] > match-history/ > [PLAYER ID] > stats/ > file name ending in .json
- Go back to SageMaker Unified Studio.
- Click on the Amazon S3 node in the visual editor and paste the S3 URI. It should auto-detect that the object is in JSON format.
- Alternatively, you can choose Browse S3 and select your file.
Note: If you are getting the error “NoSuchKey”, it can be due to spaces in the file path name. Try rerunning the Lambda function for a user without spaces in their name (such as riotID: thoxx#na1, with region: na1.)
- You should now be able to see a preview of the data with info from that match under Data preview!
Configuring Transformations
- On the left-hand menu, click the large green +. Select the Transforms tab. Here, you’ll see a bunch of options you can use to edit the data.
- Choose Drop columns.
- Click the plus symbol on the right of the S3 node and drag + drop it to connect it to the Drop columns node.
- Configuring Drop columns: Click the Drop columns node to open the editor like we did with S3 earlier.
- Here, you’ll see a list of the different fields in this data. Fields are pieces of information about that match — like the number of assists, the name of the champion, or number of deaths. It’ll also indicate the format of those fields.
- Let’s remove the championId and the gameCreation fields. Select the tick box next to those fields.
Note: Feel free to play around with the different transformation options, there are plenty more in addition to just the one we covered today!
Setting Data Target
- One last time, click the big green + in the menu, and this time select the Data targets tab.
- Choose Amazon S3 as your target.
- Drag and drop Drop columns to connect to the S3 target node.
- For the S3 URI, choose the S3 path under Project overview when you open Projects. Make sure to open this in a new tab so you don’t lose the current project you’re working on!
2. This is the S3 bucket associated with your SageMaker Unified Studio project. - For format, select JSON and leave the rest of the selections blank.
At this point, you should have an end-to-end visual flow. Now you can publish it.
- Click Save.
- Name your Job something descriptive, like
rift-rewing-data-processing. - Scroll to the bottom and click Save again to create the job.
Congratulations! You've crafted your first data processing ability - like mastering Ezreal's Q!
🎉 Step 3: Running your Data Processing Job (~5 mins)
Now that our job has been successfully created, go ahead and click Run (next to Save). You can click View Run to see the status of your job.Navigate to SageMaker Unified Studio Home > Projects > [Your Project Name] > Jobs > [Your Job Name]
- It may take a few minutes for your job to run. Check back after a bit and you should see the completed status of your job!
- On the Project overview tab, you should now be able to see the new file that was created!
- To see the output of your data processing job, navigate back to Amazon S3 in the console. You should see a folder that starts with
amazon-sagemaker-.... This folder stores the output of your job!- Open the folder and open all its subfolders until you reach
shared. This folder contains the JSON object output by our job, which will be in the formatpart-00000-... - If you open this JSON file, you’ll see the columns we chose to drop are no longer present!
- Open the folder and open all its subfolders until you reach
Conclusion
Congratulations, Summoner! Just like securing that crucial Baron buff, you've successfully executed your first visual ETL job in SageMaker Unified Studio. Dropping that column from your dataset might seem like a small CS lead, but it's actually your first step toward becoming the Challenger-tier data scientist you're meant to be!Your quick fingers that once kited enemies in bot lane are now orchestrating data transformations with equal precision. This simple column drop is just your first item purchase - your build will only get stronger from here. The possibilities for transforming your data are as vast as Runeterra itself!Remember:
- Each transformation is like adding another ability to your champion's kit
- Experiment with joins, aggregations, and feature engineering like you'd test different rune combinations
- Don't be afraid to dive the backline of your complex datasets - that's where the high-value insights are hiding
Just as you'd adapt your build path based on the enemy team comp, you can now adapt your data to meet any analytical challenge. Whether you're looking to carry with classification models or tank complex regression problems, your SageMaker skills will continue to scale into the late game.GG on your first ETL victory! The Nexus of data mastery awaits you!
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article