
From 2 Minutes to 17 Seconds: Optimizing an AWS Bedrock League of Legends AI Assistant
How I built a real-time gaming analytics platform and learned to make AI 85% faster
What You'll Learn
Building AI applications is exciting, but making them fast is where the real learning happens. In this article, I'll walk you through how I built a League of Legends analytics platform using AWS Bedrock Knowledge Bases, and more importantly, how I optimized it from painfully slow (114 seconds per query!) to production-ready (17 seconds).
By the end of this article, you'll understand:
- How to structure data pipelines for AWS Bedrock Knowledge Bases using real gaming data
- The critical difference between
retrieve()andretrieve_and_generate()(this one change cut my response time in half!) - When to choose Claude Haiku vs Sonnet for your use case
- Real-world optimization techniques that actually work
- How to build engaging UX that masks inevitable AI latency
- Common pitfalls with the Riot Games API and how to avoid them
What I Built
Ever wondered what your most-played League champion's optimal build is, or wanted quick insights without tabbing out of your browser? I created a platform that combines real-time summoner stats with an AI-powered chat assistant to help League players get instant, data-driven answers.
![[League Analytics Platform]](https://prod-assets.cosmic.aws.dev/a/34kqmGg82R6C7G1g7Ehh7QsEDxb/Scre.webp?imgSize=1000x500)
League-themed interface with summoner lookup and AI chat widget
Key features:
- Summoner Lookup: Enter any Riot ID and see profile stats plus top 3 champions with mastery details
- AI Chat Assistant: Ask about champions, items, or strategies - powered by AWS Bedrock and a Knowledge Base of 517 high-elo matches
- Smart Integration: Click any champion card and the chat automatically suggests relevant questions
- Progressive Loading: Rotating status messages keep users engaged during AI processing
But here's the thing - getting it to work was the easy part. Getting it to work fast was the real challenge.
The Performance Problem
When I first deployed my Lambda function, I was excited to test it. I typed: "What are the best items for Vayne?"
And I waited.
And waited.
114 seconds later, I got my answer.
Looking at the CloudWatch logs, I saw three tool calls, each taking 30+ seconds:
1
2
3
4
Tool #1: analyze_champion_performance → 30 seconds
Tool #2: query_match_data → 33 seconds
Tool #3: analyze_meta_trends → 28 seconds
Total: 97 seconds (plus overhead)
This wasn't going to work for a hackathon demo. Nobody waits 2 minutes for a chatbot response. Time to figure out what was wrong.
Understanding the Problem
The issue was in my architecture. Here's what I was doing:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
# My original (slow) code
def analyze_champion_performance(champion_name: str, role: str = None) -> str:
"""Analyze detailed performance metrics for a specific champion."""
query = f"champion performance analysis for {champion_name}..."
# This was the problem!
response = bedrock_client.retrieve_and_generate(
input={'text': query},
retrieveAndGenerateConfiguration={
'type': 'KNOWLEDGE_BASE',
'knowledgeBaseConfiguration': {
'knowledgeBaseId': KNOWLEDGE_BASE_ID,
'modelArn': KB_MODEL_ARN, # Another LLM call!
'retrievalConfiguration': {
'vectorSearchConfiguration': {
'numberOfResults': 8 # Way too many!
}
}
}
}
)
return response['output']['text']
What was happening:
- User asks: "What items for Vayne?"
- Main agent decides to call
analyze_champion_performancetool - Tool calls
retrieve_and_generatewhich:- Retrieves 8 document chunks from Knowledge Base
- Generates a complete response using the KB model (30 seconds!)
- Main agent receives that response
- Main agent generates another response based on the tool output (12 seconds!)
I was essentially running 3-4 separate LLM generations per question! No wonder it was slow.
The Fix: Architecture Over Everything
The breakthrough came when I realized I didn't need the Knowledge Base to generate responses - I just needed it to fetch relevant data. Here's the optimized version:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
# Optimized (fast) code
def query_match_data(query: str, max_results: int = 3) -> str:
"""Query the League of Legends match data knowledge base - returns raw data chunks."""
try:
# Use RETRIEVE only - much faster!
response = bedrock_client.retrieve(
knowledgeBaseId=KNOWLEDGE_BASE_ID,
retrievalQuery={'text': query},
retrievalConfiguration={
'vectorSearchConfiguration': {
'numberOfResults': max_results # Reduced from 8 to 3
}
}
)
# Extract just the text chunks
chunks = []
for result in response.get('retrievalResults', []):
content = result.get('content', {}).get('text', '')
if content:
chunks.append(content[:500]) # First 500 chars
# Return raw data for the main agent to synthesize
return f"Match data found ({len(chunks)} sources):\n\n" + "\n\n---\n\n".join(chunks)
except Exception as e:
return f"Error querying match data: {str(e)}"
Key changes:
retrieve()instead ofretrieve_and_generate()- Just get the data, don't generate- Reduced results from 8 to 3 - Faster retrieval, still good coverage
- Let the main agent do all synthesis - One generation instead of three
Result: 44 seconds (61% improvement!)
But I wasn't done yet.
Second Optimization: Model Selection
Looking at the timing breakdown for the 44-second queries:
- KB retrieval: ~27 seconds (retrieval of chunks)
- Agent synthesis: ~12 seconds (generating the final response)
The KB retrieval time was largely out of my control (more on that later), but that 12-second synthesis? That was using Claude Sonnet 3.5, which is thorough but not the fastest.
I switched to Claude Haiku 3.5:
1
2
# Environment variable change
MODEL_ID = "us.anthropic.claude-3-5-haiku-20241022-v1:0" # Was using Sonnet
Result: 17 seconds (another 61% improvement!)
The Trade-off
Haiku is less sophisticated than Sonnet, so I expected the quality to drop. But for factual queries about League of Legends items and strategies? Haiku was absolutely fine. The responses were still detailed, accurate, and helpful.
When to use Haiku:
- Factual queries with clear answers
- Data retrieval and basic analysis
- Cost-sensitive applications
- Speed is critical
When to use Sonnet:
- Complex reasoning required
- Creative or nuanced responses
- Quality matters more than speed
For my use case, Haiku was the right choice.
The Data Pipeline
While optimizing the AI response time, I also had to build the data pipeline that fed the Knowledge Base. This was its own learning experience.
![[League Analytics Platform]](https://prod-assets.cosmic.aws.dev/a/34kplEY7G7z60istwC5QKrlhCo5/Scre.webp?imgSize=1000x736)
Live stats from Riot Games API showing top 3 mastered champions
Collecting Match Data
I needed real League of Legends match data, so I built a pipeline:
Stage 1: Data Collection
1
2
3
4
5
6
# Pseudo-code for the collection process
for player in get_high_elo_players(): # Challenger, Grandmaster, Master
match_ids = get_player_matches(player.puuid)
for match_id in match_ids:
if not already_cached(match_id):
invoke_lambda('match_detail_fetcher', match_id)
Stage 2: Match Processing (Lambda)
- Fetch full match details from Riot API
- Transform raw data into structured format
- Calculate per-minute metrics (gold/min, CS/min, vision/min)
- Store in S3:
s3://bucket/matches/{patch}/{queue}/{matchId}.json - Cache match ID in DynamoDB to prevent duplicates
Stage 3: Aggregation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
def aggregate_match_data(bucket, patch):
matches = load_matches_from_s3(bucket, patch)
stats = {
'champion_stats': {}, # Win rates, KDA, gold efficiency
'role_meta': {}, # Position-specific trends
'item_builds': {}, # Popular build paths
'matchup_data': {}, # Head-to-head performance
'objective_impact': {} # First blood, dragon, baron correlations
}
for match in matches:
process_match(match, stats)
save_to_s3(stats, f"aggregated/{patch}/")
Result: 517 matches, 171 unique champions analyzed
Common Pitfalls with Riot API
Problem 1: Rate Limits Riot API has strict rate limits (20 requests/second, 100 requests/2 minutes). I hit these constantly during testing.
Solution: Add rate limiting and retry logic:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import time
from functools import wraps
def rate_limit(calls_per_second=1):
min_interval = 1.0 / calls_per_second
last_called = [0.0]
def decorator(func):
def wrapper(*args, **kwargs):
elapsed = time.time() - last_called[0]
wait_time = min_interval - elapsed
if wait_time > 0:
time.sleep(wait_time)
result = func(*args, **kwargs)
last_called[0] = time.time()
return result
return wrapper
return decorator
# 1 call every 2 seconds
def fetch_match_details(match_id):
# Your API call here
pass
Problem 2: Daily API Keys Development API keys expire daily and need manual renewal. This broke my automated collection several times.
Solution: Store keys in AWS Parameter Store with a monitoring script:
1
2
3
4
5
6
7
8
9
# Store key securely
aws ssm put-parameter \
--name "/riot-api/key" \
--value "RGAPI-your-key" \
--type "SecureString" \
--overwrite
# Lambda retrieves it at runtime
api_key = ssm.get_parameter(Name='/riot-api/key', WithDecryption=True)['Parameter']['Value']
Problem 3: Match ID Format Match IDs have region prefixes:
NA1_5388628816. I initially forgot to handle region mapping, which caused 404 errors.Solution: Map platform to routing region:
1
2
3
4
5
6
REGION_MAP = {
'na1': 'americas',
'euw1': 'europe',
'kr': 'asia',
# ... etc
}
Building the User Experience
Fast AI is important, but even 17 seconds feels long to a user. I needed to make the wait feel intentional.
![[League Analytics Platform]](https://prod-assets.cosmic.aws.dev/a/34kqMvp8q6Xcvp9KlPYNsTvZodN/Scre.webp?imgSize=710x1000)
Rotating status messages keep users informed while the AI searches
Progressive Thinking Messages
I added status messages that update every 15 seconds:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
const thinkingMessages = [
'🤔 Thinking...',
'📚 Searching knowledge base...',
'🔍 Analyzing champion data...',
'⚡ Processing your question...',
'🎮 Almost there...'
];
let thinkingInterval = setInterval(() => {
const indicator = document.getElementById('typingIndicator');
if (indicator) {
thinkingMessageIndex = (thinkingMessageIndex + 1) % thinkingMessages.length;
indicator.textContent = thinkingMessages[thinkingMessageIndex];
}
}, 15000); // Update every 15 seconds
This simple addition made the 17-second wait feel much more acceptable. Users knew the system was working, not frozen.
Smart Integration
Another UX win: clickable champion cards that pre-fill the chat.
![[League Analytics Platform]](https://prod-assets.cosmic.aws.dev/a/34kq8DPDoHG9OFh7EWTbUROeUyb/Scre.webp?imgSize=1000x579)
Click any champion card to automatically suggest a question
1
2
3
4
5
6
7
8
9
function askAboutChampion(championName) {
const chatInput = document.getElementById('chatInput');
chatInput.value = `What are the best items and strategies for ${championName}?`;
// Expand chat and focus
if (isMinimized) toggleChat();
chatInput.focus();
chatInput.select();
}
When someone looks up "Doublelift#NA1" and sees Vayne as his #1 champion, they can just click the Vayne card and instantly have a relevant question ready. It creates a seamless flow between the two features.
What I Would Do Differently
1. Reduce Knowledge Base Chunk Size
My KB is currently using 1000-token chunks. That 27-second retrieval time could probably be cut to 10-15 seconds by using 400-500 token chunks instead. Smaller chunks = faster vector search.
To change this, you need to recreate the data source (chunking strategy is immutable after creation):
- Delete existing data source in Bedrock console
- Add new data source with smaller chunks
- Re-sync (takes 5-30 minutes)
I didn't do this yet because I was racing against the hackathon deadline, but it's on my list.
2. Implement Streaming Responses
Right now, users wait 17 seconds and get everything at once. If I implemented streaming:
1
2
3
4
5
6
7
# Pseudo-code for streaming
def stream_response(agent_response):
for chunk in agent_response.stream():
send_to_websocket({
'type': 'chunk',
'content': chunk
})
Users would see text appearing word-by-word, which feels much faster even if the total time is the same.
3. Add Response Caching
Common questions like "best items for [popular champion]" get asked repeatedly. Caching these in DynamoDB or ElastiCache for a few hours would provide instant responses for cache hits:
1
2
3
4
5
6
7
8
9
cache_key = f"query:{hash(user_question)}"
cached = dynamodb.get_item(Key={'query_hash': cache_key})
if cached and not_expired(cached):
return cached['response']
else:
response = agent.invoke(user_question)
cache_response(cache_key, response, ttl=3600)
return response
Architecture Diagram
Here's the complete system architecture:
1
2
3
4
5
6
7
8
9
10
11
12
User Browser
↓
AWS Amplify (Static Website - HTML/CSS/JS)
↓
├─→ Lambda Function URL (riot-api-function)
│ └─→ Riot Games API (Live Summoner Data)
│
└─→ API Gateway (WebSocket)
└─→ Lambda (chat-agent-handler)
└─→ AWS Bedrock
├─→ Knowledge Base (League Match Data - 517 matches)
└─→ Claude 3.5 Haiku (Response Generation)
Data Pipeline:
1
2
3
4
5
6
7
Riot API → Lambda (match_detail_fetcher) → S3 (raw matches)
↓
Aggregation Script → S3 (processed stats)
↓
Bedrock Knowledge Base → Vector Embeddings
↓
Available for chat queries
The Numbers
Final Performance:
- Query response time: 17 seconds (down from 114s)
- Total improvement: 85% faster
- Cost per query: ~$0.02 (Bedrock + Lambda)
- KB size: 517 matches, 171 champions
Optimization Breakdown:
- Initial implementation: 114 seconds
- After switching to
retrieve(): 44 seconds (61% improvement) - After switching to Haiku: 17 seconds (additional 61% improvement)
AWS Bedrock provides detailed, data-driven insights in ~17 seconds
Key Takeaways
After building this project, here's what I learned:
1. Architecture Matters More Than Model Choice
Switching from Sonnet to Haiku saved 9 seconds. But fixing my architecture (nested LLM calls) saved 70 seconds. Always look at your architecture first.
2. Understand What Your Tools Are Actually Doing
I assumed
retrieve_and_generate was just fetching data. It wasn't - it was running a full LLM generation. Reading the docs carefully would have saved me hours of debugging.3. Start Fast, Optimize for Quality Later
I should have started with Haiku and smaller KB result sets, then scaled up if needed. Instead, I started with the "best" options (Sonnet, lots of results) and had to optimize down.
4. Real Data > Generic Content
The 517 high-elo matches made a huge difference in response quality. Users noticed and appreciated that recommendations were based on actual competitive play, not generic guides.
5. UX Can Mask Performance
Progressive loading messages transformed a "frustratingly slow" 17 seconds into an "acceptable" 17 seconds. Never underestimate the power of good feedback.
6. Measure Everything
CloudWatch Logs were invaluable. Every optimization started with: "Let me check the actual timings in the logs." Without that data, I would have been guessing.
Common Debugging Tips
Problem: Lambda timing out
1
2
3
4
5
# Check CloudWatch Logs
aws logs tail /aws/lambda/your-function-name --follow
# Look for the actual time each operation takes
# My logs showed me exactly which tool calls were slow
Problem: Bedrock Knowledge Base returning irrelevant results
- Check your chunk size (1000 tokens might be too large)
- Review your aggregated data format (is it structured well for semantic search?)
- Test queries directly in Bedrock console before blaming your Lambda code
Problem: WebSocket disconnections
1
2
3
4
5
// Add reconnection logic
socket.onclose = () => {
console.log('Disconnected, reconnecting in 3s...');
setTimeout(connectToWebSocket, 3000);
};
Problem: CORS errors on Lambda Function URL
1
2
3
4
# Ensure CORS is configured
aws lambda get-function-url-config --function-name your-function
# Should show AllowOrigins with your domain
What's Next
This project opened my eyes to how much optimization matters in AI applications. Some ideas I'm exploring:
- Multi-region support: Expand beyond NA to EUW, KR, and other regions
- Historical tracking: Store meta evolution across multiple patches
- Advanced analytics: Use SageMaker for predictive meta analysis
- Public API: Let other developers build on top of my data pipeline
Try It Yourself
The complete repository includes:
- Frontend (HTML/CSS/JS with League theming)
- Lambda functions (chat agent, Riot API integration, match processor)
- Data collection scripts (Riot API → S3 pipeline)
- Complete documentation for all components
To deploy your own version:
- Set up AWS Bedrock Knowledge Base
- Collect League match data using the provided scripts
- Deploy Lambda functions with the provided code
- Configure WebSocket API Gateway
- Deploy frontend to Amplify
Detailed instructions are in each folder's README.
Final Thoughts
I started this project thinking the hard part would be getting the AI to work. It turned out the hard part was getting the AI to work well.
Going from 114 seconds to 17 seconds taught me more about real-world AI optimization than any tutorial could. I learned to:
- Read CloudWatch logs obsessively
- Question my assumptions about what tools are doing
- Prioritize architecture over individual optimizations
- Balance quality with performance based on use case
- Build UX that works with AI latency, not against it
The skills I picked up - AWS Bedrock Knowledge Bases, Lambda optimization, WebSocket APIs, data pipeline design, and performance debugging - are directly applicable to production AI systems.
If you're building something similar, my biggest advice: deploy early, measure everything, and optimize based on real data. Your first version will be slow. That's okay. The learning happens when you make it fast.
About the Author
I'm Mike Little, working in financial services designing products and architecture. This project was my first deep dive into AWS Bedrock and Knowledge Bases, and I learned a ton in the process.
Connect with me on LinkedIn if you want to discuss AI optimization, AWS services, or League of Legends!
Note: This project uses the Riot Games API but isn't endorsed by Riot Games and doesn't reflect their views or opinions.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article