
Build a Zero-Cost Serverless AI Microservice with AWS Lambda and Open Source
Deploy an event-driven text summarizer using AWS Lambda's perpetual free tier and Hugging Face's free Serverless Inference API—with zero credit card charges.
Generative AI projects often stall early because managed APIs like Amazon Bedrock and OpenAI incur pay-per-token bills. For developers experimenting or building hobby microservices, you don't need to commit to paid infrastructure.
This guide demonstrates how to deploy an event-driven text summarization microservice using AWS Lambda’s perpetual Free Tier (1,000,000 free requests per month) combined with Hugging Face’s free Serverless Inference API running open-source models like
facebook/bart-large-cnn.Architecture Overview
1
2
3
4
5
6
7
8
9
10
11
12
13
[HTTP Client / API Gateway]
│ (JSON payload)
▼
[AWS Lambda]
│ (HTTPS POST + Bearer Token)
▼
[Hugging Face Free Inference API]
│ (Open-Source BART Model)
▼
[AWS Lambda]
│ (Clean JSON Output)
▼
[Client]
- AWS Lambda (Python 3.12): Handles the request payload, manages timeout handling, and formats the output. Runs within the 1M monthly free-tier requests.
- Hugging Face Inference API: Runs inference on free hosted open-source models without requiring EC2, container clusters, or dedicated GPU instances.
- CloudWatch: Captures operational metrics and error logs.
Prerequisites
- An AWS Free-Tier Account.
- A free Hugging Face Account.
- A Hugging Face User Access Token (generate for free at huggingface.co > Settings > Access Tokens with "Read" permissions).
Step 1: Write the Lambda Function
This implementation relies strictly on Python's built-in standard library (
urllib.request and json), eliminating the need to bundle third-party wheel files or Lambda Layers.Create a function named
FreeSummarizerFunction in the AWS Lambda console using the Python 3.12 runtime:1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
import json
import os
import urllib.request
import urllib.error
# Free open-source summarization model hosted on Hugging Face
MODEL_ID = "facebook/bart-large-cnn"
API_URL = f"[https://api-inference.huggingface.co/models/](https://api-inference.huggingface.co/models/){MODEL_ID}"
HF_API_KEY = os.environ.get("HF_API_KEY")
def lambda_handler(event, context):
try:
# Parse payload
body = json.loads(event.get("body", "{}")) if isinstance(event.get("body"), str) else event.get("body", {})
input_text = body.get("text", "")
if not input_text:
return {
"statusCode": 400,
"body": json.dumps({"error": "Field 'text' is required."})
}
payload = {
"inputs": input_text,
"parameters": {
"max_length": 130,
"min_length": 30,
"do_sample": False
}
}
# Prepare HTTP request
headers = {
"Authorization": f"Bearer {HF_API_KEY}",
"Content-Type": "application/json"
}
req = urllib.request.Request(
API_URL,
data=json.dumps(payload).encode("utf-8"),
headers=headers
)
with urllib.request.urlopen(req, timeout=25) as response:
result = json.loads(response.read().decode("utf-8"))
summary_text = result[0].get("summary_text", "")
return {
"statusCode": 200,
"headers": {"Content-Type": "application/json"},
"body": json.dumps({
"model": MODEL_ID,
"summary": summary_text
})
}
except urllib.error.HTTPError as e:
error_info = e.read().decode("utf-8")
return {
"statusCode": e.code,
"body": json.dumps({"error": "Model Provider Error", "details": error_info})
}
except Exception as e:
return {
"statusCode": 500,
"body": json.dumps({"error": str(e)})
}
Step 2: Configure Environment & Timeouts
- Go to the Configuration tab in your Lambda function.
- Select General configuration → click Edit:
- Set Memory:
128 MB(lowest cost tier, sufficient since inference executes remotely). - Set Timeout:
30 seconds(allows model spin-up buffer on free endpoints).
- Select Environment variables → click Edit:
- Add Key:
HF_API_KEY - Value: Paste your free token starting with
hf_...
Step 3: Test the Service
Create a test event in the AWS Lambda console:
1
2
3
4
5
{
"body": {
"text": "AWS Lambda is a serverless, event-driven compute service that lets you run code for virtually any type of application or backend service without provisioning or managing servers. You can trigger Lambda from over 200 AWS services and software as a service (SaaS) applications, paying only for the compute time you consume."
}
}
Click Test. Within a few seconds, the function returns the condensed output:
1
2
3
4
5
{
"statusCode": 200,
"headers": { "Content-Type": "application/json" },
"body": "{\"model\": \"facebook/bart-large-cnn\", \"summary\": \"AWS Lambda is a serverless, event-driven compute service. You can trigger Lambda from over 200 AWS services and SaaS applications. You pay only for what you consume.\"}"
}
Cost Breakdown
- AWS Lambda: 1M invocations/month & 3.2M seconds of compute time are always free.
- Model Inference: Hugging Face Serverless API provides rate-limited free access for non-commercial and prototyping workloads.
- Network / CloudWatch: Within standard AWS Free Tier allotments.
Total operational cost: $0.00
What to Explore Next
- Try other open-source models: Replace the model string with
google/flan-t5-baseormistralai/Mistral-7B-Instruct-v0.3. - Add a Function URL: Enable AWS Lambda Function URLs with CORS enabled to consume this API straight from a web frontend without requiring API Gateway charges.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article