AWS Builder Center

I built a studio that paints a new morning every day — on AWS Lambda and Amazon Bedrock

Builder Log, Part 2. Rasik Rahman S wrote Part 1 on his Daily GitHub Explorer. Same weekend challenge, same standard, different build. He shipped a briefing agent. I shipped a studio that will not let me hold the brush.

TL;DR

Aubade is an autonomous studio. Once a day it reads one city's live weather, art-directs a piece with
Amazon Bedrock, renders four candidate variants of a generative composition, measures all four, and
publishes the one that scores highest. It writes a poem about that specific sky, narrates it with
Amazon Polly, then critiques its own output and parks a style mutation for tomorrow to compete
against.
The property that matters: the model proposes and the code decides. Nothing a language model
returns is applied to the style unless it wins on measurement. That is why the style improves instead
of drifting. Craft score went from 5.78 to 8.14 over the eight mornings of 14–21 August 2026.
It runs on Amazon EventBridge Scheduler and nothing else triggers it. Here's how I built it — and the
loop that stopped it getting worse every day.

The problem

A creative tool has to satisfy four demands that pull against each other.
  • Autonomy — it has to produce without being opened, or it dies of the blank cursor.
  • Quality — unattended output that nobody checks has to be good on the day it is made.
  • Novelty — the same composition every morning is a screensaver, not a studio.
  • Safety — a generative system left alone for months must not be able to produce garbage.
Most autonomous creative agents answer the first demand and quietly drop the fourth. They wire a
model to a scheduler, apply whatever comes back, and trust the model's taste. I wanted the opposite
balance: a system that can run for a year without supervision because the space of things it is
allowed to produce contains no catastrophes.

Why a deterministic engine with the model as art director

My first version asked Amazon Nova Pro to write the artwork as Scalable Vector Graphics (SVG)
directly. It worked on the first attempt, which turned out to be the problem. I got valid SVG that
was a tidy grid of overlapping circles and rectangles. Competent, and identical in character every
time.
So I inverted the roles. The model never draws a pixel. It chooses colour, writes the poem, and
adjusts numbers. A 299-line engine does the drawing: a flow field built from summed sines,
integrated into curling strokes, layered over a graded sky with one light source.
One inversion fixed three things at once. Composition became hard-coded judgment rather than a model's
guess. Output became reproducible, because the date seeds the field. And the SVG is always valid,
because it is assembled by code rather than predicted token by token.
The same field is then drawn a second time straight into a byte buffer and encoded as a Portable
Network Graphics (PNG) file with zlib and struct, 30 lines, no dependencies.
That is the whole reason the deployment package is 23,187 bytes and there is no container image in
this project.

Architecture

Figure 1: One scheduled trigger, one function, four artefacts. Solid lines are calls made during the
run; dashed lines are the outbound weather read and telemetry.
Left to right. Amazon EventBridge Scheduler holds the only trigger, a six-field cron expression in
Universal Coordinated Time (UTC) with two retries and a one-hour maximum event age.
There is no API, no queue and no button.
AWS Lambda holds the entire agent: python3.12 on arm64, 1024 MB, 300-second ceiling.
Amazon Bedrock runs both model roles — Nova Pro as art director, Nova Lite as critic — through the
Converse API. Amazon Polly turns the poem into an MP3 file with a neural voice. Amazon DynamoDB is
the memory: one table, one item for the reigning style genome and one item per published morning, so
seven days of history is a single query.
Every artefact lands in Amazon Simple Storage Service (Amazon S3). The bucket is private, encrypted,
and owner-enforced; Amazon CloudFront fronts it with an origin access control (OAC), which is the
supported way to keep an S3 origin unreachable except through the distribution.
Amazon CloudWatch keeps every candidate score. AWS Identity and Access Management (IAM) scopes the
execution role to one table and one bucket prefix. The whole thing is one CloudFormation template
with ten resources.

The core loop

  1. Scheduler invokes the function. If a piece already exists for today, the run stands down.
  2. Pick the next of 32 cities by day of year.
  3. Read that city's current weather from Open-Meteo — temperature, cloud, wind, precipitation, code.
  4. Ask Nova Pro for a palette drawn from that sky, a title, a poem, and a nudge to the style genome.
  5. Breed four candidate genomes and render all four.
  6. Score each candidate in code and select the winner.
  7. Draw the winner twice: vector poster and raster plate.
  8. Narrate the poem with Amazon Polly.
  9. Publish artefacts and the day's JSON record to Amazon S3.
  10. Ask Nova Lite to review the piece and park its proposed mutation as tomorrow's challenger.

Design decisions that mattered

Bounded genes, not validation. All 24 style genes carry a hard range — strokes 120–1,400, curl
0.3–4.5, stroke opacity 0.06–0.85. Any value from any model is clamped before it renders, and
anything unparseable falls back to that gene's default. A model can return
strokes: 999999 or a string where a float belongs, and the worst outcome is a piece that looks
slightly different from yesterday. That is a cheaper guarantee than any amount of error handling.
Weather drives the accents in code, not in the prompt. The rain streaks, snow and fog are added
by a function that reads the World Meteorological Organization weather code, with the streak angle
set from wind speed. If the model's palette drifts off-brief, the piece still
tracks the real sky.
One table, two item shapes. pk=GENOME holds the champion, pk=DAY holds history. Recent
memory is one query with ScanIndexForward=False, not a scan.
Idempotency as the retry strategy. A run that finds today's record already present exits in 38
milliseconds. That single guard is what makes Scheduler
retries safe to enable.
Everything external degrades. A failed weather read publishes with a recorded fallback sky. A
failed director call publishes a template poem. A failed voice publishes silently. Each failure is
written into the day's record under errors, because an agent nobody watches has to keep its own
logbook.

The hard part: a self-improving loop that got worse every day

Gap 1 — the critic-driven loop degraded, and its own scores proved it.
What I saw: I built the obvious design first. After publishing, Nova Lite reviews the piece and its
proposed mutation is applied to the style for tomorrow. Seven mornings of critic scores went 4.6,
4.0, 4.1, 3.9, 3.8, 3.9, 3.7. Monotonically worse.
What I tried first: better prompts. More specific instructions, explicit target bands, a stricter
output schema. It changed the wording of the verdicts and nothing else.
The fix: stop letting the critic decide. Every morning the agent now breeds four candidates — the
reigning champion untouched, the champion plus the director's weather nudge, the champion plus
yesterday's critic mutation, and the champion plus a deterministic random walk. All four render, all
four are scored by a six-component function in code, and the highest wins.
The critic's opinion became a hypothesis that has to survive measurement.
Why it behaves that way: the critic cannot see the image. It was grading numbers and a poem in the
dark, so its mutations were guesses, and applying a guess unconditionally is a random walk. An
unseeing grader also hedges downward — it always finds something absent, so it always marks low.
Gap 2 — the fitness function had no headroom.
What I saw: my first scorer gave the untouched default genome 9.47 out of 10. There was nowhere to
climb, and every candidate scored within 0.15 of every other.
What I tried first: adding weights. That moved the numbers around without widening the spread.
The fix: retarget every component against what the engine actually produces — ink coverage as a band
around 0.45 rather than a floor, mean stroke length against 750 pixels, light source within 0.10 of
a rule-of-thirds node. The default genome now scores about 7.3. I also split
out a craft score covering only the three components the genome controls, because contrast, chroma
and novelty are decided by the weather's palette and were drowning the signal.
Why it matters: a fitness function that everything passes is not selecting anything.
Gap 3 — no image model this account could reach.
What I saw: amazon.nova-canvas-v1:0 lists in the Region with lifecycle status LEGACY, and
inference returns an access error stating the model is legacy and has not been invoked recently. The
Stability structure-control models are AWS Marketplace models; this account returned an invalid
payment instrument, then a Marketplace authorization error after I granted the subscribe action to
the role.
What I tried first: a single control-structure call did succeed once before the wall came down, and
it returned a photographic scene that followed my composition's horizon and light placement exactly.
So the payload shape was right and the entitlement was not there.
The fix: make the agent draw its own raster instead of depending on an entitlement I do not control.
There is no SVG rasteriser in a bare Lambda runtime and I would not add a container image for one
dependency, so the flow field is drawn a second time into a byte buffer and encoded as PNG in 0.83
seconds at 896 by 896 pixels. The repaint step still exists, tries two models
in order, and logs a warning when neither answers.
Why it behaves that way: model availability on Amazon Bedrock is per-account and per-provider.
Marketplace-hosted models need a valid payment instrument and the Marketplace subscribe permission,
and a provider can mark a model legacy independently of whether it still appears in
list-foundation-models.

Proof it's real

On 21 August 2026 I deleted the day's record, attached a temporary two-minute schedule, and stepped
away. CloudWatch shows the run at 07:05:21 UTC scoring all four candidates — champion 7.591,
directed 7.489, critic 6.759, explore 7.498 — selecting champion, and publishing at 07:05:33. The
next tick at 07:07:12 found the record present and stood down in 38.09 milliseconds. Wall clock
12,260.94 ms, peak memory 125 MB of 1024 MB.
Being precise about what is proven: that morning was produced by the scheduler with no human in the
path. The other seven mornings in the archive were seeded by invoking the same handler with explicit
dates, so their weather is the weather at seeding time, not at 06:00 on each date. The daily cron has
been enabled since.

What's live

  • The daily schedule, enabled, cron(0 6 * * ? *) UTC.
  • Poem, narration, vector poster and raster plate, published every morning.
  • Four-candidate selection with the winning genome persisted and the generation counter advancing.
  • The public gallery, with each day's candidate scores and the critic's verdict readable in JSON.
  • Graceful degradation on weather, director and voice failures — the weather path has fired in
    practice and recorded a fallback sky.
The repaint step is in the code and does not run on this account. It is not live, and the gallery
labels it as such.

Cost

One invocation a day at 12 seconds and 1024 MB sits inside the always-free Lambda allowance, as do
CloudFront's monthly free terabyte and DynamoDB's 25 GB. Amazon S3 grows about 290 MB a year and
Polly neural is free for 1 million characters over 12 months against roughly 6,000 used a month.
Amazon Bedrock is the only real line: two Nova calls a morning. My estimate is 0.20–0.30 USD a month
once the 12-month items lapse. It has not yet run a full billing month, so treat that as an
estimate rather than a bill.

What I learned

  • A language model is a good art director and a poor judge of its own output. Ask it to propose;
    let code decide. The downward score trend was the most useful failure of the build.
  • A fitness function everything passes is not a fitness function. Calibrate against what your
    system actually produces before you trust it to select.
  • Constraints are what let an agent be left alone. I made this reliable by shrinking what it is
    allowed to say, not by handling more errors.
  • Model availability is the least portable part of an Amazon Bedrock project. Read
    modelLifecycle.status, and design so no single model is load-bearing.
  • The standard library often clears the bar. zlib plus struct produced the images I like most
    in the project and kept deployment at zip-and-upload.
  • Separate the signal you control from the signal you do not. Splitting craft out of total
    fitness is what made progress legible.

Try it

1
2
git clone https://github.com/logesh4v/Weekend-Challenge-Set-Your-Creative-App-Free.git
cd Weekend-Challenge-Set-Your-Creative-App-Free && ./deploy.sh
One CloudFormation stack, no SAM, no bootstrap. Then invoke it once and read the log: you will see
four candidate genomes scored and one selected, which is the agent making an aesthetic decision with
nobody in the room. Live gallery: https://d2bq7r7f8dyh8z.cloudfront.net
Part 3 is Rasik's. Follow @rasikbuilderaws  and
@logesh182  if you want the rest of the series.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article