AWS Builder Center

Rebuilding an LLM movie ranker with one Jev call and no training

Netflix's GenRec ranks what you watch next by reading your history as text. This rebuilds that shape from a single Jev API call with no training, and measures what the training was worth on MovieLens.

Rebuilding an LLM movie ranker with

one Jev call and no training

Part 1 of a series. This part compares the two designs and measures the Jev version against simple baselines. Part 2 trains rankers on the same data and puts them on the same candidate sets.
Most of a recommender's work is ranking. Retrieval hands over a few dozen candidates, and something has to decide which goes first. For years that something has been a model trained on engagement logs and fed thousands of engineered features.
In July 2026 Netflix described GenRec, a ranker that does the job differently: it writes a member's history out as text, reads it once, and scores every candidate in that same pass (Netflix Tech Blog ; arXiv 2608.10257 ). One correction up front, because it circulates: Netflix did not swap its ranker for GPT. GenRec runs on Netflix's own foundation model, post-trained on Netflix data, and Netflix calls it an early step rather than a finished migration.
This post takes GenRec's shape and rebuilds it from one call to Jev, TypeSafe AI's decision model, with no training at all. The result is next-watch, a sample in the jev-samples  repo. The point of the exercise is to see how much of GenRec's design survives when you remove the part Netflix spent the most on, and to put numbers on what that removal costs.

What GenRec is

GenRec sits on an in-house Netflix foundation LLM. Training runs in two phases. Phase 1 adapts an open-source LLM to Netflix's catalog and member behaviour. Phase 2 post-trains that model for ranking, with labels from high-value engagement and reward weights that steer it toward exploration and long-term value.
The input side is where it breaks from a classic ranker. Instead of hand-built features, GenRec turns a member's history, the metadata of the titles involved and the request context into a natural-language prompt. Netflix calls this context engineering: deciding which events earn their tokens, dropping low-signal ones such as very short plays, and folding older history into a short summary. The paper reports that careful verbalization cut the prompt from about 5,000 tokens to about 1,700 with negligible loss in offline ranking quality.
The output side is just as plain. GenRec serves on vLLM in prefill-only mode: the model reads the prompt once, generates nothing, and a catalog-aware ranking head produces a score for every candidate in a single forward pass. Those scores are the ranking.
Offline, with about 40 times fewer Phase-2 labelled examples than the production ranker, GenRec scored roughly 1.6% higher relative MRR. Online, in an A/B test on about 10% of traffic for four weeks on batch-compute surfaces, it beat the production ranker by a statistically significant margin. Netflix presents all of this as an initial step toward an LLM-centred stack, not a platform-wide rollout.

The same shape, built from a Choice

Jev is a model that answers typed questions about a piece of state and returns no prose around the answers. One of its question types, Choice, takes a list of options and returns a probability for each:
1
2
3
4
5
6
"next_watch": {
"type": "choice",
"choice": "m...",
"confidence": 0.38,
"probabilities": { "m...": 0.38, "m...": 0.29, "m...": 0.24, "...": 0.0 }
}
That is already GenRec's output shape. Give it a member's history as the state and twenty titles as the options, and sorting the probabilities gives a ranked list. The model reads the state once and scores every option in the same pass, so twenty titles cost one call, not twenty. Here is how the pieces line up:
GenRecnext-watch
Verbalized history and contextA state with three named fields, each on its own character budget
Keep high-signal events, drop low-signal ones, fold old history into a summaryLikes (4 stars or more) and dislikes (2.5 or less) newest first, 3 and 3.5 ratings dropped, the rest counted into a taste summary
Candidate set from upstream retrievalA free shortlist: genre match times log popularity
Catalog-aware scoring head over the candidatesnext_watch, one Choice over twenty titles
Reward models for exploration and long-term valuediscovery, a second Choice over the same titles, blended in by a weight
Offline MRR against the production rankerMRR and Hit@k against three free rankers on held-out likes
The rows line up. One difference sits under all of them, and the rest of this post is about it: GenRec trains on Netflix's engagement logs and learns an embedding for every title in the catalog. Jev has trained on nothing here. It sees each title's name, year and genres for the first time in the call.

One member, one call

Rules handle everything that doesn't need a model. Jev handles the one judgement that does.
  1. Retrieval, free. A shortlist of twenty unrated titles. Each is scored by genre match (how well its genres fit the member's liked genres) multiplied by the log of its popularity. In recommend mode, five of the twenty slots go to the most popular titles outside the member's usual genres, so there is something new on the list.
  2. Cold start, free. A member with fewer than five likes has too little history to describe. They get a popularity list and no call.
  3. State, computed. Python writes the history into three text fields, each on its own character budget.
  4. One call, two questions. Both are Choices over the same twenty titles.
  5. Blend, computed. Python weights the two probability lists and sorts.

The state

Each member becomes three fields. Here is member 414, abridged:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
taste:
Rated 2694 titles between 2000 and 2018, 1226 of them 4 stars or higher.
Share of those likes carrying each genre: Drama 57%, Comedy 36%, Action 21%,
Thriller 20%, Romance 18%.

recent_likes:
Kings of Summer, The (2013); Comedy; rated 4
The Barkley Marathons: The Race That Eats Its Young (2015); Documentary; rated 4
Straight Outta Compton (2015); Drama; rated 4
Darkest Hour (2017); Drama, War; rated 4
The Lobster (2015); Comedy, Romance, Sci-Fi; rated 4.5
...

recent_dislikes:
Fred Armisen: Standup for Drummers (2018); Comedy; rated 2.5
The Magnificent Seven (2016); Action, Western; rated 2.5
Green Lantern (2011); Action, Adventure, Sci-Fi; rated 2
Suicide Squad (2016); Action, Crime, Sci-Fi; rated 2
Catwoman (2004); Action, Crime, Fantasy; rated 0.5
...
taste is computed in Python, because counting is not a job for the model. recent_likes holds titles rated 4 stars or more, newest first; recent_dislikes holds titles rated 2.5 or less. Ratings of 3 and 3.5 are left out because they carry little signal for their token cost, the same call GenRec makes when it drops short plays.
The budgets are 500, 2,400 and 900 characters, so a long run of likes can't crowd out the dislikes. Each field is filled line by line and stops at the last line that fits, so no title arrives cut in half. The state comes to about 1,000 tokens; with the twenty titles listed under both questions, a whole call averaged about 2,400 input tokens. GenRec's compacted prompt is about 1,700.

The questions

Question idWeightAsks
next_watch1.0Which of these titles would this member most enjoy watching next?
discovery0.25Which of these titles would widen the range of what this member watches?
next_watch is the ranker. discovery stands in for GenRec's exploration rewards, with one difference that matters later: GenRec's rewards shape the model during training, while discovery can only pull at inference time, and its weight sets how hard. Each option key is a MovieLens id such as m318, and each description is the title, year and genres, in the same format as the state. The call is a few lines with the Python SDK (abridged):
1
2
3
4
5
6
criteria = {f"m{c['id']}": option_text(movies[c["id"]]) for c in shortlist}
questions = {
qid: Choice(instructions=spec["instructions"], criteria=criteria)
for qid, spec in specs.items() # next_watch and discovery, from questions.yml
}
response = client.system_one(model="jev-1.13.0", state=state, questions=questions)
The final score for each title is the weighted mean of its two probabilities. Changing a question or a weight is an edit to a YAML file, not to the code.
Here is the full output for member 414 in recommend mode: one call, twenty shortlisted titles, the top ten printed with the three Jev columns next to the free signals:
alt text: Terminal output of next_watch.py --user 414. One Jev call took 179 ms and 2,818 input tokens and cost $0.000118. The top ten titles are listed with their jev blend, next watch and discovery probabilities, popularity and genre match. Intouchables is first at 0.240, 12 Angry Men second at 0.216, and Howl's Moving Castle third at 0.164, lifted by a discovery probability of 0.580.)
user414 results

Where the two designs part ways

GenRecnext-watch on Jev
What you need firstYears of engagement logs, a foundation-model team, GPUs for training and servingA YAML file and an API key
TrainingTwo phases; Phase 2 re-run often to stay freshNone
Candidates per rankingLarge sets, through a ranking head built for Netflix's catalogUp to 255 options per Choice; the sample uses 20
New titlesLearned at the next training runRanked from their name and description immediately, but only as well as that text allows
Where member data goesStays inside NetflixMember history is sent to TypeSafe's API
Cost per rankingNot published; Netflix calls serving cost a first-class constraint, roughly model size times context length$0.000102 per member at September 2026 pricing
Expected accuracyHigher, when you have the logsBeats simple rankers with no training; should lose to a trained model on the same data
GenRec's paper gives a rough sense of what the training is worth. Starting from the Netflix-adapted Phase-1 model instead of an off-the-shelf LLM improves offline ranking by 10 to 20%. Phase 2 adds another 35 to 50% on top, rising to about 80% two weeks later as the Phase-1 model goes stale. The paper never reports the case with no training at all, which is exactly the case next-watch measures.

A worked example

To test the ranker, the sample hides the last title a member liked and asks every ranker to place it among nineteen titles the member never rated. For member 414 the hidden title was Black Panther (2017). (Member 414 is not one of the 100 members in the evaluation below; this example comes from a separate run.) Where it landed out of twenty:
Rank of Black Panther out of 20 under each ranker for member 414: jev next_watch 2nd, genre match 13th, shortlist order 15th, popularity tied 17th to 19th.
UPLOAD IMAGE: black-panther-rank.png (alt text: Rank of Black Panther out of 20 under each ranker for member 414: jev next_watch 2nd, genre match 13th, shortlist order 15th, popularity tied 17th to 19th.)
Jev gave it 0.29, behind only The Lives of Others at 0.38.
Black Panther had only three ratings in the data at that point, and its genres (Action, Adventure, Sci-Fi) overlap only part of this member's tastes, so every counting-based ranker put it in the bottom half. Jev reads the title and the year next to a history full of recent blockbusters and art-house films, and puts it second. The model has read about films, and that knowledge is doing the work here. GenRec gets the same kind of knowledge from its Phase-1 training on Netflix's catalog; Jev brings it from pre-training on text about the world.

How well Jev ranks

The sample runs the same test on 100 MovieLens members (ml-latest-small: 100,836 ratings from 610 members over 9,742 titles). Three rules keep the test fair:
  • No reading the future. History stops before the hidden title, and ratings in the same second are dropped too, since members rate in batches. Popularity counts only ratings made before that moment.
  • Hard negatives. The nineteen other titles are drawn in proportion to their popularity, the protocol SASRec and BERT4Rec report against. That makes popularity useless on its own: a ranker has to know something about this member to beat random.
  • Ranking scored alone. The hidden title is always in the twenty, so every ranker orders the same set. Retrieval gets its own number: the free shortlist would have found the hidden title for 13 of 100 members. GenRec's paper reports its ranker the same way, over a given candidate set.
Ties count as the average over every order they allow, so no ranker gets credit for where the shuffle placed the answer.
Results from 100 members, seed 7, jev-1.13.0, 30 September 2026:
RankerMRRHit@1Hit@5WidensDistinct #1
random (expected)0.1800.0500.250n/an/a
popularity0.1780.0500.2650.18054
genre match0.2940.1320.4120.00089
shortlist order0.2530.1000.4300.00074
jev next_watch0.4310.2550.6480.02086
jev blend0.4110.2550.5650.02086
MRR is the mean of one over the hidden title's rank. Hit@k is the share of members whose hidden title made the top k. Widens is the share of members whose top pick carries none of their three most-liked genres, and Distinct #1 counts different top picks across the run.
The run made 100 calls: 242,771 input tokens, 102 ms per call on average, $0.0102 in total at $0.042 per million input tokens. That is $0.000102 per member.
Bar chart of MRR by ranker on 100 MovieLens members, seed 7: jev next_watch 0.431, jev blend 0.411, genre match 0.294, shortlist order 0.253, popularity 0.178, with a dashed line at the random expectation of 0.180.
mrr

Reading the results

Genre match is the number to beat. It reads the same genres Jev sees and scores 0.294 MRR. Jev's next_watch scores 0.431, and puts the hidden title first for 25.5 of 100 members against 13.2 for genre match and 5 for random. (The halves come from averaging over ties.) A paired bootstrap over the 100 members (10,000 resamples) puts Jev ahead by 0.137 MRR, with a 95% interval of 0.070 to 0.203. Jev ranked the hidden title higher for 64 members and lower for 26.
The gain comes from reading, not counting. Beyond the genres, Jev sees titles, years and what the member disliked. The gap over genre match is what that text is worth, with nothing trained on this data.
Popularity at random is the test working. 0.178 against 0.180 means the hard negatives did their job.
Inference-time exploration didn't earn its place. At a weight of 0.25, discovery changed no member's top pick, left Widens and Distinct #1 where they were, and lowered MRR by 0.020 (95% interval 0.011 to 0.030). It reorders the middle of the list without touching the top. For member 414 in recommend mode it put 0.58 on Howl's Moving Castle, a title outside their usual genres, and lifted it from fifth to third; across the evaluation, effects like that never reached the top. This is the clearest place the two designs differ in kind. GenRec buys exploration at training time, where a reward can reshape what the model believes. A weight at inference time has to outvote a confident next_watch answer, and 0.25 rarely does. A larger weight might buy range at some cost in accuracy; finding out means sweeping it on a different seed so it isn't tuned on the test set.

Why the two headline numbers can't sit side by side

GenRec's headline is +1.6% relative MRR offline against Netflix's production ranker. Jev's is 0.431 against 0.294, about +47% relative against genre matching. Put next to each other, those look like a rout. They aren't comparable.
GenRec's baseline is a mature production model built on thousands of features and years of logs, evaluated on Netflix's own data and surfaces. Jev's baseline is a cosine over genre tags, evaluated on a 20-title shortlist from a public dataset. Beating the second by 47% says nothing about how close you are to the first. What the two numbers do share is a direction: both show that reading a verbalized history beats the ranker that came before it, in their respective settings.

What this does and doesn't show

It shows that a generic decision model, given a verbalized history and a list of titles, covers a fair part of GenRec's shape with no training: one read of the context, a score for every candidate in the same pass, and a ranking that beats the strongest free baseline for about a hundredth of a cent per member.
It does not show that this matches GenRec. GenRec learns a representation for every title from Netflix's own logs and was tested against a production system on real members. next-watch reads titles cold, sees only names, years and genres because MovieLens has no synopses, and competes against three free rankers offline. By GenRec's own figures, training should be expected to win by a wide margin on the same data. That is the next question, and it is the subject of Part 2.
Two caveats on the numbers. They come from one public dataset, one seed and 100 members; a repeat the same afternoon on 101 members (the same 100 plus member 414) scored 0.440 MRR for Jev against 0.292 for genre match. And the 102 ms is the round trip from one laptop; TypeSafe's own latency figures have not been reproduced outside the company.

Next in the series

Part 2 trains a System One model specifically for this ranking task and puts it on the same twenty-title candidate sets, with the same scoring.
  • It reports the same model before and after that training, next to Jev, to measure what task-specific training adds over a decision model that has never seen these ratings.
The before-and-after result is the one GenRec's paper doesn't report: the same model with no training at all, which is how Jev ranks here.

Try it

The next-watch sample lives in the jev-samples  repo, with the full commands in its README.
1
2
3
4
cd samples/next-watch
uv run next_watch.py --baselines-only # free rankers, no API key
uv run next_watch.py --user 414 # recommend for one member, one call
uv run next_watch.py # the 100-member evaluation, about a cent
The evaluation prints the same table as above, plus what the run cost:
 Terminal output of uv run next_watch.py: the results table for 100 of 100 sampled members at seed 7, with jev next_watch at 0.431 MRR highlighted and genre match at 0.294, followed by 100 Jev calls, 242,771 input tokens, 102 ms per call on average and $0.01020 in total.
Results
The Netflix material above is paraphrased from the GenRec post on the Netflix Tech Blog  and the GenRec paper . MovieLens data is from GroupLens Research  at the University of Minnesota, which neither endorses nor reviewed this sample.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article