Rebuilding an LLM movie ranker with one Jev call and no training
Netflix's GenRec ranks what you watch next by reading your history as text. This rebuilds that shape from a single Jev API call with no training, and measures what the training was worth on MovieLens.
Rebuilding an LLM movie ranker with
one Jev call and no training
Part 1 of a series. This part compares the two designs and measures the Jev version against simple baselines. Part 2 trains rankers on the same data and puts them on the same candidate sets.
Most of a recommender's work is ranking. Retrieval hands over a few dozen candidates, and something has to decide which goes first. For years that something has been a model trained on engagement logs and fed thousands of engineered features.
In July 2026 Netflix described GenRec, a ranker that does the job differently: it writes a member's history out as text, reads it once, and scores every candidate in that same pass (Netflix Tech Blog ; arXiv 2608.10257 ). One correction up front, because it circulates: Netflix did not swap its ranker for GPT. GenRec runs on Netflix's own foundation model, post-trained on Netflix data, and Netflix calls it an early step rather than a finished migration.
This post takes GenRec's shape and rebuilds it from one call to Jev, TypeSafe AI's decision model, with no training at all. The result is
next-watch, a sample in the jev-samples repo. The point of the exercise is to see how much of GenRec's design survives when you remove the part Netflix spent the most on, and to put numbers on what that removal costs.What GenRec is
GenRec sits on an in-house Netflix foundation LLM. Training runs in two phases. Phase 1 adapts an open-source LLM to Netflix's catalog and member behaviour. Phase 2 post-trains that model for ranking, with labels from high-value engagement and reward weights that steer it toward exploration and long-term value.
The input side is where it breaks from a classic ranker. Instead of hand-built features, GenRec turns a member's history, the metadata of the titles involved and the request context into a natural-language prompt. Netflix calls this context engineering: deciding which events earn their tokens, dropping low-signal ones such as very short plays, and folding older history into a short summary. The paper reports that careful verbalization cut the prompt from about 5,000 tokens to about 1,700 with negligible loss in offline ranking quality.
The output side is just as plain. GenRec serves on vLLM in prefill-only mode: the model reads the prompt once, generates nothing, and a catalog-aware ranking head produces a score for every candidate in a single forward pass. Those scores are the ranking.
Offline, with about 40 times fewer Phase-2 labelled examples than the production ranker, GenRec scored roughly 1.6% higher relative MRR. Online, in an A/B test on about 10% of traffic for four weeks on batch-compute surfaces, it beat the production ranker by a statistically significant margin. Netflix presents all of this as an initial step toward an LLM-centred stack, not a platform-wide rollout.
The same shape, built from a Choice
Jev is a model that answers typed questions about a piece of state and returns no prose around the answers. One of its question types,
Choice, takes a list of options and returns a probability for each:1
2
3
4
5
6
"next_watch": {
"type": "choice",
"choice": "m...",
"confidence": 0.38,
"probabilities": { "m...": 0.38, "m...": 0.29, "m...": 0.24, "...": 0.0 }
}That is already GenRec's output shape. Give it a member's history as the state and twenty titles as the options, and sorting the probabilities gives a ranked list. The model reads the state once and scores every option in the same pass, so twenty titles cost one call, not twenty. Here is how the pieces line up:
| GenRec | next-watch |
|---|---|
| Verbalized history and context | A state with three named fields, each on its own character budget |
| Keep high-signal events, drop low-signal ones, fold old history into a summary | Likes (4 stars or more) and dislikes (2.5 or less) newest first, 3 and 3.5 ratings dropped, the rest counted into a taste summary |
| Candidate set from upstream retrieval | A free shortlist: genre match times log popularity |
| Catalog-aware scoring head over the candidates | next_watch, one Choice over twenty titles |
| Reward models for exploration and long-term value | discovery, a second Choice over the same titles, blended in by a weight |
| Offline MRR against the production ranker | MRR and Hit@k against three free rankers on held-out likes |
The rows line up. One difference sits under all of them, and the rest of this post is about it: GenRec trains on Netflix's engagement logs and learns an embedding for every title in the catalog. Jev has trained on nothing here. It sees each title's name, year and genres for the first time in the call.
One member, one call
Rules handle everything that doesn't need a model. Jev handles the one judgement that does.
- Retrieval, free. A shortlist of twenty unrated titles. Each is scored by genre match (how well its genres fit the member's liked genres) multiplied by the log of its popularity. In recommend mode, five of the twenty slots go to the most popular titles outside the member's usual genres, so there is something new on the list.
- Cold start, free. A member with fewer than five likes has too little history to describe. They get a popularity list and no call.
- State, computed. Python writes the history into three text fields, each on its own character budget.
- One call, two questions. Both are
Choices over the same twenty titles. - Blend, computed. Python weights the two probability lists and sorts.
The state
Each member becomes three fields. Here is member 414, abridged:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
taste:
Rated 2694 titles between 2000 and 2018, 1226 of them 4 stars or higher.
Share of those likes carrying each genre: Drama 57%, Comedy 36%, Action 21%,
Thriller 20%, Romance 18%.
recent_likes:
Kings of Summer, The (2013); Comedy; rated 4
The Barkley Marathons: The Race That Eats Its Young (2015); Documentary; rated 4
Straight Outta Compton (2015); Drama; rated 4
Darkest Hour (2017); Drama, War; rated 4
The Lobster (2015); Comedy, Romance, Sci-Fi; rated 4.5
...
recent_dislikes:
Fred Armisen: Standup for Drummers (2018); Comedy; rated 2.5
The Magnificent Seven (2016); Action, Western; rated 2.5
Green Lantern (2011); Action, Adventure, Sci-Fi; rated 2
Suicide Squad (2016); Action, Crime, Sci-Fi; rated 2
Catwoman (2004); Action, Crime, Fantasy; rated 0.5
...taste is computed in Python, because counting is not a job for the model. recent_likes holds titles rated 4 stars or more, newest first; recent_dislikes holds titles rated 2.5 or less. Ratings of 3 and 3.5 are left out because they carry little signal for their token cost, the same call GenRec makes when it drops short plays.The budgets are 500, 2,400 and 900 characters, so a long run of likes can't crowd out the dislikes. Each field is filled line by line and stops at the last line that fits, so no title arrives cut in half. The state comes to about 1,000 tokens; with the twenty titles listed under both questions, a whole call averaged about 2,400 input tokens. GenRec's compacted prompt is about 1,700.
The questions
| Question id | Weight | Asks |
|---|---|---|
next_watch | 1.0 | Which of these titles would this member most enjoy watching next? |
discovery | 0.25 | Which of these titles would widen the range of what this member watches? |
next_watch is the ranker. discovery stands in for GenRec's exploration rewards, with one difference that matters later: GenRec's rewards shape the model during training, while discovery can only pull at inference time, and its weight sets how hard. Each option key is a MovieLens id such as m318, and each description is the title, year and genres, in the same format as the state. The call is a few lines with the Python SDK (abridged):1
2
3
4
5
6
criteria = {f"m{c['id']}": option_text(movies[c["id"]]) for c in shortlist}
questions = {
qid: Choice(instructions=spec["instructions"], criteria=criteria)
for qid, spec in specs.items() # next_watch and discovery, from questions.yml
}
response = client.system_one(model="jev-1.13.0", state=state, questions=questions)The final score for each title is the weighted mean of its two probabilities. Changing a question or a weight is an edit to a YAML file, not to the code.
Here is the full output for member 414 in recommend mode: one call, twenty shortlisted titles, the top ten printed with the three Jev columns next to the free signals:

user414 results
Where the two designs part ways
| GenRec | next-watch on Jev | |
|---|---|---|
| What you need first | Years of engagement logs, a foundation-model team, GPUs for training and serving | A YAML file and an API key |
| Training | Two phases; Phase 2 re-run often to stay fresh | None |
| Candidates per ranking | Large sets, through a ranking head built for Netflix's catalog | Up to 255 options per Choice; the sample uses 20 |
| New titles | Learned at the next training run | Ranked from their name and description immediately, but only as well as that text allows |
| Where member data goes | Stays inside Netflix | Member history is sent to TypeSafe's API |
| Cost per ranking | Not published; Netflix calls serving cost a first-class constraint, roughly model size times context length | $0.000102 per member at September 2026 pricing |
| Expected accuracy | Higher, when you have the logs | Beats simple rankers with no training; should lose to a trained model on the same data |
GenRec's paper gives a rough sense of what the training is worth. Starting from the Netflix-adapted Phase-1 model instead of an off-the-shelf LLM improves offline ranking by 10 to 20%. Phase 2 adds another 35 to 50% on top, rising to about 80% two weeks later as the Phase-1 model goes stale. The paper never reports the case with no training at all, which is exactly the case next-watch measures.
A worked example
To test the ranker, the sample hides the last title a member liked and asks every ranker to place it among nineteen titles the member never rated. For member 414 the hidden title was Black Panther (2017). (Member 414 is not one of the 100 members in the evaluation below; this example comes from a separate run.) Where it landed out of twenty:

UPLOAD IMAGE: black-panther-rank.png (alt text: Rank of Black Panther out of 20 under each ranker for member 414: jev next_watch 2nd, genre match 13th, shortlist order 15th, popularity tied 17th to 19th.)
Jev gave it 0.29, behind only The Lives of Others at 0.38.
Black Panther had only three ratings in the data at that point, and its genres (Action, Adventure, Sci-Fi) overlap only part of this member's tastes, so every counting-based ranker put it in the bottom half. Jev reads the title and the year next to a history full of recent blockbusters and art-house films, and puts it second. The model has read about films, and that knowledge is doing the work here. GenRec gets the same kind of knowledge from its Phase-1 training on Netflix's catalog; Jev brings it from pre-training on text about the world.
How well Jev ranks
The sample runs the same test on 100 MovieLens members (
ml-latest-small: 100,836 ratings from 610 members over 9,742 titles). Three rules keep the test fair:- No reading the future. History stops before the hidden title, and ratings in the same second are dropped too, since members rate in batches. Popularity counts only ratings made before that moment.
- Hard negatives. The nineteen other titles are drawn in proportion to their popularity, the protocol SASRec and BERT4Rec report against. That makes popularity useless on its own: a ranker has to know something about this member to beat random.
- Ranking scored alone. The hidden title is always in the twenty, so every ranker orders the same set. Retrieval gets its own number: the free shortlist would have found the hidden title for 13 of 100 members. GenRec's paper reports its ranker the same way, over a given candidate set.
Ties count as the average over every order they allow, so no ranker gets credit for where the shuffle placed the answer.
Results from 100 members, seed 7,
jev-1.13.0, 30 September 2026:| Ranker | MRR | Hit@1 | Hit@5 | Widens | Distinct #1 |
|---|---|---|---|---|---|
| random (expected) | 0.180 | 0.050 | 0.250 | n/a | n/a |
| popularity | 0.178 | 0.050 | 0.265 | 0.180 | 54 |
| genre match | 0.294 | 0.132 | 0.412 | 0.000 | 89 |
| shortlist order | 0.253 | 0.100 | 0.430 | 0.000 | 74 |
| jev next_watch | 0.431 | 0.255 | 0.648 | 0.020 | 86 |
| jev blend | 0.411 | 0.255 | 0.565 | 0.020 | 86 |
MRR is the mean of one over the hidden title's rank. Hit@k is the share of members whose hidden title made the top k. Widens is the share of members whose top pick carries none of their three most-liked genres, and Distinct #1 counts different top picks across the run.
The run made 100 calls: 242,771 input tokens, 102 ms per call on average, $0.0102 in total at $0.042 per million input tokens. That is $0.000102 per member.

mrr
Reading the results
Genre match is the number to beat. It reads the same genres Jev sees and scores 0.294 MRR. Jev's
next_watch scores 0.431, and puts the hidden title first for 25.5 of 100 members against 13.2 for genre match and 5 for random. (The halves come from averaging over ties.) A paired bootstrap over the 100 members (10,000 resamples) puts Jev ahead by 0.137 MRR, with a 95% interval of 0.070 to 0.203. Jev ranked the hidden title higher for 64 members and lower for 26.The gain comes from reading, not counting. Beyond the genres, Jev sees titles, years and what the member disliked. The gap over genre match is what that text is worth, with nothing trained on this data.
Popularity at random is the test working. 0.178 against 0.180 means the hard negatives did their job.
Inference-time exploration didn't earn its place. At a weight of 0.25,
discovery changed no member's top pick, left Widens and Distinct #1 where they were, and lowered MRR by 0.020 (95% interval 0.011 to 0.030). It reorders the middle of the list without touching the top. For member 414 in recommend mode it put 0.58 on Howl's Moving Castle, a title outside their usual genres, and lifted it from fifth to third; across the evaluation, effects like that never reached the top. This is the clearest place the two designs differ in kind. GenRec buys exploration at training time, where a reward can reshape what the model believes. A weight at inference time has to outvote a confident next_watch answer, and 0.25 rarely does. A larger weight might buy range at some cost in accuracy; finding out means sweeping it on a different seed so it isn't tuned on the test set.Why the two headline numbers can't sit side by side
GenRec's headline is +1.6% relative MRR offline against Netflix's production ranker. Jev's is 0.431 against 0.294, about +47% relative against genre matching. Put next to each other, those look like a rout. They aren't comparable.
GenRec's baseline is a mature production model built on thousands of features and years of logs, evaluated on Netflix's own data and surfaces. Jev's baseline is a cosine over genre tags, evaluated on a 20-title shortlist from a public dataset. Beating the second by 47% says nothing about how close you are to the first. What the two numbers do share is a direction: both show that reading a verbalized history beats the ranker that came before it, in their respective settings.
What this does and doesn't show
It shows that a generic decision model, given a verbalized history and a list of titles, covers a fair part of GenRec's shape with no training: one read of the context, a score for every candidate in the same pass, and a ranking that beats the strongest free baseline for about a hundredth of a cent per member.
It does not show that this matches GenRec. GenRec learns a representation for every title from Netflix's own logs and was tested against a production system on real members. next-watch reads titles cold, sees only names, years and genres because MovieLens has no synopses, and competes against three free rankers offline. By GenRec's own figures, training should be expected to win by a wide margin on the same data. That is the next question, and it is the subject of Part 2.
Two caveats on the numbers. They come from one public dataset, one seed and 100 members; a repeat the same afternoon on 101 members (the same 100 plus member 414) scored 0.440 MRR for Jev against 0.292 for genre match. And the 102 ms is the round trip from one laptop; TypeSafe's own latency figures have not been reproduced outside the company.
Next in the series
Part 2 trains a System One model specifically for this ranking task and puts it on the same twenty-title candidate sets, with the same scoring.
- It reports the same model before and after that training, next to Jev, to measure what task-specific training adds over a decision model that has never seen these ratings.
The before-and-after result is the one GenRec's paper doesn't report: the same model with no training at all, which is how Jev ranks here.
Try it
The
next-watch sample lives in the jev-samples repo, with the full commands in its README.1
2
3
4
cd samples/next-watch
uv run next_watch.py --baselines-only # free rankers, no API key
uv run next_watch.py --user 414 # recommend for one member, one call
uv run next_watch.py # the 100-member evaluation, about a centThe evaluation prints the same table as above, plus what the run cost:

Results
The Netflix material above is paraphrased from the GenRec post on the Netflix Tech Blog and the GenRec paper . MovieLens data is from GroupLens Research at the University of Minnesota, which neither endorses nor reviewed this sample.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article