
HospitalityLLM: Building an Industry Model
The build story behind HospitalityLLM: how we created an industry model framework for data preparation, fine-tuning, evaluation, and internal hosting using AWS LLMOps and an 8B Small Language Model (SLM).
Series: industry-llm (1 article)
- 1HospitalityLLM: Building an Industry Model This article
By: Bharat Lakhiyani and Abhijeet Patil
When we started building HospitalityLLM, the first question sounded simple: can we create a language model that understands the work of an industry, not just the words used by that industry?
That question matters because industry knowledge is not only terminology. It is policy, workflow, escalation, compliance, service judgment, operating rhythm, and customer context. A general-purpose LLM can write a polished answer. An industry model needs to understand what a good answer means inside the business.
Hospitality was our first proving ground. Hotels are operationally dense: front desk, reservations, housekeeping, engineering, guest experience, loyalty, food and beverage, sales, events, revenue management, brand marketing, and compliance all intersect in the same guest journey. A model that helps in this environment has to do more than sound courteous. It has to know when to escalate, when to avoid overpromising, when to apply policy, when to ask for a human handoff, and how to keep the response aligned to the brand.
But the bigger goal was never limited to one hospitality use case. HospitalityLLM is the first implementation of a reusable industry model framework. The same approach can be adapted to financial services, retail, healthcare and life sciences, manufacturing, telecom, public sector, travel, energy, or any industry where valuable knowledge lives in operating procedures, customer interactions, transaction records, and expert judgment.
For broader background, the case for enterprise-owned industry LLMs is covered in Why Your Enterprise Needs Its Own LLM , and the hospitality industry motivation is covered in Why Travel and Tourism Needs More Than a Generic LLM .
This post is the build story: the challenges we hit, the fixes that worked, and how AWS LLMOps helped turn a prototype into a repeatable framework. The public name is HospitalityLLM. In this implementation, the model substrate is an 8B, fine-tuned for hospitality industry tasks.
What We Built
HospitalityLLM is an industry model implementation for hospitality, built on a framework that can be reused across industries and use cases. The public sample uses a fictional hotel brand, ZenithStays, so builders can run the pattern without exposing proprietary customer data.
The framework includes:
- An industry data contract for training and evaluation examples.
- A data preparation pipeline with privacy, quality, deduplication, and novelty gates.
- QLoRA supervised fine-tuning of Qwen3-8B on Amazon SageMaker.
- Model merge and deployment to a SageMaker real-time endpoint.
- A frozen evaluation set and benchmark loop using Amazon Bedrock as judge infrastructure.
- An internally hosted model experience for controlled business consumption.
This is important: the framework is not just a process checklist. A process tells you what steps to run. A framework gives teams reusable decision points, contracts, quality gates, evaluation controls, and deployment patterns that can be adapted to a new industry without starting from zero.

Figure 1: Industry model framework. HospitalityLLM is the first implementation, but the same pattern can support other industries and use cases.
Our LLM Command Center
- The Command Center Portal: chat compare, pipeline, observe, data, quality, lineage, HP sweep, train, register, deploy, evaluate
- Honest results: fine-tuned Qwen3 95.1, frontier Claude 94.1, base Qwen3 83.9 on a 100-point scale (an MVP run, 10 of 47+ scenarios)
- Why data is the critical piece, and the curation work behind a 5,000+ record hospitality dataset
Challenge 1: Sounding Right Was Not Enough
Our first challenge was separating fluent language from useful industry behavior.
A general model can write a polite hotel response. It can say, "We apologize for the inconvenience." That is not the same as knowing whether the issue is urgent, which team should handle it, what recovery action is allowed, which policy applies, or whether the model should stop and escalate.
The fix was to define the model around real industry tasks rather than general conversation. For hospitality, the initial task set included:
- Guest service Q&A.
- Complaint classification and response.
- Reservation intent extraction.
- Housekeeping and engineering work-order creation.
- Review summarization.
- Loyalty and property-policy Q&A.
- Sales, events, compliance, and revenue-management scenarios.
The same framing applies outside hospitality. In banking, tasks might include dispute triage, policy-guided servicing, branch operations, or advisor enablement. In retail, they might include returns handling, store associate guidance, inventory exceptions, or product-support responses. In manufacturing, they might include quality notes, work instructions, maintenance triage, or supplier issue handling.
The lesson: start with the work, not the model.
Challenge 2: Industry Data Was Messy
Useful industry data rarely arrives in training-ready form.
In hospitality, the valuable material lived across PMS service requests, CRM notes, guest transcripts, SOPs, maintenance logs, loyalty interactions, sales-event records, and internal guidance. Some records were structured. Many were not. Some were clean. Many contained private information, repeated templates, partial notes, internal identifiers, or language that should not become model behavior.
We solved this by making data preparation the center of the framework.

Figure 2: Data preparation framework with privacy, quality, deduplication, and novelty controls.
The seven gates are reusable across industries:
- Sensitive-data scrubbing: Remove or redact names, emails, phone numbers, account identifiers, loyalty IDs, reservation codes, folio numbers, or any industry-specific private identifiers.
- Format normalization: Convert exports, transcripts, documents, and manually curated examples into one consistent structure.
- Task wrapping: Turn source records into model-ready tasks that teach the model how to behave.
- Hard filtering: Remove records that are too short, too long, malformed, truncated, incomplete, or unsafe.
- Near-duplicate removal: Prevent repetitive records from dominating training.
- Quality scoring: Score examples for accuracy, completeness, register, specificity, and safety.
- Novelty probing: Compare candidate examples against the target base model and prioritize examples that add new learning signal.
The v2 schema is intentionally simple:
1
2
3
4
5
6
7
8
9
10
{
"query": "User question or operational scenario",
"response": "Expected assistant response",
"domain": "front_desk | revenue_management | guest_experience | housekeeping | food_beverage | sales_events | brand_marketing | compliance_safety",
"persona": "optional role, such as front_desk_agent or revenue_manager",
"property": "optional property identifier",
"complexity": "simple | moderate | complex",
"method": "human_curated | sop_extraction | self_instruct | evol_instruct",
"sop_references": ["optional SOP identifiers"]
}For another industry, the metadata changes, but the contract stays familiar: prompt, expected response, industry lane, persona, complexity, method, and references.
Here is a representative hospitality example:
1
2
3
4
5
6
7
8
9
10
{
"query": "A guest says the WiFi in their room keeps dropping every 20 minutes and they have client video calls today. Draft the guest response and assign a priority.",
"response": "Priority: High. I apologize for the disruption, especially with client calls scheduled today. I am escalating this to our technical support team immediately and asking the front desk to provide backup business-center access while we stabilize the in-room connection. We will follow up with you within 15 minutes and, if needed, arrange an alternate workspace or room move.",
"domain": "guest_experience",
"persona": "front_desk_agent",
"property": "ZS-DTN-01",
"complexity": "moderate",
"method": "human_curated",
"sop_references": ["SOP-GX-01"]
}That sample teaches priority, empathy, escalation, backup options, and a bounded promise. A weaker sample would only teach the model to be polite.
In the sample project, builders can run the pipeline against a supported source type:
1
2
3
4
5
python -m pipeline.run \
--input data/raw/pms_service_requests.csv \
--type csv \
--task service_request \
--field-map '{"input": "guest_message", "output": "resolution_notes"}'Challenge 3: More Data Was Not Automatically Better Data
The next trap was row count.
It is tempting to treat more examples as progress. In practice, weak data can make the model more generic, more repetitive, or more confident in the wrong behavior. The model does what the dataset teaches it. If the dataset teaches shallow templates, the model learns shallow templates.
The fix was to treat curation as an engineering control:
- Drop examples with private information.
- Drop malformed and truncated records.
- Drop vague or unsafe responses.
- Deduplicate near-identical records.
- Score examples for quality and safety.
- Keep the evaluation set frozen and separate from training.
- Use targeted augmentation only when it addresses observed gaps and does not overlap with evaluation.
Challenge 4: Training Was Easy to Launch, Hard to Trust
Once the data was ready, fine-tuning was straightforward enough to launch. Trusting the result was harder.
Without LLMOps discipline, model work becomes a pile of artifacts: one dataset here, one adapter there, an endpoint from last week, and a notebook only one person remembers how to run. That may be fine for a demo. It is not enough for a reusable industry model framework.
LLMOps gave us a control loop:
- Prepare and split the dataset.
- Upload training and evaluation data to Amazon S3.
- Launch QLoRA fine-tuning on Amazon SageMaker.
- Merge the fine-tuned adapter into the base model.
- Deploy the merged model to a SageMaker real-time endpoint.
- Evaluate the base model, fine-tuned model, and frontier baselines on the same scenarios.
- Test the deployed model interactively on realistic industry prompts.
Each run had a known dataset, known hyperparameters, known model artifact, known endpoint, and known evaluation set. When quality improved, we could connect the improvement back to the run. When a scenario failed, we could inspect whether the issue came from missing data, weak formatting, insufficient evaluation coverage, or deployment behavior.
For a walkthrough of the model lifecycle experience, see The LLM Command Center: Running the Full Domain-LLM Lifecycle in One Place .
Challenge 5: Evaluation Needed an Honest Story
An industry model should not be evaluated only by whether it sounds good. It has to be judged against the work it is supposed to perform.
For HospitalityLLM, we benchmarked on a frozen, held-out evaluation set of 278 examples that were never used in training. The judge was Claude Sonnet 4 on Amazon Bedrock at temperature 0. Deterministic checks were used for measurable categories such as numeric accuracy, policy-ID citation, and PII safety. LLM-as-judge scoring was used for comparative quality.
On a hospitality-scoped quality judge, the fine-tuned model reached 97.5% quality, with 99.3% of responses rated good at 4/5 or higher.
| Model | Overall quality | Latency p50 | Cost / request | Cost vs fine-tuned |
|---|---|---|---|---|
| Base Qwen3-8B | 34.5% | 21,396 ms | $0.00715 | 3.6x |
| Fine-tuned Qwen3-8B | 97.5% | 4,225 ms | $0.00199 | 1.0x |
| Claude Sonnet 4 | 87.8% | 4,886 ms | $0.00410 | 2.1x |
| Claude Opus 4.1 | 90.4% | 20,901 ms | $0.02160 | 10.9x |
The category view showed the same pattern:
| Category | Base | Fine-tuned | Claude Sonnet | Claude Opus |
|---|---|---|---|---|
| Factual accuracy | 28.4% | 96.0% | 89.2% | 91.4% |
| Brand voice | 31.2% | 97.0% | 85.4% | 89.8% |
| Authority | 33.6% | 98.5% | 86.8% | 89.2% |
| Empathy | 29.8% | 99.0% | 90.6% | 93.6% |
| Actionability | 30.4% | 97.5% | 90.8% | 92.4% |
| SOP compliance | 12.0% | 98.0% | 82.4% | 86.0% |
| Calculation | 18.5% | 95.5% | 84.6% | 87.8% |
| Safety, PII-clean | 92.4% | 98.5% | 92.6% | 92.8% |
These results should be read carefully. They do not mean an 8B model is generally better than frontier models. They mean an industry-tuned model can outperform larger general models on a focused, frozen industry benchmark when the data, evaluation, and operating context are aligned.
The honest caveats matter:
- The judge rewards fit to the hospitality reference style, so this measures industry fit, not general intelligence.
- Claude Sonnet 4 was both the judge and one of the benchmarked models. It did not top the table, but the caveat still matters.
- Grounding categories such as SOP citation and calculation benefited from the curated training data. Retrieval augmentation and deterministic tools remain important for production.
- No structured tool-use tasks existed in this corpus, so tool use was not scored.
- Safety was strong, but deterministic policy guardrails are still recommended for guest-data disclosure scenarios.
Challenge 6: A Model Artifact Was Not the Finish Line
The final challenge was consumption.
A model artifact in storage is not an operating capability. Business teams need a controlled way to test prompts, review behavior, and decide whether the model is ready for specific workflows. Those prompts may include sensitive operational context: customer issues, employee notes, safety incidents, contracts, pricing decisions, or policy exceptions.
By deploying the model behind an enterprise-controlled endpoint, the team can manage access, endpoint lifecycle, monitoring, data boundaries, and evaluation feedback.
This is where the work becomes more than a prototype. The model can be trained on industry data, hosted inside the enterprise boundary, evaluated against real workflows, and improved through the same LLMOps loop.
How Others Can Use the Same Framework
HospitalityLLM is one implementation. The framework is portable.
To adapt this approach to another industry:
- Choose the work, not the model first. Pick three to five high-value workflows where language, policy, and judgment matter.
- Define the industry task taxonomy. Create lanes such as service, operations, compliance, sales, support, risk, or field work.
- Identify approved source data. Start with first-party operational data, SOPs, tickets, transcripts, case notes, and expert-created examples.
- Create a training contract. Standardize every example into query, response, metadata, source method, complexity, and references.
- Run the curation gates. Scrub sensitive data, normalize formats, deduplicate, score quality, and probe for novelty.
- Freeze the evaluation set. Keep it separate from training so improvements mean something.
- Baseline before tuning. Compare the base model, tuned model, and frontier models on the same prompts.
- Deploy for controlled use. Host the model behind enterprise access controls and gather workflow feedback.
- Repeat with evidence. Use benchmark gaps to decide what data, retrieval, tools, or guardrails to add next.
This is why we call it a framework. The framework gives builders reusable assets: data contracts, curation gates, evaluation patterns, model-lifecycle automation, deployment choices, and feedback loops. The implementation changes by industry. The operating discipline stays the same.
What We Would Tell Builders
The biggest lesson is that industry model work is less about chasing a bigger model and more about building a better learning system.
Start with the business work. Define what good looks like. Build the data contract. Remove private information. Deduplicate aggressively. Keep the examples that teach useful behavior. Lock the evaluation set. Track every training run. Host the model where teams can test it safely. Then repeat.
HospitalityLLM is ultimately an industry data story. The model matters, but the durable advantage comes from the curated corpus, the frozen evaluation set, and the repeatable LLMOps framework that lets teams keep improving.
Try it Yourself
The sample project is available here:
The repository walks through the same pattern with a fictional hospitality brand: setup, data preparation, QLoRA training, model merge, deployment, evaluation, and testing. It uses Amazon SageMaker, Amazon Bedrock, Amazon S3, Jupyter notebooks, and a repeatable evaluation pattern comparing a frontier model, a base open model, and a fine-tuned model.
Replace the sample data with approved first-party industry data, keep the training and evaluation loop honest, and measure the model against the work your teams actually perform.
Series: industry-llm (1 article)
- 1HospitalityLLM: Building an Industry Model This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article