How Accurate Is Your Forecast? Benchmarking MAPE Across Hotel Departments
Every hotel produces a forecast and almost none of them grade it. This is the owner's guide to scoring the forecast the way you score RevPAR: which error metric to use, what an acceptable error band looks like by department and horizon, how to catch the bias that sank the 2025 budget cycle, and how to put a dollar figure on every point of error before it lands in the labor line.
The forecast nobody grades
Every hotel produces a forecast. Almost none of them score it. The revenue manager submits a 90-day rooms forecast on the first of the month, the director of finance rolls it into a P&L reforecast, the executive housekeeper turns it into a schedule, the executive chef turns it into a purchase order, and the general manager presents it to ownership. Thirty days later the actuals arrive, everyone looks at the variance to budget, and the forecast itself quietly disappears. Nobody asks how far off it was, whether it was off in the same direction it was off last month, or which department paid for the miss.
That is the gap this article is about. Forecast accuracy is the one operating metric that touches labor, purchasing, pricing, and owner credibility simultaneously, and it is the one metric almost no independent hotel tracks with any discipline. The 2025 budget cycle made the cost of that visible across the whole industry. According to HotStats data published by HotelData.com, US rooms revenue finished the third quarter of 2025 13.2% below budget and 11.9% below budget year to date. Even after hotels reforecast mid-year, rooms revenue still trailed the revised forecast by 4.7% in the quarter and 5.2% year to date. The forecasts were optimistic in November, still optimistic in June, and nobody had a mechanism to catch the bias until the money was gone.
Meanwhile the labor line, which is the largest controllable cost in the building and the one most directly driven by the forecast, kept rising. Hotel Management reports labor cost per occupied room at $48.32 and climbing at double-digit rates at many full-service properties, which is why a 5% scheduling error at a 300-key hotel is a six-figure problem. The 2025 Hotel Labor Costs & Trends Report found that operators protected margin in 2025 not by cutting teams but by cutting hours per occupied room 7% to 15% in guest services, housekeeping, and management. You cannot cut hours to demand you did not forecast correctly. Precision in labor is downstream of precision in forecasting, and precision in forecasting starts with measuring the error.
What follows is a practical framework for scoring the forecast: which error metric to use and when, what an acceptable error band looks like by department and by horizon, how to detect the systematic bias that MAPE hides, and how to translate a point of forecast error into dollars of labor, food, and revenue so the whole executive committee understands what accuracy is worth. It ends with a 90-day implementation plan that a revenue manager and a controller can run without buying anything.
A forecast that is never scored is not a forecast. It is a guess with a spreadsheet around it. The moment you start grading it, it starts getting better, because the people who make it finally know what better means.
Why hotels do not score their forecasts
The reasons are structural, not personal. First, the forecast has too many owners. Rooms belongs to revenue management, F&B covers belong to the outlet managers, banquet revenue belongs to catering sales, and the consolidated P&L forecast belongs to finance. Each group produces its number on its own cadence, in its own spreadsheet, using its own definition of the forecast date. There is no single version to score, so nobody scores any of them.
Second, the forecast is overwritten. The 90-day forecast submitted on September 1 for October is replaced by the one submitted on October 1, which is replaced by the 10-day forecast, which is replaced by the day-of pickup. By the time October actuals land, the original September 1 view no longer exists in the system. You cannot measure the error of a number you did not keep. This is the single most common reason forecast accuracy programs fail on day one: the property has no archive of what it said it would do.
Third, the metric is politically uncomfortable. A revenue manager who is scored on RevPAR performance has every incentive to forecast conservatively so that actuals beat the number, and a director of sales who is scored on booking pace has every incentive to forecast group aggressively so that the pipeline looks healthy. Both are rational, and both introduce bias that compounds when finance sums them. Scoring the forecast exposes those biases, which is precisely why it is valuable and precisely why it is resisted.
Fourth, most hotels have never seen what good looks like. Nobody has told the GM that a 30-day rooms forecast at 8% MAPE is excellent and 20% is a problem, so there is no target to manage to. The rest of this article fixes that.
Choosing the right error metric
MAPE, mean absolute percentage error, is the default for a reason. It is intuitive, it is unit-free, and a GM can understand "we were off by 12% on average" without a statistics lesson. Calculate it by taking the absolute difference between forecast and actual for each day, dividing by the actual, averaging across the period, and expressing the result as a percentage. A 30-day forecast that predicts 220 rooms on a night that actualizes at 200 contributes a 10% error for that day.
But MAPE has well-documented failure modes that matter in a hotel. The Institute of Business Forecasting and practitioners like Demand Planning LLC point out that MAPE explodes on low-volume days, because dividing a small absolute error by a small actual produces a huge percentage. A resort forecasting 6 banquet covers on a Tuesday that actualizes at 3 has a 100% error on that day, which will swamp the month. MAPE also treats over-forecasting and under-forecasting asymmetrically: you cannot be more than 100% under, but you can be infinitely over. And, most important for a hotel, a forecast can carry a respectable MAPE while being consistently wrong in the same direction, because the absolute value strips the sign.
The fix is not to abandon MAPE but to pair it with two companions. The first is weighted MAPE, usually written WMAPE or WAPE, which sums the absolute errors and divides by the sum of actuals. It weights busy days more heavily than slow days, which is what you want: being off by 20 rooms on a 300-room sellout night matters more than being off by 2 rooms on a 30-room night. WFM Labs recommends WAPE as the primary accuracy metric for workforce planning for exactly this reason. The second companion is bias, sometimes called mean percentage error, which is the same calculation as MAPE without the absolute value. Bias tells you the direction of the error. A forecast with 12% MAPE and +9% bias is over-forecasting almost every day. That is a much more actionable finding than 12% MAPE alone, and it is the finding that explains the 2025 budget season.
| Metric | What it measures | Best hotel use | Known weakness |
|---|---|---|---|
| MAPE | Average absolute error as a percentage of actual, day by day | Rooms occupancy and revenue at 7, 30, and 90 days; executive reporting | Explodes on low-volume days; hides direction of error |
| WMAPE / WAPE | Total absolute error divided by total actual for the period | Labor scheduling, F&B covers, any department with high day-to-day variance | Less intuitive to explain; can mask a bad week inside a good month |
| Bias (MPE) | Signed average error; positive means over-forecast, negative means under | Detecting systematic optimism or sandbagging by forecast owner | Positive and negative errors cancel; must be read alongside MAPE |
| Tracking signal | Cumulative signed error divided by mean absolute deviation | Monthly control chart; flags when a forecast has drifted out of tolerance | Requires a running history; slow to reset after a regime change |
| MAE / RMSE | Error in units (rooms, covers, dollars) rather than percent | Low-volume outlets, spa, banquet covers where MAPE is unstable | Not comparable across departments of different size |
The practical recommendation for a full-service hotel is simple. Report MAPE and bias for rooms at every horizon, report WMAPE and bias for every labor-driving department, and use MAE in units for any outlet or revenue center that regularly sees single-digit daily volume. Present all of them on one page. The point is not statistical purity. The point is that the executive committee sees the same scorecard every month and knows which number to react to.
Accuracy benchmarks by department and horizon
The honest answer to "what MAPE should we hit" is that it depends on the horizon, the department, and the demand pattern of the property. But an answer that depends on everything is useless to an owner, so here are the bands that hold across enough properties to be a fair starting point. Cross-industry guidance summarized by Hospitality Net and Lighthouse puts a MAPE under 10% as excellent, 10% to 20% as acceptable, and 15% to 25% as the realistic range for seasonal or event-driven hotels. A large city-center hotel with steady corporate demand can regularly run 10% to 12% on a 30-day rooms forecast; a 50-room independent in a ski town may find 20% is as good as it gets in shoulder months. Cloudbeds makes the same point from the demand-surface side: accuracy should be measured against the horizon at which the decision is made, not at a single arbitrary point.
Academic benchmarks are tighter than operating benchmarks because they measure the model, not the whole process. The foundational Weatherford and Kimes comparison of forecasting methods, run on Choice and Marriott data, found pickup methods, exponential smoothing, moving averages, and regression to be the most robust, and hotels still using a simple pickup model have not been proven wrong by two decades of subsequent research. More recent work published in Tourism Management shows ensemble models reaching MAPE below 5.1% on daily occupancy and cutting error more than 80% against a naive benchmark, and the Boston Hospitality Review reports similar single-digit results for well-specified daily models. Those numbers are achievable. They are also achieved by a model with clean data and no committee editing its output. The operating benchmarks below are for the forecast that actually reaches the schedule.
| Department / measure | 7-day horizon | 30-day horizon | 90-day horizon | Action threshold |
|---|---|---|---|---|
| Rooms occupied (transient-heavy urban) | Under 5% | 8-12% | 12-18% | Over 15% at 30 days |
| Rooms occupied (resort / seasonal) | Under 7% | 12-18% | 18-25% | Over 22% at 30 days |
| Rooms revenue (occupancy x ADR) | Under 6% | 10-15% | 15-22% | Over 18% at 30 days |
| Restaurant covers by meal period | 8-12% | 15-20% | Not forecast | Over 25% at 7 days |
| Banquet covers (definite business) | Under 5% | Under 8% | 10-15% | Over 10% at 7 days |
Three things about that table deserve emphasis. First, the 7-day rooms number is the one that drives the housekeeping schedule, and it should be under 5% at almost any hotel because most of the business is already on the books; a 7-day rooms MAPE above 8% is almost always a data hygiene problem, not a demand problem. Second, restaurant covers are the hardest line in the building to forecast because they depend on in-house capture, local walk-in, weather, and competitive openings, none of which the PMS sees. That is why they are the first place to deploy a model with external inputs. Third, banquet definite business should be nearly perfect a week out, and if it is not, the problem is the BEO process, not the forecast.
Detecting bias: the error that MAPE hides
Return to the 2025 numbers. Rooms revenue was 11.9% below budget year to date, and the industry-wide reforecast was still 5.2% too high. Those are not random errors. A random error would be over some months and under others and would average close to zero. A forecast that is over by 5% to 13% every month for three quarters has a bias, and the bias has a cause: an assumption baked in during budget season, in this case a double-digit RevPAR growth expectation carried into a year that actually produced the first non-recessionary RevPAR decline on record, as reported by Hotel Online from the STR and Tourism Economics forecast. The assumption was never revisited because nobody was measuring the direction of the error.
The tool for catching this is the tracking signal, borrowed from supply chain control charts and explained well by Arkieva and SupliiChain. Sum the signed errors over the last several periods, divide by the mean absolute deviation, and watch the result. A forecast with no bias hovers near zero. When the tracking signal moves beyond roughly plus or minus 4, the forecast has drifted out of control and needs to be recalibrated. The beauty of the tracking signal is that it catches a small consistent bias long before it becomes a large one. A forecast that runs 3% high every month has a MAPE that looks fine and a tracking signal that climbs steadily until someone has to explain it.
In a hotel, bias almost always has a human signature. Revenue managers under-forecast so that they beat the number. Sales directors over-forecast group so that the pipeline looks full. Outlet managers over-forecast covers so that they are staffed for the good night. Finance smooths everything toward budget so that the owner call is calm. None of these people are doing anything wrong by their own scorecard. The forecast accuracy program exists to give them a shared scorecard, and the bias column is where the conversation starts.
| Bias pattern | Typical signature | Usual root cause | Corrective action |
|---|---|---|---|
| Persistent rooms over-forecast | Positive bias at 30 and 90 days, near zero at 7 days | Budget growth assumption carried into monthly forecast; unrealistic pickup curve | Rebase pickup curve on trailing 12 months; separate budget from forecast in the system |
| Persistent rooms under-forecast | Negative bias at all horizons; actuals beat forecast every month | Revenue manager sandbagging against a performance target | Score the revenue manager on accuracy and bias, not on beating forecast |
| Group over-forecast | Large positive bias in group room nights; tentatives counted as definite | Sales pipeline optimism; wash factor not applied | Apply historical wash by market segment; forecast tentatives at weighted probability |
| Cover over-forecast in outlets | Positive bias on restaurant and bar covers, especially weekdays | Outlet managers staffing for the best case; in-house capture rate overstated | Forecast covers from actual capture rate by day of week; tie schedule to WMAPE |
| Bias that flips at month end | Under-forecast in weeks 1-3, over-forecast in week 4 | Month-end pressure to hit the number; reforecast edited by finance | Lock the forecast at submission; archive every version before finance edits |
Translating error into dollars
The reason forecast accuracy is worth an owner's attention is that every point of error has a cost, and the cost lands in three places: labor scheduled for demand that did not arrive, food purchased for covers that did not show, and revenue left on the table when the forecast was low and the hotel priced or staffed as if demand was soft. The labor cost is the largest and the easiest to compute. The collaborative study on staffing demand forecasting in the hotel industry and Deputy's guidance on data-driven hotel staffing both make the same point: labor is scheduled against the forecast, not the actual, so forecast error converts almost directly into unproductive hours when the forecast is high and into overtime and service failure when it is low. The 2026 workforce management outlook from Hotel Online notes that AI scheduling tools have lifted forecast accuracy from the mid-50s to the high 80s in percentage terms and cut overtime 20% to 40% at adopting properties, which is a direct measure of what the error was costing before.
The food cost is smaller per point but compounds fast in a hotel with multiple outlets and banquets. The Champions 12.3 study of hotel kitchens found that every dollar invested in reducing kitchen food waste returned an average of seven dollars in operating savings, and that participating properties cut food waste by at least 10% and in some cases lowered food cost by 3 points or more. A case study of hotel food service waste published in Sustainability found that the majority of waste occurs at the serving stage, in buffets and plate waste, which is exactly the waste that a better cover forecast prevents, because production is set the day before against the forecast. Orbisk reports that hotels achieving 30% to 53% waste reductions all measure by outlet and meal period, which is the same granularity a good cover forecast requires.
Here is the worked translation for a hypothetical 250-room full-service hotel at 70% occupancy, $220 ADR, one restaurant, one bar, and a modest banquet operation, using the labor CPOR from Hotel Management and typical food cost and waste figures from the studies above. The point of the exercise is not the exact dollars, which will differ at every property. The point is that a controller can rebuild this table for their own hotel in an afternoon and put a number on the accuracy program before the first meeting.
| Cost of a 5-point forecast error | Mechanism | Annual exposure | Recoverable with accuracy program |
|---|---|---|---|
| Housekeeping and front office labor | Rooms over-forecast schedules roughly 5% surplus hours against $48 labor CPOR on 63,875 occupied rooms | $150,000-$190,000 | 60-70% |
| Overtime and agency from under-forecast | Rooms under-forecast triggers premium hours at 1.5x on 8-12 days per quarter | $45,000-$70,000 | 50-60% |
| Restaurant and banquet food waste | Cover over-forecast drives overproduction at the serving stage; 5-point error on $2.4M F&B revenue at 32% food cost | $35,000-$55,000 | 40-50% |
| Revenue displacement from under-forecast | Low forecast keeps rates and restrictions loose on 15-20 nights that would have compressed | $60,000-$110,000 | 30-40% |
| Owner and lender credibility | Missed reforecast three quarters running triggers covenant reviews and asset manager intervention | Not quantified | High |
Summed, a 5-point forecast error at this hotel carries somewhere between $290,000 and $425,000 of annual exposure before the credibility line, and a disciplined accuracy program plausibly recovers half of it. That is a return worth chasing with a spreadsheet and a monthly meeting, which is all the first phase requires.
Owners do not fire general managers for missing the forecast. They fire them for missing it in the same direction three quarters in a row and not knowing why. Bias is the number that ends careers, and it is the one number almost nobody reports.
Where AI actually helps, and where it does not
The forecasting vendor market has been loud about machine learning for a decade, and some of the claims are true. McKinsey's widely cited estimate, summarized by ToolsGroup and ThroughPut, is that AI-driven forecasting reduces error 20% to 50% against conventional methods, and the mechanism is the number of inputs. A spreadsheet pickup model uses three to five variables; a gradient-boosted or ensemble model can ingest 20 to 40, including weather, local events, competitor rates, airline capacity, search trends, and cancellation velocity. In hospitality specifically, the systematic review of deep learning for resort demand forecasting and the extended additive pickup work for SME hotels both show meaningful error reduction from richer inputs, with the important caveat that the gains are concentrated at the 30-day and 90-day horizons where the booking curve is still forming. At 7 days, most of the business is on the books and a good pickup model is already close to the ceiling.
The part vendors say less about is that a model cannot fix a process. Steve Morlidge's study of eight companies, published in Foresight, found that 52% of their forecasts were worse than a naive random walk, meaning that all the human adjustment, committee review, and executive override layered on top of the statistical forecast made it worse more often than better. Forecast value added analysis, laid out step by step in SAS's white paper and the IBF series on FVA, measures each step of the process against the step before it and asks a blunt question: did this person or this meeting make the forecast more accurate, or less? Hotels that run FVA typically discover that the revenue manager's statistical forecast is the most accurate number in the building and that every subsequent edit, from the DOSM's group adjustment to the GM's gut check to finance's smoothing, degrades it. That is not an argument against human judgment. It is an argument for measuring it.
So the honest sequencing is this. First, archive and score the forecast you already have. Second, run FVA to find out which steps of your process help and which hurt, and remove the ones that hurt. Third, and only third, replace the statistical core with a model that uses more inputs, because by then you will have a clean baseline to prove it against. Hotels that skip to step three buy a better engine and bolt it to the same leaky process, and then conclude that AI does not work. Properties that want help with the modeling layer often find it useful to start with a structured forecasting assessment rather than a software purchase; that is the work our AI Revenue Optimization & Forecasting engagement is built around, and the first deliverable is always the accuracy baseline described in this article.
A 90-day implementation plan
None of this requires new software. It requires a controller, a revenue manager, a shared spreadsheet, and a standing 30-minute meeting. The plan below is the one we run with independent hotels, and the milestones are deliberately unambitious in the first month because the archive is the hard part and everything else depends on it.
| Phase | Weeks | Owner | Deliverable | Success measure |
|---|---|---|---|---|
| 1. Archive the forecast | 1-2 | Revenue manager | Every forecast version snapshotted on submission date, by day, by segment, by department, in one table | Zero overwritten forecasts after week 2 |
| 2. Build the scorecard | 3-4 | Controller | One-page monthly report: MAPE, WMAPE, bias, tracking signal at 7, 30, 90 days for rooms and each labor-driving department | First scorecard reviewed at exec committee in week 5 |
| 3. Run forecast value added | 5-8 | Revenue manager + DOF | Error measured at each process step: statistical forecast, RM adjustment, sales adjustment, GM review, finance reforecast | Each step classified as adding or destroying accuracy |
| 4. Fix the process | 9-10 | General manager | Steps that destroy accuracy removed or made advisory; forecast owners scored on accuracy and bias, not on beating the number | Bias within plus or minus 3% at 30 days |
| 5. Tie to labor and purchasing | 11-13 | DOF + department heads | Schedules and purchase orders generated from the archived forecast; error cost reported monthly in dollars | Hours per occupied room and food cost variance tracked against forecast error |
Two operating rules make this stick. The first is that the forecast is locked at submission. Finance can produce a separate reforecast, sales can produce a separate pipeline view, and the GM can annotate all of it, but the number the revenue manager submitted on the first of the month is the number that gets scored, and it does not change. The second is that accuracy and bias go on the revenue manager's and the department heads' scorecards with a weight that matters. The retail planning literature is unambiguous that planners score on what they are measured on, and hotel forecast owners are no different. Score them on beating forecast and they will sandbag. Score them on accuracy and they will forecast.
What the scorecard should look like
The monthly scorecard fits on one page and answers four questions in order. How accurate was the forecast at each horizon that drives a decision? Was it biased, and in which direction? Which step in the process added or destroyed accuracy? What did the error cost? A version we use lists rooms occupied, rooms revenue, restaurant covers by meal period, banquet covers, and total labor hours down the left side; shows 7-day, 30-day, and 90-day MAPE or WMAPE with the bias in parentheses across the top; colors any cell outside the action threshold from the benchmark table; and closes with the tracking signal trend and the dollar translation. It takes a controller two hours a month once the archive exists.
The scorecard changes the conversation at the executive committee. Instead of "why did we miss budget," the question becomes "the 30-day rooms forecast has run 6% high for four months, the tracking signal is at 5.1, and it is costing us roughly $14,000 a month in housekeeping hours; what assumption are we carrying that is wrong?" That is a question with an answer, and it is a question that makes the revenue manager, the controller, and the department heads allies rather than suspects.
Hotels that have run this for a year report the same three outcomes. Rooms MAPE at 30 days drops into the low teens or better within two quarters, mostly from removing bias rather than from any modeling change. Hours per occupied room falls because the schedule is built from a number people trust. And the owner call gets shorter, because the general manager arrives with a forecast that has a track record instead of a forecast that has a story. Forecast accuracy is not a revenue management metric. It is the operating discipline that makes every other metric in the building believable.
Frequently Asked Questions
What is a good MAPE for a hotel rooms forecast?
It depends on the horizon and the demand pattern, but the operating bands that hold across most full-service hotels are under 5% at 7 days, 8% to 12% at 30 days, and 12% to 18% at 90 days for a transient-heavy urban property, with resort and seasonal hotels running roughly 4 to 7 points wider at the longer horizons. Cross-industry guidance treats under 10% as excellent and 10% to 20% as acceptable, and most hotels that have never scored their forecast discover they are running 15% to 25% at 30 days. The number matters less than the trend and the bias: a hotel at 14% MAPE with zero bias is in far better shape than one at 11% MAPE that is over-forecasting every month. Set the first target as "measured and improving," then tighten to the benchmark bands once the archive has six months of history.
Should we measure accuracy on the forecast the revenue manager submits or the one finance publishes?
Both, separately, and that is the point. The revenue manager's statistical forecast and the finance reforecast are two different numbers produced for two different audiences, and forecast value added analysis exists to measure whether the edits between them helped or hurt. In practice the revenue manager's forecast is usually the more accurate one, and the finance version has been smoothed toward budget. Archive every version on its submission date, score each against the same actuals, and report the difference. If finance's edits are consistently degrading accuracy, finance should still publish its reforecast for the owner, but it should be labeled as a financial view, and the operating departments should schedule and purchase from the revenue manager's number.
How do we forecast accuracy for a restaurant that has walk-in business the PMS cannot see?
Forecast covers from two components and score them separately. In-house covers are a capture rate applied to the rooms forecast by day of week and meal period, and that capture rate is stable enough to model from six months of POS data tied to room charges. Local and walk-in covers are a separate series driven by day of week, weather, local events, and seasonality, and this is the series that benefits most from a model with external inputs. Use WMAPE rather than MAPE for the total because slow Tuesday lunches will otherwise dominate the error, and use MAE in covers for any meal period that regularly runs below 20 covers. The production schedule and the kitchen prep list should be generated from the combined forecast, and food waste by meal period should be logged against the forecast error so the chef can see the connection.
Our PMS and RMS overwrite the forecast every night. How do we build the archive?
Export it. Nearly every PMS and RMS can produce a daily forecast report on a schedule, and the entire archive is a matter of saving that export with the run date in the filename and appending it to one table. A revenue analyst can automate this with a scheduled report and a simple script in a day; a hotel without an analyst can do it manually in ten minutes each morning until the process is proven. The columns you need are forecast run date, stay date, segment, rooms, ADR, revenue, and, for departments, covers or hours. Do not wait for the RMS vendor to build an accuracy module; several now offer one, but the archive you own is the one you can audit, and it is the one that lets you score the forecast against every step of your own process rather than just the model's output.
Is it worth buying an AI forecasting tool before we have done this?
Usually not, and the reason is that you will have no way to prove it worked. The published error reductions from machine learning forecasting, in the 20% to 50% range, are measured against a clean baseline, and a hotel without an archived, scored forecast has no baseline. Worse, the process problems that degrade a spreadsheet forecast will degrade a model's output just as effectively if the same committee edits it afterward. Run the 90-day plan first. It costs nothing but time, it typically removes most of the bias on its own, and at the end of it you will know exactly where the remaining error sits, which is the specification you need to evaluate a tool. Hotels that arrive at the vendor conversation with six months of scorecards negotiate from strength and can hold the vendor to a measurable improvement rather than a demo.
Peter Mack is a hospitality technology strategist and founder of HospitalityOS, helping independent hotels and resorts implement AI systems that drive revenue and reduce operational costs. With 25 years in hospitality operations and technology, he has worked with properties of all types and in every region as both a General Manager, Founder, Operator, Asset Manager, and Owner.