Most operators read a crew leaderboard the way they read a sales leaderboard: the name at the top is the best performer, and the name at the bottom needs a conversation. In janitorial work that reading is usually wrong, because a raw inspection ranking tells you more about which buildings you assigned than about who is cleaning them well.
The fix is not to throw the leaderboard out. It is to rank the numbers a crew can actually control, and to stop pretending a single average score is a performance verdict.
Review a crew leaderboard by ranking three controllable numbers instead of one: inspection score per budgeted labor hour, client complaints per 10,000 cleanable square feet, and rework hours. Require at least three inspections per crew per period, and compare crews only against accounts of similar difficulty.
Below are the four beliefs that break leaderboard reviews in cleaning companies, what actually holds instead, and the specific change to make on Monday.
Myth 1: The crew with the highest inspection score is your best crew
Inspection scores are heavily determined by the building, not the crew. A 60,000 sq ft Class A office suite with carpet throughout, low occupant density and a single break room will score higher than a 42,000 sq ft medical office plaza with vinyl composition tile, exam rooms, and eleven restrooms, even if the medical crew is objectively sharper.
APPA's custodial staffing framework makes the same point from the other direction: the appropriate level of clean, and the staffing required to hit it, changes by space type and use. A leaderboard that ignores space type is ranking your account mix.
What to do instead: divide score by labor hours
Score alone is an input measure. Score per budgeted labor hour tells you how much quality a crew is producing per hour you pay for. The absolute number is meaningless. The comparison is not.
Here is the arithmetic, using two illustrative accounts with assumptions stated up front.
- Riverside Office Park, Crew B: 60,000 sq ft, 5 nights a week, 3 cleaners at 5 hours each = 15 labor hours per night. Rolling inspection average: 94.
- Northgate Medical Plaza, Crew A: 42,000 sq ft, 5 nights a week, 3 cleaners at 4 hours each = 12 labor hours per night. Rolling inspection average: 88.
On the raw leaderboard, Crew B wins by six points. Run the division: 94 / 15 = 6.27 quality points per labor hour. 88 / 12 = 7.33. Crew A is producing about 17 percent more quality per paid hour, in a harder building.
That flip is the entire argument for reworking your leaderboard. If you have been coaching Crew A and holding up Crew B as the standard, you have been coaching backward.
The difficulty tier you can build in an afternoon
You do not need a statistician. Score every account 1 to 3 on four factors, add them up, and only compare crews inside the same band.
| Factor | 1 point | 2 points | 3 points |
|---|---|---|---|
| Restroom fixture density | Low, general office | Moderate | High, public or clinical |
| Occupant traffic during service | Empty building | Partial occupancy | Occupied, 24-hour or shift work |
| Floor type mix | Mostly carpet | Mixed carpet and hard | Mostly hard surface, VCT, terrazzo |
| Spec complexity | Standard nightly | Nightly plus periodic detail | Regulated, medical, food or lab |
Four to six points is Tier 1, seven to nine is Tier 2, ten to twelve is Tier 3. Publish three short leaderboards instead of one long dishonest one.
Myth 2: Inspection averages are objective enough to rank people on
Two inspectors walking the same 30,000 sq ft building will not produce the same score, and the same inspector will not produce the same score at 6 a.m. and at 9 p.m. after a bad client call. Inspection scoring is a judgment sampled a few times a month, and the sample is usually far too small to rank on.
Three specific distortions show up constantly in leaderboard data:
- Sample size: One inspection in a period is not a measurement. A 78 that came from a single walk on the night a cleaner called out is noise, not a trend.
- Inspector bias: If the account manager who sold the job also inspects it, scores drift up. If a supervisor inspects only when a client complains, scores drift down.
- Checklist drift: A 40-item checklist and a 12-item checklist do not produce comparable percentages, and averaging across them is arithmetic with no meaning.
What to do instead: set a floor on sample size and rotate who inspects
- Require a minimum of three completed inspections per crew per review period. If a crew has fewer, they appear on the list marked "insufficient data," not at the bottom.
- Use a rolling 13-week window for ranking and a 4-week window for coaching. One quarter smooths out staffing churn and holiday weeks. One month catches problems while they are still fixable.
- Standardize the checklist within a tier. Same item count, same weighting, same photo requirement for any item scored below full marks.
- Rotate inspectors so no crew is scored by the same person more than roughly two-thirds of the time. Cross-inspection between supervisors surfaces graders who are systematically soft.
- Track inspector average alongside crew average. If one supervisor's mean score is consistently well above every other supervisor's, the problem is in the grading, not in the cleaning.
Myth 3: Clock-in punctuality and hours on site measure productivity
Time data is the easiest thing to put on a leaderboard, which is exactly why it ends up ranked. On-time clock-in percentage is a real operational metric. It is not a productivity metric, and it should never be the column that decides who wins the month.
A cleaner who clocks in on time every night and leaves 40 minutes of the spec undone outranks a cleaner who hits traffic twice a month and finishes everything. Ranking on punctuality tells your crews that arriving matters more than cleaning.
There is also a technical limit worth being honest about. Geofenced clock-in captures a point in time and a location at the start and end of a shift. It does not tell you what happened in between, and treating it as a proxy for effort is a mistake operators make constantly.
What to do instead: use time data as a gate, not a rank
Set punctuality as a pass/fail threshold for leaderboard eligibility, then rank on quality output. For example: a crew must be at or above 95 percent on-time clock-in and have zero missed clock-outs in the period to be ranked at all. Below the gate, they get a scheduling review instead of a rank.
Then use hours as a denominator, not a score. The two time-based numbers that genuinely belong on the review are:
- Hours variance: actual clocked hours divided by budgeted hours. Above 1.0 means the account is eating margin. Well below 1.0 with falling scores means the spec is not being completed.
- Rework hours: hours spent returning to correct work that was already billed. This is the single most under-tracked number in the industry and one of the most diagnostic.
For loaded labor cost in the same calculation, use your own burdened rate. BLS publishes median hourly wage data for janitors and cleaners in its Occupational Employment and Wage Statistics, which is a reasonable sanity check on your base wage, but your burden, overtime and turnover cost are yours alone.
Myth 4: Posting the leaderboard publicly motivates the crew at the bottom
Public rankings work well when everyone is running the same race. Cleaning crews are not. Publishing a company-wide list where the same three names sit at the bottom every month, largely because they were handed the hardest buildings, produces resentment and turnover in an industry that already fights both.
It also invites gaming. Once crews understand what is scored, they optimize for the scored items. Restroom mirrors get spotless while the stairwell nobody inspects goes three weeks. Crews quietly lobby to be moved off Tier 3 accounts. Supervisors stop logging small complaints because logging them hurts their own team's number.
What to do instead: rank improvement, publish selectively, coach privately
- Publish the top of the list, not the bottom. Recognize the top three within each difficulty tier. Deliver the bottom of the list one-to-one.
- Add an improvement column and weight it. Change versus the crew's own trailing 13-week baseline lets a crew at 82 that climbed from 74 beat a crew that has sat flat at 91 all year. That is the behavior you actually want.
- Randomize inspected zones. If the checklist rotates through zones unpredictably, optimizing for the inspection and doing the whole job converge.
- Never rank on complaint counts alone. Normalize complaints per 10,000 cleanable square feet per month, or you are punishing whoever cleans the biggest building.
What actually holds: the four columns worth ranking
Strip the leaderboard down to metrics a crew controls, that survive a difficulty adjustment, and that resist gaming. Four columns is enough.
| Metric | Formula | What it actually tells you | How it gets gamed |
|---|---|---|---|
| Quality per labor hour | Rolling 13-week inspection average / budgeted labor hours per service | Quality produced per paid hour, comparable inside a difficulty tier | Under-budgeting hours on paper to inflate the ratio |
| Complaint rate | Client complaints / (cleanable sq ft / 10,000) per month | What the client experiences between inspections | Supervisors logging complaints as "verbal, no action" |
| Rework hours | Hours returning to redo billed work / total hours worked | True cost of quality failures, including the ones nobody complained about | Rework recorded as regular hours |
| Improvement delta | Current 4-week average minus trailing 13-week baseline | Whether coaching is landing, independent of building difficulty | Sandbagging an early period to create headroom |
Every one of those gaming risks is manageable if you know it exists. Lock budgeted hours to the signed spec rather than to whatever the crew leader submits. Require every client contact to be logged, positive or negative. Create a rework time code and make it non-optional.
The 45-minute monthly leaderboard review
- Pull inspection scores, clocked hours and budgeted hours for the trailing 13 weeks, one row per crew per account.
- Flag any crew with fewer than three inspections and mark it "insufficient data" rather than ranking it.
- Apply the punctuality gate: below 95 percent on-time or any missed clock-out means a scheduling review, not a rank.
- Split the list into the three difficulty tiers and rank inside each tier only.
- Calculate quality per labor hour, complaint rate per 10,000 sq ft, rework percentage and improvement delta.
- Compare inspector averages against each other and note any grader more than a few points off the group.
- Pick exactly two crews to coach this month and one specific behavior for each. Not five.
- Write down what you expect to change by the next review, then check that prediction next month.
That last step is what separates a review from a ritual. If you cannot state in one sentence what you expect a number to do next month, you are not reviewing performance. You are reading a report.
Where CleanTrack360 fits
Every calculation above needs three raw inputs: inspection scores tied to a specific crew and account, clocked hours you can compare to budgeted hours, and a way to get all of it into a spreadsheet. CleanTrack360 handles the collection side. Quality inspections run on custom checklists with photo evidence and automatic scoring, geofenced GPS clock-in and clock-out runs in the crew's phone browser with a default 150 m radius you can configure per location, and reports export to CSV so you can build the quality-per-labor-hour and improvement-delta columns yourself.
Worth being precise about the limits: location is captured at clock-in and clock-out only, not tracked between them, so time data supports the punctuality gate and the hours denominator, nothing beyond that. Plans are Starter at $99/month for up to 5 team members, Pro at $199/month for up to 20, and Business at $249/month for up to 50, priced per plan rather than per user. There is a 14-day free trial and no credit card required, which is long enough to run one full inspection cycle and see what your real numbers look like.