Rectifia
← ALL POSTS
PRODUCT

Why Similar HR Cases Get Different Outcomes - and How to Fix It

August 12, 2026 · 7 min read

Picture two cases that land on the same HR Director's desk six months apart.

Case one: a mid-level manager yells at a direct report in front of the team, once, documented by two witnesses. He gets a written warning.

Case two: a different manager does almost the exact same thing - same severity, similar evidence, comparable department. She gets a "verbal coaching conversation" that isn't even written down.

Nobody decided to treat these differently on purpose. Different investigator, different day, different mood in the room, maybe a manager who happened to be a top performer and got a slightly softer read. That's not malice. It's just what happens when every case gets judged in isolation, by a person, without a memory of the fifty cases that came before it.

And it's exactly the pattern that turns into a discrimination lawsuit eighteen months later, when someone in a protected class points out that the person who got the lighter treatment doesn't look like them.

The problem nobody's software actually solves

Most workplace case management tools - NAVEX included - are built to help you handle one case well. Intake, routing, documentation, closure. That's useful, but it's solving for the wrong unit of analysis. A single case can be handled perfectly and your organization can still be dangerously inconsistent, because consistency isn't a property of one case. It's a property of the pattern across all of them.

The tools that exist today have no memory. Each investigation starts from zero.

What we actually built

Every time a case closes in Rectifia, it gets stored as a reference point - category, severity score, evidence score, department, and the action that was taken. No names, no narrative, no identifying detail. Just the shape of the decision.

When a new case comes in that looks similar - same category, severity within a configurable band, comparable department tier - the system checks it against that history. If the action an investigator is about to take looks meaningfully harsher or more lenient than what similar cases got, it surfaces a flag: this deviates from your typical pattern, in this direction, review before closing.

That's it. That's the whole mechanism. It doesn't tell the investigator what to do instead. It doesn't say the current decision is wrong. It just makes the pattern visible at the one moment it actually matters - before the case closes, while the decision is still open.

A few things we were deliberate about:

  • It flags both directions. Harsher than usual and more lenient than usual both get flagged. It would be easy to build a tool that only catches "too soft" - that's the version that looks good in a compliance deck but quietly encourages harsher discipline across the board. We didn't want that bias built into the bias checker.
  • It needs enough data before it says anything. Fewer than five comparable reference cases and the system just says "insufficient data" instead of forcing a comparison on a thin sample. A flag based on two prior cases isn't a pattern, it's noise, and treating it like a pattern would erode trust in the flag fast.
  • The AI never recommends an action. This is the same boundary that runs through the whole product - worth repeating because it's the part people assume must be different. The engine surfaces a deviation. A human decides what, if anything, to do about it. If your legal team asks whether the AI is making disciplinary decisions, the honest answer is no, and the codebase backs that up.

Why this matters more than it sounds like it does

Most HR leaders I've talked to about this nod along and then say some version of "sure, but how often does that actually come up." Fair question. Here's the uncomfortable answer: you don't usually find out it came up until someone's employment lawyer finds it for you.

Inconsistent discipline is one of the more common threads in workplace discrimination claims, precisely because it's invisible from the inside. Every individual decision felt reasonable in the room. It's only when you line up ten closed cases side by side that the pattern shows up - and by then, it's evidence in a deposition, not a flag an investigator got to see before closing the file.

There's also a quieter cost that doesn't show up in litigation: trust. Employees talk to each other. When word gets around that two people got wildly different consequences for the same thing, the reporting channel itself loses credibility, whether or not anyone can prove bias formally. People stop reporting things they think won't be handled fairly. That's the actual failure mode a Consistency Engine is trying to prevent - not just the lawsuit, but the slow erosion of "this system actually works" that happens long before anyone sues.

The honest limitation

New company, no history, no reference cases yet - the engine has nothing to compare against for a while. That's a real cold-start problem and we don't pretend otherwise. Early on, the value builds gradually as your own case history accumulates. Longer term, an opt-in, k-anonymity-protected industry benchmark pool could help newer companies get useful signal sooner - but that's a future idea, not something live today, and we won't claim it is.

If you're evaluating case management software and this is the first time you're hearing "consistency engine" as a category, that's because - as far as we can tell - nobody else in this market has built one. Worth asking your current vendor, or the next one you demo, what happens when two similar cases get two different outcomes. If the answer is "nothing, we don't track that," you've just found the gap.