How we measure ourselves
Most tools in this space will tell you they use AI to help with grievance intake. Very few will tell you how often they get it wrong.
We run a fixed set of test cases through the same pipeline a member’s real intake goes through, and we score the output against answers written before the test ran. When we change something, we run it again. Every result — the good ones and the bad ones — is recorded.
Of 908 citations produced in our strongest recorded run, zero could not be matched back to the source they claimed.
Every quote our system attributes to a decision, a contract, or a member’s own words is checked against that source, character by character. A quote that cannot be found is flagged as unverified.
That check is not an AI model deciding whether it feels right. It is a literal text match against the source the citation names.
Why this is the number we lead with. A steward can work with a gap. A steward cannot work with a citation that looks authoritative and isn’t. A case file that overstates its own support is worse than one that admits what it doesn’t have.
Two configurations, both recorded. The first is what runs in production. The second is the strongest result we have measured.
Production configuration
2.41%
Citations that fail verification
Down from 7.35% on the same generation model before our source-filtering work shipped.
9 / 10
Fabrication resistance, most recent run
Strongest recorded configuration
0.00%
Citations that fail verification
908 citations produced, none unmatched — the only perfect-fidelity run we have recorded.
10 / 10
Fabrication resistance
Figures from recorded runs, August 2026. We publish the production number alongside the best one because the gap between them is the honest picture.
Five categories, scored on every run.
Does the system invent a case, a quote, or a contract clause that does not exist? Every run includes tests designed specifically to invite one.
Of every citation produced, how many fail verification? This is the measure we watch most closely, because it is the one a steward feels directly.
When the law cuts against the member, does the system say so? Burying unfavorable authority would make a case file feel stronger and be worth less.
Is the analysis anchored in the member’s actual account and documents, rather than a plausible generic narrative?
Did the system find the specific decisions a knowledgeable federal-sector practitioner would reach for? This is our hardest measure and our lowest score — and we report it rather than quietly dropping it.
Scores are only as trustworthy as the process that produces them. This is the sequence every case goes through.
Nothing in this pipeline decides the merits of a case, predicts an outcome, or takes an action. Union staff and counsel review every file. The system’s job is to make that review faster and better-sourced — not to replace it.
Social Security numbers, bank and account numbers, government ID numbers and dates of birth are detected in every uploaded document and replaced before that text is sent for analysis. The original stays in the union’s encrypted storage. The member’s name, email and account are never included in what the system sends.
The member’s words, their documents, the union’s contract, and a library of published federal decisions.
Each quoted passage is matched character-for-character against the named source.
Authority that limits or cuts against the member’s position is identified in the same pass, held to the identical verification standard, and shown only to your designated national and legal reviewers.
Every benchmark run is kept — including the runs that scored worse than the one before. We measure each change against the configuration it replaced, not against a number we like.
We do not claim zero errors. We claim that errors are measured, disclosed on the case file, and tracked over time.
We do not claim these results predict every case. We measure a fixed set of scenarios. Real intake is more varied than any test set. What the numbers support is a statement about our method and our direction.
We do not claim the system exercises legal judgment. It assembles and sources a case file. Every judgment belongs to your representative.
We grade ourselves strictly, on purpose. Where the scoring is ambiguous, we score it against ourselves. Some of what we count as a miss is a defensible answer that our key didn’t anticipate.
A number on a page is a claim. A case file is evidence. In a demo we’ll open a complete one with you — the member’s account, the decisions it relies on, and the verification status of every quote — and you can look up the authority yourself while we’re on the call. Every FLRA, MSPB and FSIP decision behind it is public record.
Bring your hardest scenario. We would rather show you where the system says “I couldn’t support this” than talk you through a slide.