● How it is checked

How we test that Rolebird makes no unchecked claims

A promise not to invent experience is easy to make and hard to check, so it is worth little on its own. This is our method, the thresholds we hold ourselves to, and what it does not cover.

Why this matters more than it sounds

A cover letter that quietly adds a qualification you don’t have is worse than no cover letter. It is embarrassing at interview, and in New Zealand it can be grounds for dismissal later - an employer who finds out after hiring you has a straightforward case that you misled them.

So “the AI sometimes exaggerates” is not a rough edge to us. It is the failure that matters most, and it is the one we test hardest.

The gate

Rolebird’s output is shaped by a prompt. Any change to that prompt - a new instruction, a reworded rule, a model upgrade - can improve one thing and quietly break another. Prose has no unit tests, so a change that makes letters livelier can also make them looser with the truth, and nothing would catch it.

Before a prompt change ships, it is run against saved test cases and every output is graded 0–5 on nine criteria by a second, separate model call acting as judge. Two of those criteria are safety floors that cannot be traded away:

  • Nothing invented in the letter or CV changes. Every claim has to be supported by the CV or base letter the user actually supplied. Any invented fact caps the score at 2 out of 5.
  • Nothing invented in the ATS screening. Every requirement the screening reports as “matched” has to be genuinely evidenced. Telling someone they match a requirement they don’t is a fabrication with consequences, so it is held to the same standard.

Both must score 4 or higher on every single output, and the mean across all nine criteria must be at least 3.8. A fabrication cannot be averaged away against a nice turn of phrase - the safety floors are checked per output, not blended into the mean.

What the scores actually mean

Say a listing asks for experience managing a team, and your CV says you coordinated shifts of up to six staff.

  • 5 - the screening lists it as matched, and the shift coordination clearly backs that up.
  • 4 - it lists it as matched, and that is fair but thin. Coordinating shifts is not quite line management, so the wording stretches a little.
  • 2 or lower - it lists team management as matched when nothing in your CV supports it. That is a made-up match, and it caps the score automatically.

So any score above 3 means the judge found nothing made up. A 4 rather than a 5 is a match that is real but weakly evidenced - not a little bit of invention being let through.

A score is still only the judge’s opinion, so checks sit underneath it that are not opinions at all. If the judge quotes anything as unsupported, the run fails - however high the scores are. Every requirement the screening reports as matched has to carry a quote, copied word-for-word from the CV, and code verifies the quote is really there - a quote that isn’t is treated as a fabricated match, no judgment involved. And if the same requirement turns up in both the “matched” and “missing” lists, code catches it: the screening would be contradicting itself, and that is not something a model can talk its way around.

Checked on every application, not just in testing

The gate above runs when we change how Rolebird writes. Your application is not a test case, so every generation is also checked individually, before you see it:

  • Every “covered” requirement in the screening carries the words that prove it. A quote from your own CV or letter, verified by code to actually be there. A match that cannot be quoted is moved to “missing” - the screening is allowed to understate your match, and never to overstate it.
  • Every suggested CV change must quote the text in your documents that backs it. A suggestion that cannot be grounded is discarded. You may get fewer suggestions; you will not get invented ones.
  • A separate verification pass reads the finished application against your documents - the letter, the changes, the screening, the strategy notes - with its own instructions and no stake in the writing. Claims your documents do not support are removed. A letter that cannot be repaired honestly is not delivered at all: the generation fails and you can retry, because a failed generation costs us more than it costs you, and a false claim at interview costs you more than both.

The verification pass is itself a model, reading with different instructions than the one that wrote. It is not infallible - where the honest boundary of a phrase sits, two careful readers can disagree. What cannot happen silently is a claim with nothing in your documents behind it: it is removed, or the generation fails.

The other seven

The safety floors stop it lying. These are what stop it being useless - a letter that invents nothing but reads like every other AI letter has failed differently.

  • Voice match - does it sound like the writer’s own base letter, in formality, warmth and sentence length, rather than a template?
  • Opening - specific and true, not “I am writing to apply” or “excited to apply”.
  • Evidence - concrete results and decisions, with “team player” and “proven track record” cut.
  • CV reasoning - each suggested change explains why it helps for this role, rather than restating the job ad back at you.
  • Prioritisation - the few high-impact changes are marked as such, and the list is not padded with trivia all marked the same.
  • Strategy - candid positioning advice, including the real gap and how to handle it, rather than encouragement.
  • NZ English and format - local spelling and tone, one page, and placeholders where personal details belong.

The full rubric, with the exact wording each criterion is graded against, lives in the codebase alongside the gate that enforces it. For how the ATS screening in particular is tested, and why part of that check is code rather than a model, see how we test that the ATS screening tells you the truth.

What this does not prove

Being straight about the limits is the only thing that makes the rest worth reading.

  • The gate is a sample; the per-application checks are not a guarantee either. The gate runs against a small set of saved cases and catches a prompt change that makes fabrication more likely. The checks above run on every individual generation - but the verifier is a model too, and a borderline phrase (“a track record of”, over one strong result) can read as fair to one careful reader and as a stretch to another. What we can promise is narrower and checkable: nothing reaches you unchecked, and no claim survives with nothing in your documents behind it.
  • The judge is a model, not a person. It is a different model from the one that did the writing - the gate refuses to run if the two are ever configured the same, because a model grading its own work shares its own blind spots. It sees the source material and the output, but not the writing prompt, and scores against an explicit rubric. It can still simply be wrong.
  • It does not check facts about the world. It checks that the output is supported by what you gave us. If your own CV overstates something, Rolebird will faithfully carry it through.

Which is the real reason Rolebird shows you a list of suggested CV changes rather than silently rewriting your CV, and why nothing is ever submitted on your behalf. The last check is you, and it should be.

And the same for what we publish

The gate above is about what the model writes for you. Our research on the NZ job market is held to a different standard, for a similar reason: advice built on a number nobody checked is worth as little as a letter built on experience you don’t have.

It works by triangulation - reading Stats NZ, RBNZ, MBIE, Seek, Trade Me and JobAdder against each other rather than repeating whichever figure was published most recently. Where sources disagree, the report says so instead of picking the convenient one. Of 102 falsifiable claims extracted from 22 sources, the important ones were cross-checked against multiple independent sources: 21 confirmed, 3 sent back for manual re-checking, and 1 refuted and removed.

That last number is the one worth noticing. Reporting what was thrown out is the part nobody does, and it is the only way to know the rest was actually checked. It is a synthesis of published sources, not primary research - we ran no survey and hold no dataset of our own behind those figures.

Why any of this matters to your job hunt

Most of what makes a job search miserable is not knowing whether your experience is normal. Forty applications and no replies feels like a verdict on you. Knowing that applications per advertised position are up 261% since 2022 - averaging 43 candidates a job, and 150–300 on contested ones - and that 41% of employers say they still cannot find the right person, tells you something very different: the market is noisy, and the answer is not to send more.

That is what the research is for - not motivation, but calibration. Knowing what is actually happening lets you decide where to spend your effort, and it is the same reasoning that shaped the product: fewer applications, each one worth reading.

See it yourself

Every generated application comes with an ATS screening showing which requirements are matched and which are missing - including the ones you don’t meet. A tool trying to flatter you would leave those out. The same standard applies to your data: how your CV is de-identified before it ever reaches us. Read more about what we do with your documents, or the guides.