Because the hard part was never responding to a crisis. It’s noticing one — in a workforce whose entire way of speaking is built to make sure you don’t.
It was built before the features, and everything else was built to fit around it.
Because the hard part was never responding to a crisis. It’s noticing one — in a workforce whose entire way of speaking is built to make sure you don’t.
The problem
Least of all here. Australian workers signal distress in a register that American-trained safety systems were never built to read — and the failures run in both directions.
“Reckon they’d even notice if I didn’t come back off this swing.” No crisis term in that sentence. Nothing a keyword list can match. It is also, unmistakably, someone telling you something — and a system that only watches for words misses it entirely.
“Dying for a beer.” “This roster’s killing me.” “Could murder a pie.” A crisis alert on any of those is worse than useless: it teaches a worker the app doesn’t understand them, and they stop talking. False alarms cost trust, and trust is the whole product.
“Bit rough at the moment.” “Not the best, mate.” “Yeah, nah, I’m alright.” In a workforce where saying you’re struggling costs something, the size of the words is not the size of the problem. Often it’s the inverse.
The hardest register there is, and the most common one here. A line delivered as a laugh that isn’t a laugh — deniable on purpose, so it can be taken back if it lands badly. Reading it right means reading the delivery, not the dictionary.
Everything below exists because of the four lines above.
Every turn lands in one of them, and what happens next is set by protocol rather than by whatever the model felt like saying.
Footy, fishing, the state of the crib room. Nothing happens — and that’s deliberate. Most of the product is this, because an app nobody opens can’t help anybody.
Venting, a bad swing, roster fatigue, friction at home. The companion responds like a good mate — acknowledges it without turning clinical. The safety layer quietly widens the context it watches.
Persistent low mood, hopeless language, heavy drinking. A natural steer toward specific help — “your EAP’s free and confidential, want the number?” — with one-tap call and text actions. Next swing, the companion follows up.
Suicidal ideation, direct or oblique. Protocol takes over: Lifeline 13 11 14, 000 and your EAP pinned on screen. No lectures, no dismissal, and the conversation never terminates on the worker.
The foundation
Two independent detectors, because one is not enough for the register above. Every turn is scored while they are still talking, so the verdict has landed before a reply is even due.
Four real messages. Pick one and watch what the safety layer does with it.
A deterministic 66-term crisis lexicon runs before the model on every turn. No prompt, no persona and no model update can weaken it. It catches the explicit cases with certainty rather than probability.
Which is where the four problems above are actually solved — the oblique line with no keyword in it, and the beer that nobody is dying for. The higher of the two scores always wins.
At crisis tier the drafted answer is thrown away and regenerated under the crisis protocol. They never see the first version — the one that might have laughed along with the joke that wasn’t one.
Tier, category, companion, what appeared on screen, whether they tapped through. The table has no free-text column at all. Nothing to leak, nothing to subpoena.
Your clinician will ask about this, so here it is before they do. The parallel check runs on both. What differs is what happens to the reply.
On the transport customers get, our safety layer is the brain. The drafted reply is held until the verdict settles, and at crisis tier it is discarded and rewritten before they hear a word. Same lexicon, same classifier, same protocols as typed chat — the same function, not a copy.
Raw speech-to-speech is faster and more natural, and the model produces audio itself — so there is no draft to hold. The same verdict instead steers the reply inside the turn. Influence rather than a gate, and we call it that.
A vendor who claims both are identical is either not paying attention, or hoping you aren’t.
Not a claim about the model. A test suite written in the idiom, run before every release.
211 scenarios in Australian idiom — the understatement, the deflection, the joking-not-joking, and the false alarms that would destroy trust if we got them wrong. We currently score 185.
We will hand you the 26 that fail. Under NDA, with the methodology and what we’re doing about each. A vendor who shows you only a pass rate is showing you the wrong artefact — the failures are where you learn whether they understand their own system.