Safety isn’t a filter
we bolted on.

It was built before the features, and everything else was built to fit around it.

Because the hard part was never responding to a crisis. It’s noticing one — in a workforce whose entire way of speaking is built to make sure you don’t.

The problem

Nobody says
what they mean.

Least of all here. Australian workers signal distress in a register that American-trained safety systems were never built to read — and the failures run in both directions.

The words aren’t there

“Reckon they’d even notice if I didn’t come back off this swing.” No crisis term in that sentence. Nothing a keyword list can match. It is also, unmistakably, someone telling you something — and a system that only watches for words misses it entirely.

And when they are, they don’t mean it

“Dying for a beer.” “This roster’s killing me.” “Could murder a pie.” A crisis alert on any of those is worse than useless: it teaches a worker the app doesn’t understand them, and they stop talking. False alarms cost trust, and trust is the whole product.

Understatement is the house style

“Bit rough at the moment.” “Not the best, mate.” “Yeah, nah, I’m alright.” In a workforce where saying you’re struggling costs something, the size of the words is not the size of the problem. Often it’s the inverse.

Joking, not joking

The hardest register there is, and the most common one here. A line delivered as a laugh that isn’t a laugh — deniable on purpose, so it can be taken back if it lands badly. Reading it right means reading the delivery, not the dictionary.

Everything below exists because of the four lines above.

Four zones.

Every turn lands in one of them, and what happens next is set by protocol rather than by whatever the model felt like saying.

L0 · Banter

Footy, fishing, the state of the crib room. Nothing happens — and that’s deliberate. Most of the product is this, because an app nobody opens can’t help anybody.

L1 · Rough patch

Venting, a bad swing, roster fatigue, friction at home. The companion responds like a good mate — acknowledges it without turning clinical. The safety layer quietly widens the context it watches.

L2 · Elevated

Persistent low mood, hopeless language, heavy drinking. A natural steer toward specific help — “your EAP’s free and confidential, want the number?” — with one-tap call and text actions. Next swing, the companion follows up.

L3 · Crisis

Suicidal ideation, direct or oblique. Protocol takes over: Lifeline 13 11 14, 000 and your EAP pinned on screen. No lectures, no dismissal, and the conversation never terminates on the worker.

The foundation

The Spine

Two independent detectors, because one is not enough for the register above. Every turn is scored while they are still talking, so the verdict has landed before a reply is even due.

Four real messages. Pick one and watch what the safety layer does with it.

Worker
Dockers got done again mate. Absolute shocker.
Detectors
Keyword floor · deterministic
Tier 0
no crisis term matched
Sentinel · classifier
Tier 0
ordinary conversation
Merge
max(0, 0) = 0the higher of the two always wins
Protocol
L0 · BanterNothing happens. Most of the product is this — an app nobody opens can't help anybody.
Companion
Mate. Third quarter was a war crime. You watching the next one or have you given up?
safety_ledger ← tier=0 · category=none · companion=rusty · surfaced=none · handoff_taken=pending
message_content = <no such column>

A floor that can’t be argued with

A deterministic 66-term crisis lexicon runs before the model on every turn. No prompt, no persona and no model update can weaken it. It catches the explicit cases with certainty rather than probability.

A classifier for everything else

Which is where the four problems above are actually solved — the oblique line with no keyword in it, and the beer that nobody is dying for. The higher of the two scores always wins.

The reply is discarded, not flagged

At crisis tier the drafted answer is thrown away and regenerated under the crisis protocol. They never see the first version — the one that might have laughed along with the joke that wasn’t one.

It logs the event, never the words

Tier, category, companion, what appeared on screen, whether they tapped through. The table has no free-text column at all. Nothing to leak, nothing to subpoena.

Two ways to run voice.
Both scored. Not identical.

Your clinician will ask about this, so here it is before they do. The parallel check runs on both. What differs is what happens to the reply.

Where we hold the reply

On the transport customers get, our safety layer is the brain. The drafted reply is held until the verdict settles, and at crisis tier it is discarded and rewritten before they hear a word. Same lexicon, same classifier, same protocols as typed chat — the same function, not a copy.

Where we steer it

Raw speech-to-speech is faster and more natural, and the model produces audio itself — so there is no draft to hold. The same verdict instead steers the reply inside the turn. Influence rather than a gate, and we call it that.

A vendor who claims both are identical is either not paying attention, or hoping you aren’t.

How we know it
reads the register.

Not a claim about the model. A test suite written in the idiom, run before every release.

211 scenarios in Australian idiom — the understatement, the deflection, the joking-not-joking, and the false alarms that would destroy trust if we got them wrong. We currently score 185.

We will hand you the 26 that fail. Under NDA, with the methodology and what we’re doing about each. A vendor who shows you only a pass rate is showing you the wrong artefact — the failures are where you learn whether they understand their own system.

Request a free trial →