trigger-discipline
The pipeline flags. You decide.
You are the enforcement reviewer at an AI platform's trust and safety desk. Each archetype in this queue is grounded in a published threat report, and each innocent look-alike — a pentester, a phishing-awareness trainer, a CTF student, a detection engineer, a journalist, an SRE, a novelist, a user who merely shares a VPN with an actor — was written to be indistinguishable from an actor on content.
Trigger discipline is the rule that keeps a finger off the trigger until the decision has actually been made. The trigger here is BAN, and this queue is built to make you reach for it early: most of the innocents were written to look exactly like the people you are looking for.
For each account you make one call: BAN, MONITOR, or CLEAR. Content is free to look at and misleading — what someone asks an AI tells you almost nothing; how they behave tells you almost everything. Behavioral and infrastructure evidence takes time to open, and from the second shift on that time runs on the shift clock.
The rules in full — policy, scoring, what each panel costs, keys
The policy
The policy here is the instrument's own, and it refuses two things. A ban that cites nothing beyond content is bounced: tick the evidence rows that carry your accusation, and at least one must come from a panel that is not content — behavior, infrastructure, or the scorer. A ban below the confidence floor is bounced the same way: every ban carries a confidence band, and the floor is “likely”. A refusal costs nothing. The desk hands the file back and waits.
Scoring
| Ban a threat-actor account | +10 |
| Clear a benign account | +5 |
| Monitor a threat-actor account | +2 |
| Monitor a benign account | −2 |
| Miss an actor (clear or leave undecided) | −10 |
| Ban a benign account | −25 |
Maximum here: . The asymmetry is the point: a missed actor returns, and a wrongly banned person mostly cannot argue their way back in. Monitor is the hedge — it keeps an actor in sight and puts an innocent under watch. The hedge is never the winning play, by design. Time is not part of the score.
Evidence, and how long it takes
| Content | instant | prompt excerpts, category, refused or completed — what a content filter sees |
| Account file | 1h | signup age, email kind, payment, phone verification |
| Behavior | 2h | session timeline, cadence, volume, refusal rate, category mix |
| Network | 2h | infrastructure per session, plus which other queue accounts share it |
| Pipeline read | 2h | the scorer's risk breakdown; cluster assessment and policy decision if one exists |
Each panel opens once per account and stays open. Every evidence row in an open panel can be cited; a ban stands on its citations.
Keyboard: B ban · M monitor · C clear · G flag a policy gap · 1–5 evidence panels · Enter continue · arrow keys move through the queue. While the band picker is open, number keys choose a band and Esc cancels.
Seed . Add ?seed=N to the URL to reorder the queue.