Getting the most out of Nima
Best Practices for T&S Leads, Policy Ops, Moderators and Engineers
How to use this guide
Nima delivers the most value when four roles adopt it deliberately: T&S Leads who own enforcement outcomes, Policy Ops who turn policy into rules, moderators who apply policy to the cases automation can't settle, and engineers who connect the platform to Nima. This guide gives the practices that matter most for each role, and the rituals that connect them.
Persona | Typical titles | Owns | Lives in Nima |
|---|---|---|---|
T&S Lead | Head of Trust & Safety, T&S Director, Community Safety Lead | Policy, risk posture, team, regulatory exposure | Dashboards, queue health, strike design, reporting |
Policy Ops | Policy Operations Specialist, T&S Analyst, Moderation Ops Manager | Rules, thresholds, integrations, quality | Rule engine, AI providers, Rule Testing, custom attributes |
Moderator | Content Moderator, Trust & Safety Agent, Review Specialist | Case decisions, escalations, decision notes | Review queues, item context, actions and escalation |
Engineer | Backend Engineer, Integration Engineer, Platform Engineer | API integration, webhooks, metadata flow | API endpoint, webhook configuration, data fetch |
In smaller teams one person plays several roles. Read the tracks that apply and use the Working together section as your own checklist.
T&S Lead: role and goals
The T&S Lead succeeds with Nima when enforcement is consistent, proportionate and explainable to users, executives and regulators. set the policy frame, the risk appetite and the operating cadence that Policy Ops works within.
What success looks like:
- Every enforcement action traces to a named policy in Nima
- Queues reflect reviewer skills and decision types, with SLAs the team actually meets
- Graduated consequences run through strike engines rather than ad hoc bans
- Transparency and statement-of-reasons data comes out of Nima without manual rework
T&S Lead
Phase | Focus | Outcomes to reach |
|---|---|---|
Phase 1 | Foundations | Policy taxonomy agreed and loaded into Nima; severity level per policy; initial queue map signed off; named Policy Ops owner |
Phase 2 | Proportionality | Strike engine ladders defined (warning, restriction, suspension, ban); SLAs set per queue; appeal path live |
Phase 3 | Governance | First governance review: KPI dashboard agreed with leadership; |
T&S Lead: best practices
Start from policy, not tooling. Map Nima policies one-to-one to your community guidelines before anyone builds a rule.
Design queues around decisions, not volume. Give each queue one decision type or reviewer skill: CSAM escalation, hate speech in a given language, appeals. Tight scope keeps reviewer context sharp and makes misrouting visible.
Make enforcement proportionate. Use strike engines to encode graduated consequences. This cuts repeat-offender workload and gives users and regulators a consistent, explainable enforcement story.
Build compliance in from day one. Consistent policies and actions make statement-of-reasons, action logs and transparency reporting a by-product rather than a quarterly scramble.
Run governance on a cadence. Hold a monthly review with Policy Ops, not just post-incident reviews.
T&S Lead: KPIs and questions to ask
KPI | Why it matters | Healthy signal |
|---|---|---|
Automated action overturn rate | Proxy for over-enforcement | Falling, and explained per policy |
Appeal rate and appeal success rate | User-facing fairness | Stable or falling success rate |
Queue SLA attainment | Reviewer capacity and routing quality | Met on priority queues |
Repeat-offender rate | Whether strikes change behavior | Falling after strike ladder launch |
Policy Ops: role and goals
Policy Ops succeeds with Nima when every rule is traceable, tested and tuned from evidence rather than instinct. You translate the T&S Lead's policies and risk appetite into rules, signals, thresholds and routing.
What success looks like:
- Key platform signals flow into Nima as custom attributes
- AI provider has a clear job, and you know which score drives which rule
- A labeled test set can be used and every rule change runs through batch testing before activation
Policy Ops
Phase | Focus | Outcomes to reach |
|---|---|---|
Phase 1 | Signals and routing | Custom attributes defined (account age, verification, surface); AI providers mapped to content types; rules route into the agreed queues |
Phase 2 | Evidence | Baseline batch test run on every rule in dry run mode; first threshold adjustments shipped with before/after results |
Phase 3 | Feedback loop | Weekly overturn review running; overturned cases added to the golden set; first rule hygiene pass complete |
Policy Ops: best practices
Invest early in custom attributes. Provide as much context to the system as you can
Treat thresholds as hypotheses. A threshold set by gut feel drifts toward over-enforcement or under-enforcement. Every threshold should have a recorded rationale and a test result behind it.
Gate changes with batch testing. Run every meaningful rule or threshold change against the golden set before activation. Read the results by type: under-enforcement (false negatives) and over-enforcement (false positives) first, then over-routing and under-routing. Use ad-hoc testing to sanity-check single edge cases, not to approve changes.
Close the loop with reviewer decisions. Overturned automated actions are your richest tuning signal. Review overturn rates per rule and per model weekly, and add each overturned case to the golden set.
Keep the rule set clean. Once a month, retire rules that haven't fired, merge overlapping ones, and re-validate thresholds against fresh traffic.
Moderator: role and goals
Moderators succeed with Nima when their decisions are accurate, consistent and well documented, because every decision they make is both an enforcement action and a training signal. You apply policy to the cases automation can't settle, and your overturns tell Policy Ops where rules are wrong.
What success looks like:
- Decisions cite the correct policy and action every time
- QA agreement with calibrated reference decisions is high and stable
- Unclear cases are escalated or flagged, not guessed
- Queue SLAs are met without sacrificing accuracy
Moderator
Phase | Focus | Outcomes to reach |
|---|---|---|
Phase 1 | Policy and tooling | Policy taxonomy and severity levels learned; review screen, actions and escalation paths walked through with a lead |
Phase 2 | Shadowed review | Decisions made on a calibration set and compared with reference answers; disagreements discussed |
Phase 3 | Live queue, low severity | Working one queue with QA sampling on every shift; notes on overturns becoming routine |
Phase 4 | Full queue assignment | Assigned queues at target SLA; first calibration session joined with peers |
Moderator: best practices
Read the context before the content. Check the custom attributes and history Nima shows alongside the item, such as account age.
Decide on policy, not instinct. Pick the policy that the content actually violates and the action the strike ladder prescribes. When personal judgment and policy disagree, follow policy and flag the gap.
Tag precisely. The policy you select feed statement-of-reasons notices, transparency reports and strike counts.
Escalate rather than guess. Use the escalation queue for severe harms, legal edge cases and anything the policy doesn't clearly cover. A fast escalation beats a confident wrong call.
Report what the rules miss. If you see a pattern the automation isn't catching, such as a new slang term or evasion technique, raise it. Moderators are usually the first to spot under-enforcement.
Moderator: quality signals and wellbeing
Signal | What it tells you | Healthy signal |
|---|---|---|
QA agreement rate | Accuracy against calibrated decisions | High and stable across queues |
Decision time per item | Pace versus care | Stable, not falling at the expense of accuracy |
Escalation rate | Whether edge cases reach the right people | Steady; spikes point to unclear policy |
Overturns with a note | Quality of feedback to Policy Ops | Every overturn annotated |
Wellbeing is part of accuracy. Rotate moderators between high-exposure and low-exposure queues, cap continuous time on severe-harm queues, and use any blurring or grayscale display options your setup provides.
Engineer: role and goals
The Engineer owns what flows into Nima and how decisions flow back to the platform. Nima is designed to keep that work small, so the job is less about building an integration and more about configuring one well.
What success looks like:
- Content and metadata from every surface reach Nima through the single endpoint
- Enforcement decisions arrive in your platform's own format and are applied without a translation layer
- The metadata Policy Ops relies on is present on the first request and refreshed through data fetch where it goes stale
Engineer
Phase | Focus | Outcomes to reach |
|---|---|---|
Phase 1 | Scoping | Content types, surfaces and metadata fields agreed with Policy Ops; webhook body format defined |
Phase 2 | Integration | Events flowing through the single endpoint in staging; metadata ingested; webhook decisions applied on the platform |
Phase 3 | Production | Data fetch set up for metadata that goes stale; monitoring live; production cutover, starting with one surface |
Engineer: best practices
Integrate once, through a single endpoint. All content types, surfaces and events go through one API endpoint, so there's one contract, one auth setup and one place to monitor. Adding a new surface or content type means sending a new payload shape, not building a new integration.
Shape webhooks to fit your platform. Customise the webhook body so enforcement decisions arrive in the format your services already expect: your field names, IDs and action codes. That removes the translation layer on your side, and decisions can go straight to the service that applies them.
Send metadata up front, refresh it when needed. Include the metadata workflows rely on, such as account age and verification status, in the first API request, and Nima ingests it along with the content. When that context may have gone stale by the time a rule runs or a moderator reviews the case, data fetch refreshes it from your systems.
Working together
The roles meet at five handoffs: platform data becomes signals, policy becomes rules, rules become reviewed decisions, evidence becomes policy changes, and incidents become test cases.
Ritual | Cadence | Led by | Output |
|---|---|---|---|
Overturn and appeals review | Weekly | Policy Ops | Golden set updates, tuning backlog |
Governance review | Monthly | T&S Lead | Policy changes, risk appetite updates, retired rules |
Incident retro | After each notable incident | T&S Lead | New test cases, rule or queue changes |
Quarterly configuration audit | Quarterly | Both | Rule hygiene, provider fit, SLA and strike ladder review |
Moderator calibration session | Weekly | Policy Ops with moderators | Aligned decisions on hard cases, policy clarifications, new test cases |
Signal review | Monthly | Engineer with Policy Ops | New metadata fields, webhook body changes, data fetch needs |