AI NewsModels & agentsAnnouncement
Musubi released PolicyLM-1.7B, an open-weights model that reads your policy
Musubi published PolicyLM-1.7B, an Apache-2.0 content moderation classifier that reads a written policy and scores a message against each category in a median of 35 milliseconds on one 24 GB GPU.

Image: Musubi
Why it mattersTeams running live chat, comments or user posts can now self-host a model that reads their own policy text and labels messages against it at classifier speed and cost.
A team running a live chat, a comment thread or any surface that posts user text has to label the posts somehow, and the usual options are a hard-wired classifier that cannot read a written policy or a general chat model that costs real money on every message.
Musubi released PolicyLM-1.7B, an Apache-2.0 model that reads a written policy at inference time and returns a score from 0 to 1 for every category in the policy. The model has 1.7 billion parameters, is published on Hugging Face as musubilabs/policylm-1.7b at revision v1.2, and ships with 23 built-in Aegis 2.0 safety categories for teams that do not want to write their own. TechCrunch covered the launch.
What it reads and what it answers
A policy for PolicyLM is up to 16 categories per call, each with a violation rule, a not-a-violation rule and an exception override, up to 1,662 tokens for the whole policy. The model scores every category in one pass and returns which rules a message breaks when it flags one. The scores are calibrated enough that Musubi ships two preset cutoffs: a precision preset at 0.335 for the custom policy mode (0.69 for the built-in taxonomy), and a balanced preset at 0.275 (0.45) for places where a missed violation costs more than a false flag. Musubi also reports that about one in two single-clause edits to a category changed a decision at the balanced cutoff on its own benchmark.
The model is fine-tuned from BidirLM-1.7B-Embedding on three training sets: NVIDIA's Nemotron-Safety-Guard-Dataset-v3, Alibaba's XGuard-Train-Open-200K and ToxicityPrompts' PolyGuardMix. It is evaluated on messages in 19 languages. On Musubi's custom-policy benchmark, which the company ran itself, PolicyLM-1.7B scores higher than every other model Musubi tested under 20 billion parameters. On its own measurement, median latency is 35 milliseconds per short chat message on one 24 GB NVIDIA L4, and 22 milliseconds on an H100. It also runs on a laptop CPU or Apple silicon.
What it is not
Musubi names what the model is not for: child-safety enforcement (route that to dedicated tooling), being the only self-harm safeguard, a security boundary against adversarial users, and moderating assistant responses. It also warns that when a decision needs a written reason, such as for appeals, bans or takedowns, a larger chat model is a better fit. The model works offline, needs no second download, and takes neither trust_remote_code nor image or conversation-history input.
Musubi co-founder Filip Jankovic told TechCrunch the goal is to let platform managers label content proactively instead of reacting after a flag. Jankovic traces the idea to a 2024 project called GLiNER rather than to TypeSafe's recent Jev decision model, and Musubi's own announcement says the two aim at related jobs: "If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself." The team using this on live chat can swap out a cloud moderation call for a self-hosted classifier that reads a policy document the team controls.
Source
- Introducing PolicyLM-1.7B, Musubi
- musubilabs/policylm-1.7b, Hugging Face
- How AI decision models could change content moderation, Russell Brandom, TechCrunch
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.