Content moderation for Microsoft Agent Framework
Arcjet content moderation detects harmful content in untrusted text before it
is stored, displayed, or forwarded. It is a Guard rule – call it from
guard() / Guard, not protect().
What is Arcjet?
Arcjet is the AI agent runtime security platform. Discover the agents running in your organization, enforce policy across every action, prompt, and tool call, and keep the evidence to prove what happened. Detect prompt injection, authorize agent tool calls, redact PII, and block bots and abuse.Quick start
Section titled “Quick start”Content moderation is a Guard rule. Wrap the tool with GuardTool and
GuardModerateContent.
For more information about adapter options, see the Microsoft Agent Framework agent guard.
package main
import ( "context" "encoding/json"
"github.com/arcjet/arcjet-go" "github.com/arcjet/arcjet-go/agentframework" "github.com/microsoft/agent-framework-go/tool" "github.com/microsoft/agent-framework-go/tool/functool")
type messageArgs struct { Text string `json:"text"`}
func postMessage(_ context.Context, in messageArgs) (map[string]bool, error) { return map[string]bool{"ok": true}, nil}
func guardedPost(guard *arcjet.GuardClient) (tool.FuncTool, error) { moderate := must(arcjet.GuardModerateContent(arcjet.GuardModerateContentOptions{ Mode: arcjet.ModeLive, })) fn := functool.MustNew( functool.Config{Name: "post_message", Description: "Post a user message"}, postMessage, ) return agentframework.GuardTool(guard, fn, agentframework.ToolPolicy{ Action: "message.posted", Rules: agentframework.Args(func(_ context.Context, in messageArgs) ([]arcjet.GuardRuleInput, error) { return []arcjet.GuardRuleInput{moderate.Text(in.Text)}, nil }), })}
func must[T any](v T, err error) T { if err != nil { panic(err) } return v}Language examples
Section titled “Language examples”npm install @arcjet/guardimport { launchArcjet, moderateContent } from "@arcjet/guard";
const arcjet = launchArcjet({ key: process.env.ARCJET_KEY! });const moderate = moderateContent();
const decision = await arcjet.guard({ label: "tools.chat", rules: [moderate(userMessage)],});
if (decision.conclusion === "DENY" && decision.reason === "MODERATE_CONTENT") { throw new Error("Harmful content detected – rephrase your message");}
const result = moderate.result(decision);// `detected` is true when harmful content was found. Billing is undefined// when the service does not report usage. Content moderation uses text_units.console.log(result?.detected, result?.billing?.unit, result?.billing?.count);pip install arcjetimport os
from arcjet.guard import ModerateContent, launch_arcjet
arcjet = launch_arcjet(key=os.environ["ARCJET_KEY"])moderate = ModerateContent()
decision = await arcjet.guard( label="llm.output", rules=[moderate(text)],)
if decision.conclusion == "DENY" and decision.reason == "MODERATE_CONTENT": raise RuntimeError("Harmful content detected – rephrase your message")
result = moderate.result(decision)# `detected` is True when harmful content was found. Billing is None# when the service does not report usage. Content moderation uses text_units.print(result.detected if result else None)if result and result.billing: print(result.billing.unit, result.billing.count)go get github.com/arcjet/arcjet-go@latestmoderation, err := arcjet.GuardModerateContent(arcjet.GuardModerateContentOptions{ Mode: arcjet.ModeLive, // required})if err != nil { return err}
decision, err := guard.Guard(ctx, arcjet.GuardRequest{ Label: "tools.generate", Rules: []arcjet.GuardRuleInput{moderation.Text(userMessage)},})if err != nil { return err}if decision.IsDenied() && decision.Reason == arcjet.ReasonModerateContent { return errors.New("content flagged by moderation")}
// Billing is optional. Content moderation usage is measured in text_units.if result := moderation.Result(decision); result != nil && result.Billing != nil { fmt.Printf("charged %d %s\n", result.Billing.Count, result.Billing.Unit)}Set Mode on every Guard rule. An empty Mode returns ErrInvalidMode.
Keep the response generic. Do not leak detector details or explain exactly what was flagged.
What next?
Section titled “What next?” Content moderation intro When to use content moderation and how the result is shaped.
Agent guards Protect tool calls and other actions without an HTTP request.
Prompt injection detection Block hostile instructions before they reach your model.
Sensitive information Detect PII locally before it leaves your application.