Skip to content

AI budget control for Go

AI providers bill per token. Without per-user limits, a single user can exhaust your entire monthly budget – through prompt attacks, runaway loops, or just heavy legitimate use.

A token bucket rate limit maps directly onto how AI billing works. You estimate the cost of each request in tokens, deduct it from the user’s bucket, and deny requests when the bucket is empty. The bucket refills over time, giving each user a sustained allowance without sharp rate-limit cliffs.

Alternatively, you can use a fixed window or sliding window limit to enforce a hard cap on spend per user per day, week, or month. For details on different approaches, see the rate limiting algorithms reference.

Arcjet handles bucket state across all instances of your application – no Redis or external state store required.

Token bucket rate limiting tracks an AI token budget per user. Interval is a time.Duration. Pass the tokens this request consumes with WithRequested.

main.go
package main
import (
"encoding/json"
"log"
"net/http"
"os"
"time"
"github.com/arcjet/arcjet-go"
)
var aj = must(arcjet.NewClient(arcjet.Config{
Key: os.Getenv("ARCJET_KEY"),
Rules: []arcjet.Rule{
arcjet.TokenBucket(arcjet.TokenBucketOptions{
Mode: arcjet.ModeLive,
Characteristics: []string{"userId"},
RefillRate: 2000,
Interval: time.Hour,
Capacity: 5000,
}),
},
}))
type chatRequest struct {
Message string `json:"message"`
}
func chat(w http.ResponseWriter, r *http.Request) {
var body chatRequest
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
http.Error(w, "bad request", http.StatusBadRequest)
return
}
userID := "user-123"
estimate := (len(body.Message) + 3) / 4
if estimate < 1 {
estimate = 1
}
decision, err := aj.Protect(
r.Context(),
r,
arcjet.WithRequested(estimate),
arcjet.WithCharacteristics(map[string]string{"userId": userID}),
)
if err != nil {
log.Printf("arcjet: %v", err)
} else if decision.IsDenied() {
http.Error(w, "AI usage limit exceeded", http.StatusTooManyRequests)
return
}
_ = json.NewEncoder(w).Encode(map[string]string{"reply": "..."})
}
func main() {
http.HandleFunc("/chat", chat)
log.Fatal(http.ListenAndServe(":8000", nil))
}
func must[T any](v T, err error) T {
if err != nil {
log.Fatal(err)
}
return v
}

characteristics: ["userId"] - Tracks the bucket per user. Replace "userId" with the characteristic that identifies a unique user in your application, such as a session token, API key, or authenticated user ID. Pass the value to aj.protect() as a named argument.

refillRate and interval - Set the sustained allowance. refillRate: 2_000, interval: "1h" gives each user 2,000 tokens per hour. Adjust to match your AI provider’s pricing and your cost targets. These are hard coded in this example, but you can also calculate them dynamically based on user subscription level or other factors. Pass the calculated values to the rule.

capacity - The maximum tokens a user can accumulate. Setting capacity: 5_000 with refillRate: 2_000 lets users burst up to 5,000 tokens if they haven’t used their allowance recently.

The example uses a characters / 4 heuristic (~1 token per 4 characters for common English text). This is a reasonable starting point – it avoids introducing extra dependencies and works well enough for budget enforcement where a small margin of error is acceptable.

For accurate counts, use a tokenizer: