Skip to content

Rate limiting reference

Arcjet rate limiting lets you define rules which limit the number of requests a client can make over a period of time.

Each rate limit is configured on an exact path with a set of client characteristics and algorithm specific options.

Tracks the number of requests made by a client over a fixed time window. Options are explained in the Configuration documentation. For more details about how the algorithm works, see the fixed window algorithm description.

// Options for fixed window rate limit
// See https://docs.arcjet.com/rate-limiting/configuration
type FixedWindowRateLimitOptions = {
// "LIVE" will block requests. "DRY_RUN" will log only
mode?: "LIVE" | "DRY_RUN";
// How the client is identified. See https://docs.arcjet.com/fingerprints
characteristics?: string[];
// Time window the rate limit applies to (for example, "1h", "60s", or seconds as a number)
window: string | number;
// Maximum number of requests allowed in the time window
max: number;
};

Tracks the number of requests made by a client over a sliding window so that the window moves with time. Options are explained in the Configuration documentation. For more details about how the algorithm works, see the sliding window algorithm description.

// Options for sliding window rate limit
// See https://docs.arcjet.com/rate-limiting/configuration
type SlidingWindowRateLimitOptions = {
// "LIVE" will block requests. "DRY_RUN" will log only
mode?: "LIVE" | "DRY_RUN";
// How the client is identified. See https://docs.arcjet.com/fingerprints
characteristics?: string[];
// The time interval in seconds for the rate limit
interval: number;
// Maximum number of requests allowed over the time interval
max: number;
};

Based on a bucket filled with a specific number of tokens. Each request withdraws a token from the bucket and the bucket is refilled at a fixed rate. Once the bucket is empty, the client is blocked until the bucket refills. Options are explained in the Configuration documentation. For more details about how the algorithm works, see the token bucket algorithm description.

// Options for token bucket rate limit
// See https://docs.arcjet.com/rate-limiting/configuration
type TokenBucketRateLimitOptions = {
// "LIVE" will block requests. "DRY_RUN" will log only
mode?: "LIVE" | "DRY_RUN";
// How the client is identified. See https://docs.arcjet.com/fingerprints
characteristics?: string[];
// Number of tokens to add to the bucket at each interval
refillRate: number;
// The interval in seconds to add tokens to the bucket
interval: number;
// The maximum number of tokens the bucket can hold
capacity: number;
};

When using a token bucket rate limit, each request must specify the number of tokens it wishes to withdraw from the bucket. This is done by passing a requested property to the protect function.

For how to specify the number of tokens to request, see the token bucket request example.

The amount of tokens to deduct from the bucket is specified in the requested option as a positive integer when calling the Arcjet protect function.

// Deduct 50 tokens from the bucket
// The value for `requested` must be a positive integer
const decision = await aj.protect(req, { requested: 50 });

Rate limit rules use characteristics to identify the client and apply the limit across requests. The default is to use the client’s IP address. However, you can specify other characteristics such as a user ID or other metadata from your application.

In this example we define a rate limit rule that applies to a specific user ID. The custom characteristic is userId with the value passed as a prop on the protect function. You can use any string for the characteristic name and any string, number or boolean for the value.

To identify users with different characteristics, such as IP address for anonymous users and a user ID for logged in users, you can use withRule (JS) or with_rule() (Python) to create augmented clients that use different characteristics. See the example in the custom characteristics section.

Per route versus middleware

Rate limit rules can be configured in two ways:

  • Route handlers: The rule is defined in the route handler (previously known as an API route) itself. This lets you configure the rule alongside the code it is protecting which is useful if you want to use the decision to add context to your own code. However, it means rules are not automatically applied to every request.
  • Middleware/Proxy: The rule is defined in the middleware (renamed to proxy in Next.js 16). This lets you configure rules in a single place or apply them globally to all routes, but it means the rules are not located alongside the code they are protecting and can miss route specific context .

If you use a platform that performs health checks or liveness probes, ensure that Arcjet is not enabled for those routes. These requests don’t have all of the metadata that Arcjet requires to make security decisions.

Per route

If you define your rate limit within an API route Arcjet assumes that the limit applies only to that route. If you define your rate limit in middleware, either use the Next.js matcher config to choose which paths to execute the middleware for, or use request.nextUrl.pathname.startsWith.

Rate limit only on /api/*

You can use conditionals in your Next.js middleware to match multiple paths.

Middleware

Avoid double protection

If you use Arcjet in middleware/proxy and individual routes, you need to be careful that Arcjet is not running multiple times per request. This can be avoided by excluding the API route from the middleware matcher.

For example, if you already have a rate limit defined in the API route at /api/hello, you can exclude it from the middleware by specifying a matcher in /proxy.ts:

Pages and server actions

Arcjet can be used inside Next.js middleware, API routes, pages, server components, and server actions. Client components cannot be protected because they run on the client only.

See the Next.js SDK reference for examples of pages and page components and server actions.

Arcjet provides a single protect function that is used to execute your protection rules. This requires a RequestEvent property which is the event context as passed to the request handler.

This function returns a Promise that resolves to an ArcjetDecision object. This contains the following properties:

  • id (string) – The unique ID for the request. This can be used to look up the request in the Arcjet dashboard. It is prefixed with req_ for decisions involving the Arcjet cloud API. For decisions taken locally, the prefix is lreq_.
  • conclusion (ArcjetConclusion) – The final conclusion based on evaluating each of the configured rules. If you wish to accept Arcjet’s recommended action based on the configured rules then you can use this property.
  • reason (ArcjetReason) – An object containing more detailed information about the conclusion.
  • results (ArcjetRuleResult[]) – An array of ArcjetRuleResult objects containing the results of each rule that was executed.
  • ip (ArcjetIpDetails) – An object containing Arcjet’s analysis of the client IP address. For more information, see the SDK reference.

To check whether a rate limit rule returned a deny conclusion, use decision.isDenied() and decision.reason.isRateLimit() (JS) / decision.is_denied() and decision.reason_v2.type == "RATE_LIMIT" (Python).

You can iterate through the results and check whether a rate limit was applied:

for (const result of decision.results) {
console.log("Rule Result", result);
}

This example logs the full result as well as each rate limit rule:

When using a token bucket rule, pass an additional requested prop (a positive integer) to the protect function. This is the number of tokens the client is requesting to withdraw from the bucket.

With a rate limit rule enabled, you can access additional metadata in every Arcjet decision result:

  • max (number): The configured maximum number of requests applied to this request.
  • remaining (number): The number of requests remaining before max is reached within the window.
  • window (number): The total amount of seconds in which requests are counted.
  • reset (number): The remaining amount of seconds in the window.

These can be used to return RateLimit HTTP headers (draft RFC) to offer the client more detail.

In JavaScript, use setRateLimitHeaders from @arcjet/decorate. In Python, use set_rate_limit_headers from the arcjet package. Both write RateLimit and RateLimit-Policy headers from the decision.

When several rate limit results are present, the tightest remaining budget is advertised. If two policies share the same max, no headers are written.

This would result in draft RFC response headers similar to the following:

...
< RateLimit: limit=10, remaining=5, reset=9
< RateLimit-Policy: 10;w=10
...

Arcjet is designed to fail open so that a service issue or misconfiguration does not block all requests. The SDK also times out and fails open after 2000 ms by default. However, in most cases, the response time is less than 20 ms to 30 ms.

If there is an error condition when processing the rule, Arcjet returns an ERROR result for that rule and you can check the message property on the rule’s error result for more information.

If all other rules that were run returned an ALLOW result, then the final Arcjet conclusion is ERROR.

Arcjet runs the same in any environment, including locally and in CI. You can use the mode set to DRY_RUN to log the results of rule execution without blocking any requests.

We have an example test framework you can use to automatically test your rules. Arcjet can also be triggered based using a sample of your traffic.

For details, see the Testing section of the docs.

Examples

Rate limit by IP address

The following example shows how to configure a rate limit on a single API route. It applies a limit of 60 requests per hour per IP address. If the limit is exceeded, the client is blocked for 10 minutes before being able to make any further requests.

Applying a rate limit by IP address is the default if no characteristics are specified.

Rate limit by IP address with custom response

The following example is the same as the preceding one. However this example also shows a customized response rather than the default.

Rate limit by AI tokens

If you are building an AI application you may be more interested in the number of AI tokens rather than the number of HTTP requests. Popular AI APIs such as OpenAI are billed based on the number of tokens consumed and the number of tokens is variable depending on the request, such as conversation length or image size.

The token bucket algorithm is a good fit for this use case because you can vary the number of tokens withdrawn from the bucket with every request.

The following example configures a token bucket rate limit using the openai-chat-tokens library to track the number of tokens used by a gpt-3.5-turbo AI chatbot. It sets a limit of 2,000 tokens per hour with a maximum of 5,000 tokens in the bucket. This allows for a reasonable conversation length without consuming too many tokens.

See the arcjet-js GitHub repo for a full example using Next.js.

Rate limit by API key header

APIs are commonly protected by keys. You may wish to apply a rate limit based on the key, regardless of which IPs the requests come from. To achieve this, you can specify the characteristics Arcjet uses to track the limit.

The following example shows how to configure a rate limit on a single API route. It applies a limit of 60 requests per hour per API key, where the key is provided in a custom header called x-api-key. If the limit is exceeded, the client is blocked for 10 minutes before being able to make any further requests.

If you specify different characteristics and do not include ip.src, you may inadvertently rate limit everyone. Be sure to include a characteristic which can narrowly identify each client, such as an API key as shown here.

Global rate limit

Using Next.js middleware lets you set a rate limit that applies to every route:

Middleware runs on every route so be careful to avoid double protection if you are configuring Arcjet directly on other routes.

Response based on the path

You can also use the req NextRequest object to customize the response based on the path. In this example, we’ll return a JSON response for API requests, and a HTML response for other requests.

Rewrite or redirect

The NextResponse object returned to the client can also be used to rewrite or redirect the request. For example, you might want to return a JSON response for API route requests, but redirect all page route requests to an error page.

Wrap existing handler

All the examples on this page show how you can inspect the decision to control what to do next. However, if you just wish to send a generic 429 Too Many Requests response you can delegate this to Arcjet by wrapping your handler withArcjet.

Discussion