Skip to content

Rate Limiting in Spring Boot with Bucket4j and Redis

Published: at 12:00 PM

Rate Limiting in Spring Boot with Bucket4j and Redis

Table of Contents

Open Table of Contents

Brief

My goal is to make posts like this the SIMPLEST place on the internet to learn how to do things that caused me trouble. “Spring Boot rate limiting” is one of those searches where the top results are either a HashMap of counters that resets whenever the pod restarts, or a fifteen-dependency API-gateway tutorial. The middle path — correct, distributed rate limiting inside your app with Bucket4j and Redis — is genuinely simple, and this post builds it on Spring Boot 4 with a k6 load test to prove the limits actually hold.

Demo repo, as always, at github.com/StevenPG/DemosAndArticleContent.

Token Buckets in 90 Seconds

Bucket4j implements the token bucket algorithm: a bucket holds up to capacity tokens and refills at a fixed rate; each request consumes one token; empty bucket means rejected request. Two properties make it the standard answer:

The naive alternative (fixed windows: “100 requests per minute, reset at :00”) has the classic boundary exploit — 100 requests at 11:59:59 plus 100 at 12:00:01 — which token buckets don’t.

Why Redis Is Not Optional

An in-memory limiter enforces its limit per instance. Three replicas behind a load balancer means your “100 req/min” is actually ~300, the counters vanish on every deploy, and the effective limit changes when you scale. If a rate limit matters enough to build, it matters enough to be shared state — and Redis is the natural home: fast, with atomic compare-and-swap, and TTLs to garbage-collect idle buckets.

Bucket4j ships a Redis integration that stores serialized bucket state and updates it atomically, so N app instances all draw from the same bucket.

The Build

dependencies {
    implementation("org.springframework.boot:spring-boot-starter-web")

    implementation("com.bucket4j:bucket4j-core:8.10.1")
    implementation("com.bucket4j:bucket4j-redis:8.10.1")
    implementation("io.lettuce:lettuce-core")
}

Lettuce (the client Spring Data Redis uses under the hood) connects us to Redis; bucket4j-redis provides the LettuceBasedProxyManager that makes remote buckets look like local ones:

@Configuration
public class RateLimitConfiguration implements WebMvcConfigurer {

    @Bean(destroyMethod = "shutdown")
    RedisClient redisClient(@Value("${rate-limit.redis-uri}") String redisUri) {
        return RedisClient.create(redisUri);
    }

    @Bean
    LettuceBasedProxyManager<byte[]> proxyManager(RedisClient redisClient) {
        return LettuceBasedProxyManager.builderFor(redisClient)
                .withExpirationStrategy(ExpirationAfterWriteStrategy
                        .basedOnTimeForRefillingBucketUpToMax(Duration.ofMinutes(2)))
                .build();
    }
}

The expiration strategy is the detail everyone forgets: without it, every client that ever hits you leaves a bucket key in Redis forever. This strategy sets each key’s TTL to “time until the bucket would be full again, plus a bit” — an idle bucket expires exactly when expiring it is indistinguishable from a fresh one.

The Interceptor: Two Tiers of Limits

Real APIs don’t have one limit. The demo implements the common shape — authenticated callers get a generous per-key limit, anonymous traffic shares a strict per-IP limit:

@Component
public class RateLimitInterceptor implements HandlerInterceptor {

    private final LettuceBasedProxyManager<byte[]> proxyManager;

    // ...

    @Override
    public boolean preHandle(HttpServletRequest request, HttpServletResponse response,
                             Object handler) throws Exception {
        String apiKey = request.getHeader("X-Api-Key");
        String bucketKey = apiKey != null && !apiKey.isBlank()
                ? "rl:key:" + apiKey
                : "rl:ip:" + clientIp(request);
        BucketConfiguration config = apiKey != null ? apiKeyLimit() : anonymousLimit();

        BucketProxy bucket = proxyManager.builder()
                .build(bucketKey.getBytes(StandardCharsets.UTF_8), () -> config);

        ConsumptionProbe probe = bucket.tryConsumeAndReturnRemaining(1);
        if (probe.isConsumed()) {
            response.setHeader("X-Rate-Limit-Remaining",
                    String.valueOf(probe.getRemainingTokens()));
            return true;
        }

        long retryAfterSeconds = probe.getNanosToWaitForRefill() / 1_000_000_000 + 1;
        response.setStatus(HttpStatus.TOO_MANY_REQUESTS.value());
        response.setHeader("Retry-After", String.valueOf(retryAfterSeconds));
        response.setContentType(MediaType.APPLICATION_JSON_VALUE);
        response.getWriter().write("""
                {"error": "rate limit exceeded", "retryAfterSeconds": %d}
                """.formatted(retryAfterSeconds));
        return false;
    }

    /** 100 requests/minute refilled gradually, burst capacity 120. */
    private BucketConfiguration apiKeyLimit() {
        return BucketConfiguration.builder()
                .addLimit(Bandwidth.builder()
                        .capacity(120)
                        .refillGreedy(100, Duration.ofMinutes(1))
                        .build())
                .build();
    }

    /** 20 requests/minute for anonymous traffic. */
    private BucketConfiguration anonymousLimit() {
        return BucketConfiguration.builder()
                .addLimit(Bandwidth.builder()
                        .capacity(20)
                        .refillGreedy(20, Duration.ofMinutes(1))
                        .build())
                .build();
    }
}

Register it for your API paths (registry.addInterceptor(rateLimitInterceptor).addPathPatterns("/api/**")) and you’re enforcing.

Details that make this production-shaped rather than demo-shaped:

Proving It with k6

A rate limiter you haven’t load-tested is a hypothesis. The repo includes a k6 script running two scenarios simultaneously — anonymous traffic at 3x its limit, and API-key traffic just under its own:

export const options = {
  scenarios: {
    anonymous: {
      executor: "constant-arrival-rate",
      rate: 60,
      timeUnit: "1m",
      duration: "2m", // limit is 20/min
      preAllocatedVUs: 10,
    },
    with_api_key: {
      executor: "constant-arrival-rate",
      rate: 90,
      timeUnit: "1m",
      duration: "2m", // limit is 100/min
      preAllocatedVUs: 10,
      exec: "withApiKey",
    },
  },
};

Before reaching for k6, the fastest sanity check is a bash loop against the anonymous endpoint (burst capacity 20) — the first 20 requests should sail through, then the bucket runs dry:

for i in $(seq 1 25); do curl -s -o /dev/null -w '%{http_code}\n' localhost:8080/api/quote; done
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
200
429
429
429
429
429

Exactly what the bucket math predicts: 20 tokens spent, then five straight rejections. No timing games, no ambiguity — the limiter just works.

Run docker compose up -d, ./gradlew bootRun, then k6 run k6/rate-limit-test.js. Over the 2-minute run the counters land where the bucket math says they must: the anonymous scenario sends 120 requests against a budget of 20/min + 20 burst and gets roughly 60 accepted / 60 rejected, while the API-key scenario — under its refill rate the whole time — comes through with zero 429s. Watching k6 report exactly the numbers you predicted from capacity and refillGreedy is the moment this stops feeling like configuration and starts feeling like arithmetic.

     execution: local
        script: k6/rate-limit-test.js
        output: -

     scenarios: (100.00%) 2 scenarios, 20 max VUs, 2m30s max duration (incl. graceful stop):
              * anonymous: 1.00 iterations/s for 2m0s (maxVUs: 10, gracefulStop: 30s)
              * with_api_key: 1.50 iterations/s for 2m0s (maxVUs: 10, exec: withApiKey, gracefulStop: 30s)


  █ TOTAL RESULTS

    checks_total.......: 302     2.516511/s
    checks_succeeded...: 100.00% 302 out of 302
    checks_failed......: 0.00%   0 out of 302

    ✓ status is 200 or 429

    CUSTOM
    responses_200..................: 240    1.999877/s
    responses_429..................: 62     0.516635/s

    HTTP
    http_req_duration..............: avg=6ms    min=1.47ms med=5.97ms max=11.15ms p(90)=8.65ms p(95)=9.12ms
      { expected_response:true }...: avg=6.47ms min=1.72ms med=6.69ms max=11.15ms p(90)=8.72ms p(95)=9.14ms
    http_req_failed................: 20.52% 62 out of 302
    http_reqs......................: 302    2.516511/s

running (2m00.0s), 00/20 VUs, 302 complete and 0 interrupted iterations
anonymous    ✓ [======================================] 00/10 VUs  2m0s  1.00 iters/s
with_api_key ✓ [======================================] 00/10 VUs  2m0s  1.50 iters/s

240 accepted, 62 rejected, zero failed checks — the anonymous scenario (60 req/min against a 20/min limit) is exactly where the token math says it should land, and every single response was a clean 200 or 429, never a timeout or a 500. That’s the whole point of testing a rate limiter under load: it should fail predictably, not fail badly.

There’s also a Testcontainers integration test asserting the precise 20-accepted/5-rejected boundary, in the spirit of the Testcontainers post coming later this week.

What About Filters, Gateways, and Resilience4j?

Fair questions, quick answers:

Summary

Distributed rate limiting on Spring Boot 4 is: Bucket4j for correct token-bucket math, Redis (via LettuceBasedProxyManager) so every instance enforces the same budget, an interceptor that keys buckets by API key or IP, and honest 429s with Retry-After. The whole thing is ~150 lines, adds one Redis round trip per request, and — per the k6 run — does exactly what the math says under 3x overload.

Clone the demo and watch your own 429s roll in.


Next Post
The Ultimate Guide to Spring Security 7 (Migrating from 6)

Related Posts