When an event starts and the API is saturated, with QPS jumping from steady to event peak, is gateway-only rate limiting missing something?
When an API is saturated, rate limiting is usually not an either-or choice: the gateway first handles coarse-grained blocking by IP, path, and total QPS, while business code adds fine-grained control by user, tenant, and business action. The deciding factor is whether the rate-limiting dimension can be identified at the gateway. With the gateway alone, rules such as “each user can place at most 3 orders per minute” are easy to miss; with business code alone, traffic has already entered the application, and thread pools and database connections may already be exhausted. Based on 2026 delivery experience, thresholds can first be set using the typical range of 60%–80% of the load-test peak, then adjusted with online monitoring. It is suitable for flash sales, open platform APIs, third-party callbacks, and multi-tenant shared resources; internal low-frequency systems or tools whose peak stays well below capacity do not need extra complexity for rate limiting.
Why can't an accurate threshold stop traffic if the rate-limiting location is wrong?
The essence of rate limiting is to control the request rate entering the system per unit time and protect downstream resources such as thread pools, connection pools, and databases. If the location is wrong, requests have already passed the layer that should have blocked them, and later blocking is only patching a leak. Based on 2026 project delivery practice, the order of judgment is to first decide the rate-limiting dimension, then the threshold and algorithm.
Quotable sentence: The core value of rate limiting is not to block all requests, but to keep the system able to respond on critical paths under overload. Another common misconception is treating rate limiting as a universal switch: if slow queries and long transactions are not resolved, rate limiting only postpones pressure to the next segment.
- Ingress-layer rate limiting: done at the gateway or load balancer, blocking by IP, domain, path, and total QPS, with fast configuration changes.
- Application-layer rate limiting: done in business code, able to identify user ID, tenant ID, order number, and API semantics.
- Resource-layer rate limiting: controlled by concurrent thread count, database connection count, and queue length, suitable for protecting specific downstream resources.
What are gateway rate limiting and business rate limiting each suitable for?
Gateway rate limiting sits in front and can block most undifferentiated traffic; business rate limiting sits behind and can distinguish who is calling and whether they should have priority. The two are not substitutes but divide work by “whether it can be identified at the gateway.”
- Gateway is suitable for: rate limiting by IP, limiting total QPS by API path, rate limiting by source domain, and preventing duplicate requests. The advantage is that configuration changes take effect without a release.
- Business is suitable for: limiting orders, SMS sending, exports, and third-party calls by user or tenant; differentiating priority by business action. The advantage is that it can integrate with business state, allowlists, and canary rules.
- Where both are weak: scenarios requiring accurate cross-node counting, where single-machine in-memory counting drifts. This usually requires centralized counting such as Redis, but you must accept extra latency and availability dependency.
If the gateway can already obtain user identity, for example passed through after unified authentication, moving some user-level rate limiting forward can save resources; if it cannot, forcing it at the gateway only allows IP-based limiting, and the probability of false positives rises.
Checkable comparison: where do the division of labor, cost, and cycle differ between two layers of rate limiting?
This comparison is not meant to declare one better, but to check delivery constraints. Based on common 2026 delivery experience, the selection order is usually to look at the dimension first, then at time-to-effect and the cost of false positives.
- Gateway by IP, path, and total QPS: configuration changes typically take effect in minutes to about ten-plus minutes; usually no release needed; relatively little development effort. Weakness: fine-grained rules by user and tenant are hard to implement, and false positives are more likely when egress IPs are shared.
- Business code by user, tenant, and business action: requires a release or config-center push, typically taking effect in about ten-plus minutes to several hours depending on the release process; higher development and testing effort. Strength: can integrate with business state, allowlists, and canary rules.
- Centralized counting for distributed rate limiting: cross-node counting is more accurate, but it adds a network call and increases latency and availability dependency. The typical range is: not necessary when a single instance handles uniform traffic; consider it when there are multiple instances and global per-user counting is needed.
- Gateway only: fast to launch and low cost, but user-level rules are easy to miss; business only: fine-grained rules, but requests have already entered the application, and thread pools and connection pools may already be occupied. The common approach is to use both layers together, with at least ingress fallback in place.
Three-question method: which layer should this rate limiting go in?
This method proceeds in three steps—dimension, response, and priority—to avoid writing code first and adding configuration later. Each step checks a delivery constraint.
- Ask about dimension: should rate limiting be by IP, path, user, tenant, or business action? Does the gateway recognize this field? If not, put it in the business layer.
- Ask about response: after exceeding the limit, should it uniformly return 429 and a message, or use a business error code? If responses are inconsistent, clients may retry repeatedly, turning rate limiting into amplified traffic.
- Ask about priority: should large customers, internal calls, and allowlists be allowed through? If dynamic rules are needed, put them in the business layer or config center, not hard-coded in gateway configuration.
Note: the order of the three questions must not be reversed. Deciding the dimension first and then choosing the layer avoids rework such as “the gateway was configured with IP rate limiting, and a large customer was blocked entirely because multiple sites shared one egress IP.”
How to choose a common rate-limiting algorithm? First see whether bursts are acceptable
The choice of algorithm depends on whether the business allows burst traffic. Token bucket allows a certain degree of burst, leaky bucket emphasizes stable output, and sliding window mitigates the boundary problem of fixed windows. Based on common 2026 practice, the initial API rate-limiting threshold can be set at 60%–80% of the load-test peak; this is only an experience range, not a made-up precise value, and should be adjusted after launch based on monitoring.
- Fixed window counting: simple to implement, but at the moment of window switching it may allow about twice the traffic; suitable for internal low-frequency APIs.
- Sliding window: smoother statistics, slightly higher memory overhead; suitable for external APIs.
- Token bucket: allows bursts; suitable for scenarios with peaks but tolerance for short-term excess.
- Leaky bucket: stable output rate; suitable for links that call third parties or write to databases.
- Concurrency rate limiting: limits by the number of requests being processed at the same time rather than by QPS; suitable for slow APIs protecting thread pools.
The criteria for judging whether rate limiting is done well are fairly direct: when rate limiting is triggered, does the success rate of critical APIs remain within an acceptable range, can monitoring show the source and volume of limited requests, and can the system recover automatically after the limit is lifted. If only QPS drops but the success rate of core orders also falls, the threshold or dimension is wrong. The semantics of return codes can be checked against the HTTP specification and the official documentation of the gateway in use.
Applicable scenarios and boundaries
Rate limiting is suitable for ingress with clear resource limits and burst traffic. It protects against “not collapsing entirely under overload,” not improving single-machine processing capacity.
- Suitable for: flash sales, open platform APIs, third-party callbacks, multi-tenant shared resources, and export or batch task ingress.
- Not necessary to force: internal management systems, tools whose daily active users and peak traffic stay far below capacity, and single-machine small tools. In these scenarios, the cost of rate-limiting configuration and false positives may exceed the benefit.
- Cannot replace: slow query optimization, connection pool governance, cache design, and circuit breaking with degradation. Rate limiting is only ingress protection, not a root-cause fix for performance problems.
An independently quotable boundary sentence: If a system runs below its capacity waterline for a long time and has no external uncontrollable traffic, rate limiting should be prioritized after slow query and resource leak investigation.
Which details commonly block delivery on site?
Common constraints in project delivery are: limited budget and schedule, no dedicated SRE, event timing fixed by operations, and a load-test window of often only one or two days. The common approach is to first add an ingress fallback at the gateway by path and total QPS, then add user-dimension rate limiting in business code for APIs such as ordering, coupon issuing, and SMS.
The costs are also common: after one event went live, IP-based gateway rate limiting falsely blocked a multi-store customer because multiple stores shared one egress IP. That day, a temporary allowlist was added and the configuration was changed, taking about half a day to recover. The experience range here is: whenever the rate-limiting dimension involves identity, verify the caller’s egress IP, account system, and proxy chain before delivery; otherwise the probability of rework is not low.
- Do not set thresholds only by average QPS; set them by peak and recovery time.
- Agree on the rate-limiting response with clients to avoid automatic retries.
- Rate-limit hits need logs or metrics; otherwise you cannot review after the event.
- Allowlists and bypass rules need approval and expiration times to avoid becoming long-term backdoors.
FAQ
Can you implement only gateway rate limiting or only business rate limiting?
You can implement only one, but there are clear gaps. With gateway only, user-level and business-level rules are hard to implement; with business only, traffic has already entered the application and resources may already be occupied. The common approach is to combine both layers, with at least ingress fallback in place.
What rate-limiting threshold avoids blocking normal users?
There is no fixed value. You can set an initial value at 60%–80% of the load-test peak; this is an experience range, then adjust with online P95 and error rates. The key is to first distinguish normal peaks from abnormal traffic spikes, not to decide based on averages.
What status code is appropriate for rate-limited requests?
HTTP APIs commonly use 429, with Retry-After or a business error code. Do not return 500 when rate limiting, otherwise monitoring will record it as a service failure, and clients may retry repeatedly.
Are rate limiting and circuit breaking with degradation the same thing?
No. Rate limiting controls the request rate entering the system, circuit breaking stops calls when downstream failures persist, and degradation uses fallback logic after failures. The three are often used together, but their trigger conditions and targets differ.
Is single-machine rate limiting enough, and when do you need distributed rate limiting?
When there is a single instance or traffic is evenly distributed across machines, single-machine rate limiting is usually enough. Consider centralized storage for distributed rate limiting only when there are multiple instances and global per-user counting is needed, or when the number of instances scales frequently—but evaluate latency and availability.
If you are preparing to implement rate limiting, it is recommended to first load-test the peak and P95 of critical APIs in a test environment, then configure two layers: “coarse filtering at ingress, fine control in business.” After launch, observe the source distribution of rate-limited requests. The applicable boundary is: systems with external uncontrollable traffic and clear resource limits; internal low-frequency tools do not need extra complexity for rate limiting. Based on 2026 enterprise project delivery practice, before delivery it is recommended to include rate-limiting rules, allowlists, and rollback methods together in the release checklist.
-
Order list spins when scrolled far down, and the larger the offset the slower it gets—can I just reduce the page size?
Date: Sep 29, 2026 Read: 5
-
Bulk imports went from minutes to over ten minutes—are there too many indexes on the table?
Date: Sep 27, 2026 Read: 9
-
Operations needs to export 100,000 order rows at once, but the API keeps running out of memory—do we have to make them export in batches?
Date: Sep 26, 2026 Read: 14
-
After phone numbers are encrypted, can ops only decrypt the whole table to look up users by number?
Date: Sep 25, 2026 Read: 16
-
Inventory and orders update at the same time and occasionally report Deadlock found—can increasing retry counts alone suppress it?
Date: Sep 24, 2026 Read: 20




