Why a working integration can become overloaded
An integration with a model may work correctly under normal traffic and still degrade when many requests arrive at once. If the application sends every request immediately, without controlling how many are in progress or waiting, a burst can accumulate work faster than the service can complete it. The visible result may be higher latency, requests that are no longer useful by the time they finish, rate-limit responses, or a bill that deviates from expectations.
Concurrency matters because a request does not occupy resources for a fixed instant. Generation can take more or less time depending on the task, input size, and requested output. For that reason, requests per second alone do not describe the pressure an integration creates. Two workloads with the same request rate can have different service times and resource consumption.
The goal of load control is not to keep every possible resource busy at any cost. It is to protect the application’s objectives: complete useful work, respect priorities, limit waiting, and prevent a temporary demand spike from turning the system into an unending queue. The design decisions that follow are general recommendations; they do not describe a universal limit or replace the documented conditions for a particular API, model, or access channel.
It helps to separate two problems that are often conflated. Admission determines, before a request is sent, whether the application accepts the work, holds it briefly in a queue, or rejects it. Retry policy applies after a failure or an uncertain result. This guide focuses mainly on the first decision and on controlled waiting: automatic retries do not fix an admission policy that lets in more work than the system can process.
What to measure before setting limits
Start by observing a representative workload rather than picking a concurrency number by guesswork. Record the arrival rate and the number of simultaneous requests, but also include input size and generated output where available. For an admission decision made before sending, the output is not yet known; it can be measured afterward to improve estimates, but it should not be treated as an exact fact about the future.
Measure end-to-end latency and, if you can instrument the phases separately, distinguish queue time from the time between sending and receiving a response. For streaming generated responses, also record time to the first chunk and total duration. A good time to first chunk does not mean the whole task finished quickly; nor does a long total duration, by itself, prove that the queue caused the delay.
Use percentiles such as p95 and p99 alongside averages. An average summarizes a trend, but can hide a minority of requests that wait much longer. Pair these measures with counts of completed, canceled, expired, and rejected requests; errors returned by the provider; token usage where available; and cost per completed task if your system can calculate it reliably.
Little’s Law provides a way to check the relationship between the average number of units in a system, the average arrival rate, and the average time each unit spends in it: L = λW. Used carefully, it helps explain why, if the arrival rate remains steady while time in the system increases, the average number of jobs present can also grow. It is not a formula for predicting percentiles, spikes, costs, or an API’s limit on its own; it describes a relationship between averages under the relevant conditions.
Break measurements down by model, endpoint, provider, region, or project when those differences matter to your configuration. Do not combine short and long tasks in a single series if that hides problems affecting one category. Also keep production traffic separate from tests, and record configuration changes so you can connect a latency change with a plausible cause.
Metrics and the decisions they inform
Use these measurements as complementary signals. None of them, in isolation, proves what the safe limit is.
| Measure | What it helps you observe | Caution |
|---|---|---|
| Requests per interval | Arrival rate and bursts | Does not reflect the duration or size of each task |
| In-flight concurrency | Work occupying capacity at a given moment | Define which states count as “in progress” |
| Input and output tokens | Observed size of requests and responses | Future output is unknown at admission time |
| Queue time and end-to-end latency | Local waiting and the complete user experience | Do not attribute time measured before sending to the provider |
| p95, p99, and expirations | Long queues and tasks that lose their usefulness | Interpret percentiles alongside the number of observations |
| Rejections, errors, and cost per task | Operational and economic consequences | Define consistently what counts as a completed task |
Do not confuse rate, concurrency, tokens, and your own capacity
A rate limit controls how many requests or counted units may be submitted during a period. A concurrency limit restricts how many operations the application keeps active at the same time. A token-based limit concerns the volume of text processed or requested under the service’s rules. An application-level limit is an additional decision: for example, how many jobs the local queue accepts or how long it allows them to wait.
These constraints solve different problems. A system can stay within an average rate and still accumulate a short burst; it can have only a few simultaneous requests that take a long time; or it can receive a small number of requests with lengthy inputs. The application must know the published constraints for the channel it uses and also set local controls that protect its user experience.
Do not transfer quotas from one product to another by analogy. Vertex AI documentation describes quotas and limits whose scope may depend on the service, project, and region. The OpenAI API reference includes response headers related to request and token limits. These examples show why you need to check the documentation that applies to your specific account and configuration; they do not establish shared figures or a single rule for all endpoints.
A 429 response does not always identify a single cause or solution, either. Vertex AI documentation distinguishes situations related to shared capacity from those related to provisioned capacity. Therefore, record the type of response and consult the relevant documentation before deciding what it means. In particular, do not turn every 429 response into an instruction to retry immediately: that behavior belongs to post-failure handling and can increase pressure if admission remains open.
Design admission: accept, wait, or reject
An admission policy should have three explicit outcomes. Accept means the request can start under the local limits and known quotas. Wait means it is held temporarily in a queue with a defined capacity and deadline. Reject means the system does not promise to process it at that time and returns a response that lets the client application decide what to do.
An unbounded queue is not a safe solution: it can turn visible overload into increasingly long waits and stored work that arrives too late to be useful. Set a maximum number of items, a waiting budget, and an expiration condition. If the queue reaches its limit, reject new work or apply a replacement rule justified by the product; do not hide the problem by adding storage without an operational limit.
The rejection response should be consistent with your service interface. Explain that the task was not admitted or indicate that capacity is temporarily occupied, without claiming that the model processed it. When the application can try again later, provide a clear, documented signal so the client can make that decision. Avoid promising an exact wait time if your measurements cannot support it.
Admission can combine conditions: available concurrency, a queue below its maximum, a per-user quota, and remaining waiting time. Evaluate the conditions before reserving resources, and release the reservation when the task completes, is canceled, or expires. If a request is canceled by the client but continues to occupy a local slot, the concurrency metric no longer represents the work that is actually active.
Set the priority order explicitly. For example, you could reserve a fraction of capacity for interactive tasks and limit batch jobs, provided those categories exist in your product and the policy does not leave the lower-priority category without service indefinitely. The goal is not to invent a universal priority, but to reflect the value and deadline of each class of work.
Admission decision for each request
Recommended local flow; calibrate thresholds for each service and workload.
- 01Validate that the task is admissible and determine its class, user, and useful deadline.
- 02Check available concurrency, the local quota, and remaining queue capacity.
- 03If the task can start, reserve capacity and send it; measure waiting and processing separately.
- 04If it cannot start but there is room and waiting budget, enqueue it with an expiration time.
- 05If there is no room or the task can no longer be completed in time, reject it explicitly.
- 06When it completes, is canceled, or expires, release the reservation and record the outcome to help adjust the policy.
Weight work without pretending you know its exact cost
Counting requests is a simple rule, but it treats tasks alike even when they may differ substantially. One alternative is to assign each request an estimated weight using signals available before sending: estimated input tokens, expected output length, task class, or historical duration measurements for that type of operation. The weight can help prioritize work, limit budgets, or decide whether a task fits in the queue.
An estimate is not a guarantee. The generated response may be shorter or longer than expected; duration may vary even when inputs are similar in size. So calibrate the estimate against observed results, retain safety margins, and review prediction errors. If actual output is recorded, use it to improve future policies rather than presenting a value that was unavailable at admission time as if it had been known.
One practical option is to set separate limits by class or reserve an expected amount of work per window. Another is to use an internal credit system in which each request consumes an estimated weight and capacity is replenished according to the local policy. In either case, document how the weight is calculated, how it is corrected, and what happens when a request exceeds the estimate. Do not present local token accounting as if it were identical to the provider’s quota.
Compare policies using the same set of requests and the same workload. If a weighted policy reduces peak waiting time for small jobs, also check what happens to large tasks: they could be repeatedly pushed back. Success criteria should consider aggregate throughput, latency by class, and the share of tasks completed within their deadline—not just a favorable metric for the shortest requests.
Priority, fairness, and protection against blocking
A single queue may be enough for a product with equivalent tasks, but it has risks when it mixes urgent work with long-running jobs. A lengthy task near the front can make other tasks wait even when they are short. In addition, if one client generates a disproportionate share of demand, it can consume shared capacity and harm others.
You can separate queues by service class, set per-user quotas, or limit the number of simultaneous tasks per client. You can also reserve capacity for interactive traffic and process batch work using whatever headroom remains. These are design options, not a guarantee of fairness: the priority, reservation size, and selection rule must be tested against real usage patterns.
Define what fairness means in your case. It might mean that every client gets an opportunity to make progress, that interactive tasks meet a deadline, or that no user consumes the entire local budget. A strict per-user quota can protect against hogging, but it can also leave capacity unused when other users are inactive. A quota that is too flexible may fail to protect smaller clients during a burst.
To prevent a low-priority class from being postponed indefinitely, consider maximum wait times, priority aging, or minimum service opportunities. Measure queue time by class and client, not just the global average. If categories have different deadlines, record how many finish on time and how many expire. A policy that improves latency for one class at the expense of making another useless should be presented as an explicit product decision.
Choosing a queueing rule
These options illustrate design trade-offs; the right choice depends on the service’s objectives.
| Rule | Can help with | Risk to measure |
|---|---|---|
| One shared queue | Keeping implementation simple when tasks are comparable | A long task or high volume from one client delaying everyone else |
| Separate queues by priority | Protecting work with different deadlines | The lower-priority class being postponed |
| Per-client quota | Limiting local capacity hogging | Wasting capacity if other clients cannot use what the quota leaves unused |
| Weighted budget | Distinguishing tasks by their expected cost | Misclassifying requests or penalizing large tasks too heavily |
Waiting budgets, cancellation, and expiration
Each task should have a useful deadline defined by the product, even if that deadline is not exposed directly to the user. If a request remains queued beyond the point at which it can still provide value, keeping it only accumulates stale work. Set a maximum waiting budget and check it again before sending the task to the provider.
Expiration in the queue and cancellation after sending are not the same thing. Before sending, removing a task can prevent work from starting when it is no longer needed. After sending, the effect of cancellation depends on the integration’s capabilities and the documented behavior of the endpoint. Do not assume that closing a connection stops remote processing or reverses consumption. Instrument what the system can observe and describe the limitations.
Record how many requests expire before starting and how many are canceled in progress. If many expire in the queue, review queue size, admission rate, and waiting budget. If requests are canceled after being sent, investigate whether an earlier expiration policy or an interface that lets users withdraw work could reduce wasted tasks. Do not infer cost savings without measurements to support that conclusion.
Expiration can also help manage priorities: an urgent task that has already missed its deadline should not keep its place at the front simply because it arrived first. However, automatically discarding work can have functional consequences. Define whether an expiration is reported, whether the task is kept for later execution, or whether it is deleted; the behavior should be predictable for anyone integrating with the service.
Reproducible load testing and operations
Test the policy before applying it broadly. Build a representative set of tasks that includes different input and output lengths and the priority classes the product will actually use. Run a baseline workload, then increase concurrency in stages and add controlled bursts. Keep the request set constant when comparing variants so the differences do not depend on testing different tasks.
Define stop and rollback criteria in advance. For example, you might stop increasing load if p99 exceeds the agreed objective, if expirations or errors rise, or if queue time exceeds the product’s budget. The specific values should come from your objectives, measurements, and API conditions, not from a universal figure. Record the configuration, date, client version, and test parameters so the test can be repeated.
Break latency down into components: local waiting time, time to the first response where relevant, and total duration. Also compare completed requests, rate-limit responses, cancellations, percentiles by class, and cost per completed task. If you look only at total throughput, you might miss a failing class; if you look only at an error response, you might miss that the problem was a local queue that began growing before the provider was called.
In production, review the published quotas for the specific channel and monitor the headers or response fields documented by that service. Those signals can help identify limits and feed observability, but their availability and meaning depend on the API. Do not assume a header exists in every response or remains unchanged. Keep track of when you reviewed the documentation and verify changes before updating controls.
A cautious rollout starts with conservative local limits, observability, and a way to return to the previous configuration. Admission can then be increased gradually if results support that decision. If metrics worsen, reduce incoming work, shorten the queue, or adjust task classes; do not automatically respond to rising latency by expanding the queue, because that can hide overload and extend waits.
Testing and adjustment sequence
A useful comparison should be repeatable and have defined exit criteria.
- 01Define latency objectives, maximum wait, completion rate, and rollback conditions.
- 02Prepare tasks of different lengths, priority classes, and a documented baseline workload.
- 03Increase concurrency gradually and add a controlled burst without changing the task set.
- 04Separate queue time, initial response, and total duration; record percentiles and results by class.
- 05Compare rejections, expirations, errors, and cost per completed task with the previous configuration.
- 06Keep or roll back the change according to the objectives; repeat the test after relevant changes to the model, endpoint, or quota.
Limits of this approach and practical criteria
There is no concurrency figure that can be safely transferred to every integration. Quotas and conditions may vary by service, model, project, region, and channel. Quotas can also change, and an observed result in one test does not prove that the same capacity will be available in another configuration or at another time. Consult the applicable official documentation and verify limits before turning them into permanent rules.
There is also no single formula that converts tokens or requests per second into guaranteed latency. The relationship between arrivals, time in the system, and average population helps reason about queue growth, but it does not replace a representative test or predict every response. Output sizes and observed durations provide information; prior estimates support cautious admission, not an exact outcome guarantee.
Before approving a policy, check that you know how many requests are in progress and waiting; that the queue has a maximum and an expiration rule; that admission accounts for relevant differences in workload size; and that there is a clear response when a task is not accepted. Also make sure priorities do not leave a class without service, cancellations are measured in the correct state, and metrics separate local waiting from remote processing.
Finally, preserve the distinction between load control and error recovery. Admission determines how much demand enters and how much work waits. Retries determine how to respond after a failure or uncertain result. A robust system needs compatible policies for both phases, but one does not replace the other: first avoid accepting more work than you can manage; then define, separately, how to respond to failures.
Open questions
- The provided sources do not include specific quota, concurrency, or token figures for the services mentioned; these must be verified for each account, model, endpoint, and region.
- The availability and meaning of limit headers or fields may vary between responses and change over time.
- The effect of canceling a request that has already been sent depends on the API and is not determined by the supplied sources.
- Each application’s workload, durations, and latency objectives are unspecified; thresholds must be derived from its own tests.
- Recommendations about queues, priorities, estimated weights, and testing are design criteria, not performance guarantees.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction