Rate Limiting
Rate Limits
We use rate limiting to safeguard the stability of our API. There are two main types of rate limits to consider: Environment Rate Limits and product-specific Quotas. An API key can also have its own rate limit within its environment’s limit.
Environment Rate Limits
Every environment has a rate limit that determines the number of API requests that can be made within a specific timeframe.
The default rate limiter for production environments allows up to 300 requests per minute timeframe for all API requests from any API key within the environment. The request count starts when the first request is made, allowing up to 300 requests within the first 60 seconds. The count resets at the start of the next minute.
You can see the rate limits by navigating to API > Rate Limits within the Persona Dashboard, with the appropriate environment selected.
Any request over the limit will return a 429 Too Many Requests error.
Our server responses also return your API limit, remaining requests, and seconds until the limit resets as headers. If you curl the endpoints with the -vvv flag, you’ll see the headers as such.
API Key Rate Limits
An API key can have its own rate limit, which can’t be set higher than its environment’s rate limit. A request made with that key counts against both the key’s limit and the environment’s limit, and it’s rejected if either limit has been reached. The key’s limit caps how much of the environment’s capacity that key can use. It doesn’t add capacity on top of the environment’s limit.
All API keys in an environment share the environment’s limit, so traffic on one key uses capacity that other keys rely on. If you run several integrations on separate API keys, size their limits so their combined traffic fits within the environment’s limit.
For a key with its own rate limit, the RateLimit-* headers describe whichever limit is closer to being reached for that request: the key’s or the environment’s. RateLimit-Limit can therefore change from one response to the next. A 429 Too Many Requests response reports the limit that rejected the request, so a RateLimit-Limit that doesn’t match your key’s limit means the environment’s limit was reached. For a key without its own rate limit, the headers always describe the environment’s limit.
Ramping Up Traffic
If your environment’s rate limit is above 1,000 requests per minute, Persona may also limit how quickly your traffic can grow toward that limit. This protects both your integration and our API from sudden spikes, such as a new launch, a backfill, or a queue of retries all starting at once.
When a ramp-up limit applies to your environment, it works like this:
- You can always send up to 1,000 requests per minute. Traffic at or below 1,000 requests per minute is never affected by the ramp-up limit, and an environment whose rate limit is 1,000 requests per minute or less never has a ramp-up limit.
- Each minute, you can send a fixed number of requests more than you sent the minute before. This number is your ramp rate. A typical ramp rate is 100 requests per minute: if you sent 1,500 requests in the previous minute, you can send up to 1,600 this minute, never more than your rate limit.
- Your allowance shrinks gradually when your traffic drops. While you keep sending requests, the allowance falls by at most your ramp rate each minute. After about two minutes with no requests at all, it returns to 1,000 requests per minute.
You can view your environment’s ramp-up limit on the Rate Limits page within the Persona Dashboard, with the appropriate environment selected.
With a ramp rate of 100 requests per minute, reaching a rate limit of 2,000 requests per minute takes at least 10 minutes of steadily increasing traffic, and reaching 3,000 takes at least 20 minutes:
In general, ramping from 1,000 requests per minute to a rate limit of L takes (L - 1000) / ramp rate minutes. Requests over the ramp-up allowance return a 429 Too Many Requests error, even when your RateLimit-Remaining header shows capacity left, because the RateLimit-* headers describe your full rate limit rather than your current ramp-up allowance.
To reach your rate limit without being throttled:
- Start at or below 1,000 requests per minute and increase your send rate in steps no larger than your ramp rate each minute. Check your environment’s ramp-up limit on the Rate Limits page in the Persona Dashboard before you start.
- Hold each step for a full minute before increasing again, so that each minute’s traffic builds the allowance for the next.
- Keep traffic flowing. If a high-volume job pauses for more than a minute or two, ramp up again from 1,000 requests per minute when it resumes rather than jumping back to full speed.
- Spread requests evenly across each minute instead of sending them in bursts at the start of the minute.
- Back off on
429responses. Rejected requests still count toward your limits, so slow down rather than retrying immediately. See Best Practices for a retry strategy.
If you’re planning a launch or a large one-time job that needs to reach a high request rate quickly, contact support@withpersona.com ahead of time.
Product-specific Quotas
Within an environment, there are product-specific rate limits that determine the number of API calls that can be made to create an object (for a particular product) within a specific timeframe. These product-specific rate limits are referred to as Quotas.
As an example, for a particular production environment, there may be an Inquiry Quota which allows up to 300 requests per minute timeframe for Inquiry Created requests from any API key within the environment. The request count starts when the first request is made, allowing up to 300 requests within the first 60 seconds. The count resets at the start of the next minute.
Most environments have Quotas for the following: Inquiry creation, Verification creation, Report creation, Case creation, and Transaction creation.
You can see Quotas by navigating to API > Rate Limits within the Persona Dashboard, with the appropriate environment selected.
Any request over the limit will return a 429 Too Many Requests error.
Our server responses also return your API limit, remaining requests, and seconds until the limit resets as headers. If you curl the endpoints with the -vvv flag, you’ll see the headers as such.
Best Practices
- Structured Retry Logic - To react to a
429, we recommend implementing an exponential backoff strategy. This approach helps manage rate limits effectively, ensuring your requests are processed as soon as capacity becomes available. A sample backoff interval could start with delays of 5 seconds, then progressively increase to 10, 20, 40 seconds, and so forth. - Monitoring
RateLimitresponse headers - Rate limits and quotas are available in the dashboard, with additional details provided dynamically in response headers for each API request. We recommend monitoring theRateLimit-Remainingheader and, if it drops below 15% of theRateLimit-Limit, begin throttling requests to avoid reaching the limit. Read both headers from every response rather than cachingRateLimit-Limit, since for an API key with its own rate limit it can report either the key’s or the environment’s limit. This is especially useful for operations that are high volume. You should also consider building alerting in your system to notify you when you’re approaching your rate limit. - Ramping up gradually - When starting a high-volume job or launch, start at or below 1,000 requests per minute and increase your send rate step by step, as described in Ramping Up Traffic, instead of starting at your full rate limit.
- Pacing batch jobs - For high-volume jobs, such as backfills and bulk updates, pace requests using
RateLimit-RemainingandRateLimit-Resetrather than sending as fast as possible and retrying on429. Rejected requests still count toward the limit that rejected them, so a job that keeps exceeding a limit keeps it exhausted for the rest of the window, including for your other traffic that shares it. - Monitoring
Quotaresponse headers - Rate limits and quotas are available in the dashboard, with additional details provided dynamically in response headers for each API request. For API requests that create an object within Persona, we recommend monitoring theQuota-Remainingheader and, if it drops below 15% of theQuota-Limit, begin throttling requests to avoid reaching the limit. Similarly toRateLimitresponse headers, it can be useful to build alerting into your system.

