> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.withpersona.com/2020-05-18/rate-limiting/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.withpersona.com/_mcp/server. # Rate Limiting > Understand environment rate limits, API key rate limits, product quotas, and rate limit response headers. ## Rate Limits We use rate limiting to safeguard the stability of our API. There are two main types of rate limits to consider: Environment Rate Limits and product-specific Quotas. An API key can also have its own rate limit within its environment's limit. ### Environment Rate Limits Every environment has a rate limit that determines the number of API requests that can be made within a specific timeframe. The default rate limiter for production environments allows up to 300 requests per minute timeframe for all API requests from any API key within the environment. The request count starts when the first request is made, allowing up to 300 requests within the first 60 seconds. The count resets at the start of the next minute. You can see the rate limits by navigating to [API > Rate Limits](https://app.withpersona.com/dashboard/api-rate-limits) within the Persona Dashboard, with the appropriate environment selected. Any request over the limit will return a `429 Too Many Requests` error. Our server responses also return your API limit, remaining requests, and seconds until the limit resets as headers. If you `curl` the endpoints with the `-vvv` flag, you'll see the headers as such. ``` ... < RateLimit-Limit: 300 < RateLimit-Remaining: 280 < RateLimit-Reset: 53 ``` ### API Key Rate Limits An API key can have its own rate limit, which can't be set higher than its environment's rate limit. A request made with that key counts against both the key's limit and the environment's limit, and it's rejected if either limit has been reached. The key's limit caps how much of the environment's capacity that key can use. It doesn't add capacity on top of the environment's limit. All API keys in an environment share the environment's limit, so traffic on one key uses capacity that other keys rely on. If you run several integrations on separate API keys, size their limits so their combined traffic fits within the environment's limit. For a key with its own rate limit, the `RateLimit-*` headers describe whichever limit is closer to being reached for that request: the key's or the environment's. `RateLimit-Limit` can therefore change from one response to the next. A `429 Too Many Requests` response reports the limit that rejected the request, so a `RateLimit-Limit` that doesn't match your key's limit means the environment's limit was reached. For a key without its own rate limit, the headers always describe the environment's limit. ### Product-specific Quotas Within an environment, there are product-specific rate limits that determine the number of API calls that can be made to create an object (for a particular product) within a specific timeframe. These product-specific rate limits are referred to as Quotas. As an example, for a particular production environment, there may be an Inquiry Quota which allows up to 300 requests per minute timeframe for Inquiry Created requests from any API key within the environment. The request count starts when the first request is made, allowing up to 300 requests within the first 60 seconds. The count resets at the start of the next minute. Most environments have Quotas for the following: Inquiry creation, Verification creation, Report creation, Case creation, and Transaction creation. You can see Quotas by navigating to [API > Rate Limits](https://app.withpersona.com/dashboard/api-rate-limits) within the Persona Dashboard, with the appropriate environment selected. Any request over the limit will return a `429 Too Many Requests` error. Our server responses also return your API limit, remaining requests, and seconds until the limit resets as headers. If you `curl` the endpoints with the `-vvv` flag, you'll see the headers as such. ``` ... < Quota-Limit: 300 < Quota-Remaining: 280 < Quota-Reset: 53 ``` ## Best Practices * **Structured Retry Logic** - To react to a `429`, we recommend implementing an exponential backoff strategy. This approach helps manage rate limits effectively, ensuring your requests are processed as soon as capacity becomes available. A sample backoff interval could start with delays of 5 seconds, then progressively increase to 10, 20, 40 seconds, and so forth. * **Monitoring `RateLimit` response headers **- [Rate limits and quotas](/rate-limiting) are available [in the dashboard](https://app.withpersona.com/dashboard/api-rate-limits), with additional details provided dynamically in response headers for each API request. We recommend monitoring the `RateLimit-Remaining` header and, if it drops below 15% of the `RateLimit-Limit`, begin throttling requests to avoid reaching the limit. Read both headers from every response rather than caching `RateLimit-Limit`, since for an [API key with its own rate limit](#api-key-rate-limits) it can report either the key's or the environment's limit. This is especially useful for operations that are high volume. You should also consider building alerting in your system to notify you when you're approaching your rate limit. * **Pacing batch jobs** - For high-volume jobs, such as backfills and bulk updates, pace requests using `RateLimit-Remaining` and `RateLimit-Reset` rather than sending as fast as possible and retrying on `429`. Rejected requests still count toward the limit that rejected them, so a job that keeps exceeding a limit keeps it exhausted for the rest of the window, including for your other traffic that shares it. * **Monitoring `Quota` response headers **- [Rate limits and quotas](/rate-limiting) are available [in the dashboard](https://app.withpersona.com/dashboard/api-rate-limits), with additional details provided dynamically in response headers for each API request. For API requests that create an object within Persona, we recommend monitoring the `Quota-Remaining` header and, if it drops below 15% of the `Quota-Limit`, begin throttling requests to avoid reaching the limit. Similarly to `RateLimit` response headers, it can be useful to build alerting into your system. > Understand environment rate limits, API key rate limits, product quotas, and rate limit response headers.