Guide
Rate limits
Requests to https://mcp.jethost.com/mcp are metered per connection and per account. Normal day-to-day use with an AI assistant stays well within the limits; they exist to keep the service fast and fair when a client loops or runs away.
HTTP 429Retry-AfterRateLimit-* headers
Limits #
| Bucket | Allowance |
|---|---|
| Per connection | 100 requests per 60 s — one access token. Starts afresh when the client refreshes its token. |
| Per account & application | 200 requests per 60 s — every connection one application holds to your account, combined. Survives token refresh. |
- Every POST to the MCP endpoint counts — tool calls and protocol messages alike (initialize, tools/list, ping, notifications).
- Windows are fixed: a window opens with the first request and closes 60 s later, when the allowance is restored in full.
- Signed-in customers see the same numbers in the client area under Usage limits.
Response headers #
Every authenticated response from the MCP endpoint reports the bucket closest to running out, so a client can slow down before it is refused:
| RateLimit-Limit | Allowance of that bucket per window. |
| RateLimit-Remaining | Requests left in the current window. |
| RateLimit-Reset | Seconds until the window resets. |
When you hit a limit #
- The server answers 429 Too Many Requests with a Retry-After header (seconds), the RateLimit-* headers of the bucket that refused, and a JSON body — see the example.
- Wait at least Retry-After seconds, then continue. Retrying sooner is refused again and does not shorten the wait.
- A rate limit is always a 429. A timeout, a dropped connection or a 5xx is not a rate limit — treat it as a transient failure and retry with exponential backoff.