Skip to main content
Rate limits are enforced at the Vercel Edge runtime using Upstash Redis counters. All limits are sliding 60-second windows unless noted.

Default public API rate limit

Applies to all /api/* routes that don’t have a stricter override. Implemented by api/_rate-limit.js (legacy api/*.js edge functions) and server/_shared/rate-limit.ts (the gateway and .ts edge functions).

MCP server

See MCP for details.

Per-plan API rate limits

Authenticated REST API keys (wm_…) are limited per account, not per IP — a key behind a shared egress IP is not throttled by other tenants’ traffic, and all of an account’s keys share one allowance.
  • Per-minute is a hard burst limit — exceeding it returns 429 immediately.
  • Daily included is your plan’s allowance; it resets at 00:00 UTC. Requests beyond it are rejected with 429 — the sold plan limit is authoritative, with no overage headroom. Usage is metered against the same counter that enforces, so the Settings notice and the 429 agree.
  • The per-minute burst and daily allowance are both per account (shared across all of an account’s wm_… keys), so issuing more keys does not raise your limit. (Operator-issued Enterprise keys are the exception — each is rate-limited independently.)
  • Need a higher limit? Contact support to raise your plan’s allowance.

Dashboard AI quota

Dashboard and direct REST AI operations use a separate daily budget from MCP. The counter resets at 00:00 UTC. Free and signed-out dashboard users keep the normal keyword/cache fallback for feed enrichment; they do not consume the paid direct-AI budget. These limits are independent of the MCP allowances above. Signed-out callers are rejected outright, and Pro-gated AI routes deny free accounts before any spend. Separately from these plan allowances, a small non-plan safety floor of 50 requests/day applies to any caller whose paid entitlement cannot be confirmed at request time — a lapsed subscription, or a transient entitlement-lookup outage. It exists so an outage degrades gracefully instead of rejecting paying customers; it is not an allowance any plan includes, and it is never larger than the smallest paid allowance.

Stock backtest provider-work quota

GET /api/market/v1/backtest-stock is not LLM-backed. On a cache miss it fetches Yahoo Finance history for a caller-chosen symbol, so it must not share llm:direct-usage or dashboardAiCallsPerDay. Independently of the 60 requests / 60 s route policy: The 200 ceiling is four full 50-symbol Pro watchlist hydrations. Cached repeats and invalid symbols do not consume the budget. Exceeding it returns 429 with Retry-After until the next 00:00 UTC. If the quota store cannot prove a reservation, the route fails closed with 503 rather than fetching Yahoo.

OAuth endpoints

Matches the implementations in api/oauth/register.js, api/oauth/authorize.js, and api/oauth/token.ts. For /api/oauth/token, the limiter key is client_secret hash for client_credentials, then client_id when present, and only falls back to caller IP when neither credential identifier is available. All three grant types (authorization_code, refresh_token, client_credentials) fail open when the Upstash limiter is unconfigured or throws. Token persistence still fails closed when Redis storage is down; a limiter-only 503 would abort MCP client handshakes during an SDK timeout even when the pipeline path still works. client_credentials keeps the operator env-key allowlist as a second gate. Degradation is observable: a bounded/deduplicated [rate-limit] redis-error log and Sentry capture, X-RateLimit-Mode: degraded on the response (listed in Access-Control-Expose-Headers so cross-origin JS can read it), and a usage reason of rate_limit_degraded. Genuine budget exhaustion remains HTTP 429 rate_limit_exceeded. Exceeding any of these during the OAuth flow will cause the MCP client to fail the connection handshake — wait 60 s and retry.

Provider proxies

Routes that fetch a third-party host on our behalf carry their own per-IP budget, so a single scripted caller cannot drive unbounded traffic to a provider we do not control. These budgets are per IP, not aggregate: they bound any one caller, but they do not cap total egress across all callers. The two edge handlers (/api/skills/fetch-agentskills, /api/youtube/live) enforce their budgets in-handler via checkScopedRateLimit/checkRateLimit; /api/reverse-geocode mirrors its per-IP budget as a literal constant per api/*.js constraints. /api/infrastructure/v1/reverse-geocode is a gateway RPC, so the gateway enforces its per-IP budget through checkEndpointRateLimit (fail-closed on Redis outage). After a shared-cache miss, both reverse-geocode handlers also use one fail-closed provider-wide Redis bucket capped at 1 request per second before they call Nominatim; cache hits do not consume that aggregate budget.

Write endpoints

Other write endpoints (/api/brief/share-url, /api/notification-channels, /api/create-checkout, /api/customer-portal, etc.) fall back to the default per-IP limit above.

Bootstrap / health / version

These mostly use the default public API limit. Cache headers vary by endpoint:
  • GET /api/bootstrap — only the explicitly-marked ?...&public=1 URLs are shared-cacheable. ?tier=fast&public=1 / ?tier=slow&public=1 use browser max-age=60 / max-age=300 and CDN s-maxage=600 / s-maxage=7200. Single-key public URLs: on-demand keys (?keys=<onDemandName>&public=1) inherit the slow profile — browser max-age=300, CDN s-maxage=7200 — unless the key declares its own, which every key published more often than that shield does: correlationCards (browser max-age=60, CDN s-maxage=300), chinaDecisionSignals (browser max-age=60, CDN s-maxage=900), canadaRoads (browser max-age=60, CDN s-maxage=900), albertaRoads (browser max-age=60, CDN s-maxage=900), manitobaRoads (browser max-age=60, CDN s-maxage=900), marketCorrelationSeries (browser max-age=60, CDN s-maxage=900), imdCycloneMarine (browser max-age=60, CDN s-maxage=900), bcOpen511 (browser max-age=60, CDN s-maxage=1800), flightDelays (browser max-age=60, CDN s-maxage=1800), and forecasts (browser max-age=300, CDN s-maxage=3600); ?keys=weatherAlerts&public=1 uses Cache-Control: public, s-maxage=600, stale-while-revalidate=120, stale-if-error=900 with the fast-tier CDN shield. Every other shape — key-authenticated, session-authenticated, the unmarked ?tier=... URLs, and the anonymous ?keys=weatherAlerts path — uses Cache-Control: no-store and emits no CDN cache headers, so a credentialed URL can never be answered from a shared cache. User API key validation also has a fail-closed fixed 60 s per-IP pre-validation limit of 600 attempts.
  • GET /api/healthprivate, no-store, max-age=0 plus CDN-Cache-Control: no-store.
  • GET /api/versionpublic, s-maxage=300, stale-while-revalidate=60, stale-if-error=3600.

Rate limit response headers (self-throttle before a 429)

Every /api/* response — success or error — advertises the IETF RateLimit header fields so an agent can pace itself before it trips a 429:
  • RateLimit-Policy — the applicable quota (q) over a window of w seconds for the default sliding window. Stricter per-endpoint, per-plan, and OAuth limits (see the tables above) apply on those routes.
  • RateLimit-Limit — the same quota as a bare integer, for parsers that predate the structured-field draft.
These are static advertisements, so they add no latency on the hot path. The legacy X-RateLimit-* names are also emitted for back-compat.

Response when limited

An HTTP 429 additionally carries the live per-window counters (remaining is 0; the reset and Retry-After are delta-seconds) plus the combined RateLimit member:
Note the IETF RateLimit-Reset (and the t value in the combined RateLimit member) is seconds remaining, whereas the legacy X-RateLimit-Reset is an absolute epoch in milliseconds. For a daily-ceiling 429 the Retry-After counts down to the next 00:00 UTC.

Retry guidance

  • Respect Retry-After. Don’t pound on a 429.
  • For batch work, pace yourself: at 600 req/min/IP the default gives you ~10 req/s headroom.
  • For MCP, 60/min is generous for conversational use but tight for scripted batch fetches — prefer the REST API for batch.
  • Spurious 429s often mean you’re sharing an egress IP (corporate proxy, CI runner). Contact support for a per-key limit bump if needed.

Customer notifications and paid-plan caps

API and MCP plan caps are tracked against the product catalog limits that ship with entitlements: When a paid user approaches or exceeds one of these limits, WorldMonitor records a compact Convex rollup and opens a current account notice in Settings. The daily count is read from the same per-account meter that governs enforcement, so a warning reflects the same usage number the plan is metered against. Daily limits warn at 80% and switch to over-limit at 100%; burst limits only notify on sustained pressure, not a single isolated spike. If the notice is still current, a Resend-backed lifecycle sends an email at a bounded cadence. The email and dashboard notice explain the current usage, the relevant plan limit, and the available options: reduce traffic, wait for reset, upgrade when a self-serve path exists, or contact support when the next tier is not self-serve. WorldMonitor does not automatically upgrade a user, charge for overages, or move a customer into API Business because they crossed a cap. Any future hard enforcement for paid plans must first pass the internal apiPlanLimitNotices.getEnforcementReadiness gate: no stale usage source, no pending/failed email, and no blocked self-serve upgrade path.

Hard caps (not soft limits)

  • Webhook callback URLs must be HTTPS (except localhost).
  • api/download file sizes capped at ~50 MB per request.
  • POST /api/scenario/v1/run-scenario globally pauses new jobs when the pending queue exceeds 100 — returns 429.
  • api/v2/shipping/webhooks TTL is 30 days — re-register to extend.