Default public API rate limit
Applies to all
/api/* routes that don’t have a stricter override. Implemented by api/_rate-limit.js (legacy api/*.js edge functions) and server/_shared/rate-limit.ts (the gateway and .ts edge functions).
MCP server
See MCP for details.
Per-plan API rate limits
Authenticated REST API keys (wm_…) are limited per account, not per IP — a key behind a shared egress IP is not throttled by other tenants’ traffic, and all of an account’s keys share one allowance.
- Per-minute is a hard burst limit — exceeding it returns 429 immediately.
- Daily included is your plan’s allowance; it resets at 00:00 UTC. Requests beyond it are rejected with 429 — the sold plan limit is authoritative, with no overage headroom. Usage is metered against the same counter that enforces, so the Settings notice and the 429 agree.
- The per-minute burst and daily allowance are both per account (shared across all of an account’s
wm_…keys), so issuing more keys does not raise your limit. (Operator-issued Enterprise keys are the exception — each is rate-limited independently.) - Need a higher limit? Contact support to raise your plan’s allowance.
Dashboard AI quota
Dashboard and direct REST AI operations use a separate daily budget from MCP. The counter resets at 00:00 UTC.
Free and signed-out dashboard users keep the normal keyword/cache fallback for feed enrichment; they do not consume the paid direct-AI budget. These limits are independent of the MCP allowances above.
Signed-out callers are rejected outright, and Pro-gated AI routes deny free accounts before any spend. Separately from these plan allowances, a small non-plan safety floor of 50 requests/day applies to any caller whose paid entitlement cannot be confirmed at request time — a lapsed subscription, or a transient entitlement-lookup outage. It exists so an outage degrades gracefully instead of rejecting paying customers; it is not an allowance any plan includes, and it is never larger than the smallest paid allowance.
Stock backtest provider-work quota
GET /api/market/v1/backtest-stock is not LLM-backed. On a cache miss it fetches Yahoo Finance history for a caller-chosen symbol, so it must not share llm:direct-usage or dashboardAiCallsPerDay. Independently of the 60 requests / 60 s route policy:
The 200 ceiling is four full 50-symbol Pro watchlist hydrations. Cached repeats and invalid symbols do not consume the budget. Exceeding it returns 429 with
Retry-After until the next 00:00 UTC. If the quota store cannot prove a reservation, the route fails closed with 503 rather than fetching Yahoo.
OAuth endpoints
Matches the implementations in
api/oauth/register.js, api/oauth/authorize.js, and api/oauth/token.ts.
For /api/oauth/token, the limiter key is client_secret hash for client_credentials, then client_id when present, and only falls back to caller IP when neither credential identifier is available.
All three grant types (authorization_code, refresh_token, client_credentials) fail open when the Upstash limiter is unconfigured or throws. Token persistence still fails closed when Redis storage is down; a limiter-only 503 would abort MCP client handshakes during an SDK timeout even when the pipeline path still works. client_credentials keeps the operator env-key allowlist as a second gate. Degradation is observable: a bounded/deduplicated [rate-limit] redis-error log and Sentry capture, X-RateLimit-Mode: degraded on the response (listed in Access-Control-Expose-Headers so cross-origin JS can read it), and a usage reason of rate_limit_degraded. Genuine budget exhaustion remains HTTP 429 rate_limit_exceeded.
Exceeding any of these during the OAuth flow will cause the MCP client to fail the connection handshake — wait 60 s and retry.
Provider proxies
Routes that fetch a third-party host on our behalf carry their own per-IP budget, so a single scripted caller cannot drive unbounded traffic to a provider we do not control. These budgets are per IP, not aggregate: they bound any one caller, but they do not cap total egress across all callers.
The two edge handlers (
/api/skills/fetch-agentskills, /api/youtube/live) enforce their budgets in-handler via checkScopedRateLimit/checkRateLimit; /api/reverse-geocode mirrors its per-IP budget as a literal constant per api/*.js constraints. /api/infrastructure/v1/reverse-geocode is a gateway RPC, so the gateway enforces its per-IP budget through checkEndpointRateLimit (fail-closed on Redis outage). After a shared-cache miss, both reverse-geocode handlers also use one fail-closed provider-wide Redis bucket capped at 1 request per second before they call Nominatim; cache hits do not consume that aggregate budget.
Write endpoints
Other write endpoints (
/api/brief/share-url, /api/notification-channels, /api/create-checkout, /api/customer-portal, etc.) fall back to the default per-IP limit above.
Bootstrap / health / version
These mostly use the default public API limit. Cache headers vary by endpoint:GET /api/bootstrap— only the explicitly-marked?...&public=1URLs are shared-cacheable.?tier=fast&public=1/?tier=slow&public=1use browsermax-age=60/max-age=300and CDNs-maxage=600/s-maxage=7200. Single-key public URLs: on-demand keys (?keys=<onDemandName>&public=1) inherit the slow profile — browsermax-age=300, CDNs-maxage=7200— unless the key declares its own, which every key published more often than that shield does:correlationCards(browsermax-age=60, CDNs-maxage=300),chinaDecisionSignals(browsermax-age=60, CDNs-maxage=900),canadaRoads(browsermax-age=60, CDNs-maxage=900),albertaRoads(browsermax-age=60, CDNs-maxage=900),manitobaRoads(browsermax-age=60, CDNs-maxage=900),marketCorrelationSeries(browsermax-age=60, CDNs-maxage=900),imdCycloneMarine(browsermax-age=60, CDNs-maxage=900),bcOpen511(browsermax-age=60, CDNs-maxage=1800),flightDelays(browsermax-age=60, CDNs-maxage=1800), andforecasts(browsermax-age=300, CDNs-maxage=3600);?keys=weatherAlerts&public=1usesCache-Control: public, s-maxage=600, stale-while-revalidate=120, stale-if-error=900with the fast-tier CDN shield. Every other shape — key-authenticated, session-authenticated, the unmarked?tier=...URLs, and the anonymous?keys=weatherAlertspath — usesCache-Control: no-storeand emits no CDN cache headers, so a credentialed URL can never be answered from a shared cache. User API key validation also has a fail-closed fixed 60 s per-IP pre-validation limit of 600 attempts.GET /api/health—private, no-store, max-age=0plusCDN-Cache-Control: no-store.GET /api/version—public, s-maxage=300, stale-while-revalidate=60, stale-if-error=3600.
Rate limit response headers (self-throttle before a 429)
Every/api/* response — success or error — advertises the IETF RateLimit header fields so an agent can pace itself before it trips a 429:
RateLimit-Policy— the applicable quota (q) over a window ofwseconds for the default sliding window. Stricter per-endpoint, per-plan, and OAuth limits (see the tables above) apply on those routes.RateLimit-Limit— the same quota as a bare integer, for parsers that predate the structured-field draft.
X-RateLimit-* names are also emitted for back-compat.
Response when limited
An HTTP 429 additionally carries the live per-window counters (remaining is0; the reset and Retry-After are delta-seconds) plus the combined RateLimit member:
RateLimit-Reset (and the t value in the combined RateLimit member) is seconds remaining, whereas the legacy X-RateLimit-Reset is an absolute epoch in milliseconds. For a daily-ceiling 429 the Retry-After counts down to the next 00:00 UTC.
Retry guidance
- Respect
Retry-After. Don’t pound on a 429. - For batch work, pace yourself: at 600 req/min/IP the default gives you ~10 req/s headroom.
- For MCP, 60/min is generous for conversational use but tight for scripted batch fetches — prefer the REST API for batch.
- Spurious 429s often mean you’re sharing an egress IP (corporate proxy, CI runner). Contact support for a per-key limit bump if needed.
Customer notifications and paid-plan caps
API and MCP plan caps are tracked against the product catalog limits that ship with entitlements:
When a paid user approaches or exceeds one of these limits, WorldMonitor records a compact Convex rollup and opens a current account notice in Settings. The daily count is read from the same per-account meter that governs enforcement, so a warning reflects the same usage number the plan is metered against. Daily limits warn at 80% and switch to over-limit at 100%; burst limits only notify on sustained pressure, not a single isolated spike.
If the notice is still current, a Resend-backed lifecycle sends an email at a bounded cadence. The email and dashboard notice explain the current usage, the relevant plan limit, and the available options: reduce traffic, wait for reset, upgrade when a self-serve path exists, or contact support when the next tier is not self-serve.
WorldMonitor does not automatically upgrade a user, charge for overages, or move a customer into API Business because they crossed a cap. Any future hard enforcement for paid plans must first pass the internal
apiPlanLimitNotices.getEnforcementReadiness gate: no stale usage source, no pending/failed email, and no blocked self-serve upgrade path.
Hard caps (not soft limits)
- Webhook callback URLs must be HTTPS (except localhost).
api/downloadfile sizes capped at ~50 MB per request.POST /api/scenario/v1/run-scenarioglobally pauses new jobs when the pending queue exceeds 100 — returns 429.api/v2/shipping/webhooksTTL is 30 days — re-register to extend.
