Self-Hosting

Rate Limits

Default rate limits, per-endpoint behavior, run limits and how to tune them.

Appstrate uses rate-limiter-flexible. With REDIS_URL set (Tier 2 and up), buckets are shared across instances. Without Redis, buckets are process-local and reset on restart.

How buckets are keyed

Limits are per endpoint and per identity, so one caller hitting two endpoints has two counters. The key combines the HTTP method, the matched route pattern (not the literal URL, so /api/files/{id}/content is one bucket for all ids) and the identity:

CallerIdentity in the key
Dashboard sessionThe user id
API key (apst_...)apikey:<key id>
Unauthenticated routeip: and the client IP
Run-bound internal routesThe run id carried by the verified run token
Inbound MCP serverThe API key, end-user or user id, falling back to the client IP

The client IP comes from X-Forwarded-For only as far as TRUST_PROXY allows. Behind a reverse proxy, set TRUST_PROXY to the number of proxies that append to the header, or every caller is counted as the proxy. See Docker Compose.

On top of per-endpoint limits, two organization-wide limits run when a run is launched:

  • Runs per minute per organization (per_org_global_rate_per_min), covering agent runs, inline runs, scheduled runs and remote runs (POST /api/runs/remote).
  • Concurrent runs per organization (max_concurrent_per_org).

Run limits (PLATFORM_RUN_LIMITS)

A JSON object, validated strictly at boot. Unknown keys fail the boot. The defaults apply when the variable is unset or {}, so the system is never unlimited out of the box.

PLATFORM_RUN_LIMITS='{"timeout_ceiling_seconds":1800,"per_org_global_rate_per_min":200,"max_concurrent_per_org":50}'
KeyDefaultMeaning
timeout_ceiling_seconds1800Maximum runtime of any run. A longer timeout declared by an agent is clamped to it
per_org_global_rate_per_min200Run launches per minute per organization. Over the limit: 429 with code org_run_rate_limited
max_concurrent_per_org50Concurrent runs per organization. Over the limit: 429 with code org_run_concurrency_exceeded. The check is atomic per organization
agent_memory_ceiling_mb1536Memory ceiling per run, in MiB
agent_cpu_ceiling2vCPU ceiling per run

The last two bound what an agent manifest may request. See the resource guide.

Inline run limits (INLINE_RUN_LIMITS)

These cap POST /api/runs/inline and POST /api/runs/inline/validate.

KeyDefaultMeaning
rate_per_min60Inline runs per minute, per identity on that endpoint
manifest_bytes65536Maximum size of the inline manifest
prompt_bytes200000Maximum size of the agent prompt
max_skills20Maximum number of skills in an inline manifest (0 allowed)
retention_days30Days before the temporary package behind an inline run is emptied (manifest and prompt) and its run logs deleted. The run record is kept

Proxy limits

Two JSON variables, validated the same strict way, cap the platform's model and credential proxies:

  • LLM_PROXY_LIMITS for /api/llm-proxy/*: rate_per_min (default 60) and max_request_bytes (10 MiB).
  • CREDENTIAL_PROXY_LIMITS for /api/credential-proxy/proxy: rate_per_min (100), max_request_bytes (10 MiB), max_response_bytes (50 MiB) and session_ttl_seconds (3600).

Neither the root Compose file nor the files in examples/self-hosting/ forward these two variables, so with Docker Compose add a bare - LLM_PROXY_LIMITS or - CREDENTIAL_PROXY_LIMITS line under appstrate.environment before setting them in .env (see Docker Compose).

API_BODY_LIMIT_BYTES (default 10 MiB) caps request bodies globally. Durable files are capped separately by FILE_MAX_BYTES (default 100 MiB), and ORG_STORAGE_QUOTA_BYTES sets an optional per-organization storage quota. See Environment Variables.

Per-endpoint limits

These are the fixed per-minute limits set in code. The list is representative. The route files under apps/api/src/routes/ are authoritative.

EndpointLimit per minute
POST /api/agents/{scope}/{name}/run20
POST /api/runs/remoteper_org_global_rate_per_min
PATCH /api/runs/{id}/sink/extend30
GET /api/runs/{id}/logs120
POST /api/runs/inline, /api/runs/inline/validateINLINE_RUN_LIMITS.rate_per_min
POST /api/agents/{scope}/{name}/schedules10
GET /api/agents/{scope}/{name}/bundle30
POST /api/packages/import, /import-bundle, /import-github10
GET /api/packages/{scope}/{name}/files, .../files/content, .../{version}/download50
POST /api/uploads20
PUT /api/uploads/_content (by IP)60
GET /api/files, GET /api/files/{id} and its content120
DELETE /api/files/{id}, POST /api/files/{id}/keep60
GET /preview/files/{id} (by IP)120
POST /api/end-users, PATCH and DELETE /api/end-users/{id}60
GET /api/end-users, GET /api/end-users/{id}300
POST /api/webhooks, PATCH and DELETE /api/webhooks/{id}10
POST /api/webhooks/{id}/rotate5
GET /api/webhooks, /{id}, /{id}/deliveries300
POST /api/proxies/{id}/test, POST /api/models/test, POST /api/models/{id}/test, POST /api/model-provider-credentials/test and /{id}/test5
POST /api/model-provider-credentials/discover6
GET /api/models/openrouter10
POST /api/auth/bootstrap/redeem (by IP)5
Inbound MCP endpoint, per envelope120
Run-bound /internal/* routes200 per run, per route

The OAuth, OIDC and sign-in pages of the oidc module carry their own limits, set in the module. Authentication endpoints under /api/auth/* use Better Auth's own limiter, which these variables do not configure.

Response on limit

HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 15
RateLimit: limit=20, remaining=0, reset=15
RateLimit-Policy: 20;w=60

{
  "type": "https://docs.appstrate.dev/errors/rate-limited",
  "title": "Rate Limited",
  "status": 429,
  "detail": "Too many requests. Please try again shortly.",
  "code": "rate_limited",
  "retry_after": 15,
  "request_id": "req_..."
}

RateLimit and RateLimit-Policy are also sent on successful responses from the user-facing rate-limited routes (not the run-bound /internal/* routes), so clients can back off before they hit the limit. Respect Retry-After (seconds), and use exponential backoff for repeated 429 responses. The organization-wide run limits answer with the codes org_run_rate_limited and org_run_concurrency_exceeded shown above.

Specifics

  • Realtime (SSE). /api/realtime/* has no per-message limit. The stream fans events out as they arrive.
  • Outbound webhook deliveries run in a background worker, outside the HTTP pipeline, so these limits do not apply to them. Size your receiver for burst retries. See Webhooks.
  • Idempotency-Key. The rate limiter runs before the idempotency check, so a replayed request still consumes a point.

Monitoring

At LOG_LEVEL=debug the platform writes an access line per request. Alert on sustained 429 responses to catch abusive clients or limits set too low.

Bypasses

There is no built-in bypass: no admin exemption and no IP allowlist. For a tenant that needs different limits, raise PLATFORM_RUN_LIMITS for the instance, or apply a policy in the reverse proxy in front of Appstrate.

On this page