Troubleshooting
Resolve common boot, run, network and configuration problems when self-hosting Appstrate.
Start with docker compose logs appstrate --tail 200 (or appstrate logs for an installer install) and curl http://localhost:3000/health. Boot errors name the variable at fault.
A stack written by appstrate install runs under a derived Compose project name. For it, use appstrate logs and appstrate status, or add --project-name "$(jq -r .projectName .appstrate/project.json)" to the docker compose commands on this page.
The container will not start
Five secrets are required and validated at boot:
| Variable | Rule |
|---|---|
BETTER_AUTH_SECRET | Required |
CONNECTION_ENCRYPTION_KEY | 32 bytes, base64-encoded |
UPLOAD_SIGNING_SECRET | At least 16 characters |
RUN_TOKEN_SECRET | At least 16 characters |
CONNECT_SESSION_SECRET | At least 16 characters |
The shipped Compose files also fail fast on POSTGRES_PASSWORD, on MINIO_ROOT_PASSWORD in the files that run MinIO (the root file, the Tier 3 template and examples/self-hosting/docker-compose.yml), and on POSTGRES_USER and MINIO_ROOT_USER in the root file. See Docker Compose for a generator.
Other cross-field checks that stop the boot:
TRUST_PROXY=falsebehind an HTTPSAPP_URLin production. SetTRUST_PROXYto your proxy depth, usually1.APP_URLis not an absolute origin. It must behttp://orhttps://with no path, query or fragment, and HTTPS in production (loopback excepted).USERCONTENT_URLon the same host asAPP_URL, or not HTTPS in production.- Image versions disagree. The platform,
PI_IMAGEandSIDECAR_IMAGEmust carry the same version. Pin them withAPPSTRATE_VERSION. See Upgrading. PLATFORM_RUN_LIMITS,INLINE_RUN_LIMITS,LLM_PROXY_LIMITSorCREDENTIAL_PROXY_LIMITScontains an unknown key or a bad value. The error names the key.RUN_ADAPTERnames a backend nobody registered. The error lists the registered ones.firecrackerneeds thefirecrackermodule inMODULES.- A model in
SYSTEM_PROVIDER_KEYSis outside its provider's offer, or a provider needs an API shape the platform does not serve. The error names the entry.
A variable in .env has no effect
Two common causes:
- Compose does not forward it. The container only receives the variables named under
appstrate.environmentin the Compose file. Add a bare- NAMEline. Even the current files do not list every variable (LLM_PROXY_LIMITSandFILE_MAX_BYTESare two that they do not), and a file from before 1.0.0-beta.65 lacks more. See Compose forwards only what it lists. - The name is wrong. Appstrate does not recognise renamed variables, and an unknown key is stripped without a warning. Compare the spelling with Environment Variables and the operator notes of the release.
Database connection refused
- Check that
DATABASE_URLpoints to a reachable PostgreSQL, and that the database container is healthy:docker compose ps. - On the Docker tiers the database sits on an internal network and publishes no port. Test from inside:
docker compose exec postgres pg_isready(serviceappstrate-postgresin the root file). - If
DATABASE_URLis unset, Appstrate uses PGlite inPGLITE_DATA_DIR. A single PGlite directory must be opened by one process only. - A password with
@,/or:breaks the connection string Compose builds. Use hex or URL-safe passwords.
Agents fail to run
- Wrong backend.
RUN_ADAPTERdefaults toprocess(host subprocesses, no isolation). The shipped Compose files setdocker. - Docker socket. The platform container must reach
/var/run/docker.sock. Check the group mapping (DOCKER_GID) and the socket's permissions. See Isolation and Security. - Runtime images missing. Check
docker images | grep appstrate. Pull the same version ofappstrate-pi,appstrate-sidecarand the fiveappstrate-mcp-runner-*images fromghcr.io/appstrate. A missing MCP runner image fails the integration that needs it, because the sidecar never pulls images. - Local integration refused under
process. Asource.kind: "local"integration (for example@appstrate/github-git) does not start withRUN_ADAPTER=process. SetINTEGRATION_RUNTIME_ADAPTER=dockeror useRUN_ADAPTER=docker. - "Run never started executing". The runner posted no event within
RUN_BOOT_DEADLINE_SECONDS(default 300). Look for a slow image pull, a missing image or a container that exits at boot. - Run failed as stalled. No heartbeat for
RUN_STALL_THRESHOLD_SECONDS(default 60). Check the run's container logs and the host's resources. - "Server restarted while run was in progress". A restart finalizes in-flight runs as failed. Retry them.
- Model on a private endpoint fails. Add its host to
EGRESS_ALLOW_INTERNAL_HOSTS. The host must resolve from the API process, wherelocalhostis the API's own loopback.
Outbound calls refused
| Symptom | Meaning |
|---|---|
403 URL targets a blocked network range, and Target refused (SSRF) in the sidecar log | The target is private, loopback or link-local. To allow it, name the host literally in the integration's authorized_uris and list it in EGRESS_ALLOW_INTERNAL_HOSTS |
403 with unauthorized_target (platform credential proxy), or a sidecar refusal as unauthorized | The integration declares no authorized_uris, or the URL is outside them. An empty list authorizes nothing |
502 Target host could not be resolved | The sidecar could not resolve the host. It needs working DNS, even when it sends through PROXY_URL |
Docker network pool exhausted
Docker create network appstrate-exec-... failed: 400
{"message":"all predefined address pools have been fully subnetted"}Each run uses one network, and Docker's default address pool holds about 31 on a stock host. Reclaim unused networks with docker network prune. For a lasting fix, carve smaller subnets in /etc/docker/daemon.json (Docker Desktop: Settings, Docker Engine) and restart the daemon:
{
"default-address-pools": [
{ "base": "172.20.0.0/16", "size": 24 },
{ "base": "10.200.0.0/16", "size": 24 }
]
}Each /16 base then yields about 256 networks. Appstrate also reclaims orphaned appstrate-exec-* networks after a failure and retries once.
400 on API calls: missing organization or space
Requests that belong to an organization need an X-Org-Id header, and space-scoped routes (agents, runs, schedules, end-users, API keys, notifications, packages, integrations, files and uploads) also need X-Space-Id. Without it you get 400. An API key carries its own organization and space, so it needs neither header. With an API key, X-Org-Id is ignored (the key is bound to its organization) and only X-Space-Id is checked: a value that contradicts the key's space returns 403.
Sign-in and first-owner problems
- Sign-up is closed (
signup_disabled). The instance runs in closed mode. Ask for an invitation, or see AUTH_MODES.md. /claimanswers 410. The bootstrap token is only redeemable while the instance has no organization. If you setAUTH_BOOTSTRAP_TOKENand it does nothing, check that your Compose file forwards the variable to the container (the files shipped before 1.0.0-beta.65 do not).- The owner address cannot register. An address named in
AUTH_BOOTSTRAP_OWNER_EMAILorAUTH_PLATFORM_ADMIN_EMAILSis not created by the plain sign-up form. Claim it at/claimwhile the instance has no organization (withAUTH_BOOTSTRAP_OWNER_EMAILset, the claim accepts that address only), or use a magic link (SMTP required) or a Google or GitHub sign-in whose provider asserts the address as verified. The server log states the reason. - A magic link or a Google or GitHub sign-in fails right after an upgrade. Better Auth 1.7.7, shipped in 1.0.0-beta.65, changed how these in-flight values are stored. A link mailed before the restart is refused and a social sign-in started before it must be started again. Ask for a new link. See Upgrading.
- Redeem rate-limited.
/api/auth/bootstrap/redeemallows 5 attempts per minute per IP. Behind a proxy, setTRUST_PROXY.
OAuth callback fails
APP_URL must match the public URL exactly, including the scheme and no stray port:
# correct
APP_URL=https://appstrate.example.com
# wrong: missing scheme
APP_URL=appstrate.example.com
# wrong: internal port when a proxy terminates TLS on 443
APP_URL=http://appstrate.example.com:3000Register the same redirect URI at the provider (Google, GitHub or an OIDC client).
Uploads and live streams misbehave behind a proxy
413that Appstrate never logs. The proxy's body limit is lower than your upload. Raise it to at least 100 MiB (nginx:client_max_body_size 100m).- Chat or run logs arrive in one batch at the end. The proxy buffers or compresses
text/event-stream. Disable that for the Appstrate location. - Everyone is rate-limited together.
TRUST_PROXYis not set to your proxy depth, so every caller shares the proxy's IP. See Rate Limits.
Redis fallbacks
Without REDIS_URL (Tier 0 and 1), the queue, pub/sub, cache and rate limiter are in-process:
- The scheduler's cron evaluator polls every 30 seconds, and queued jobs are lost on restart.
- Rate-limit counters reset on restart.
- Several instances do not share state. Set
REDIS_URLfor more than one instance.
With Redis, keep its volume if you care about scheduled runs and queued webhook deliveries.
MinIO crash-loops
FATAL Unable to initialize backend: Unable to write to the backend means the miniodata volume is not owned by uid 65532. Re-own it once with the stack stopped. The exact commands, with a snapshot step first, are in the self-hosting README.
Stored objects that will not delete
Deleting a file, a run workspace, a space or an organization removes its database rows at once and queues the stored objects in a deletion outbox. A background worker purges them. Every STORAGE_DELETION_WORKER_INTERVAL_MS (default 60000, see Environment Variables) it claims a batch of due jobs and deletes the objects. A failed job is retried later, with a delay that doubles after each attempt (about one minute at first, capped at 6 hours, with up to 10% jitter). A job is never abandoned: deletion is retried for as long as it takes.
After 8 attempts a job that is still pending is called a dead letter. The threshold only makes the job visible, it does not stop the retries. Dead letters are what to look at when disk or bucket usage does not go down after deletions. With the @appstrate/module-observability module, the gauges appstrate.storage_deletion.backlog, appstrate.storage_deletion.oldest_pending_age_seconds and appstrate.storage_deletion.dead_letters report the same state.
List the jobs with the platform-admin API:
curl "https://your-instance/api/admin/storage-deletion-jobs?status=dead" \
-b cookies.txt| Query parameter | Values |
|---|---|
status | pending (default, every job not yet completed), dead (pending with 8 or more attempts) or completed |
limit | 1 to 200, default 50 |
startingAfter | The id of the last job of the previous page. Follow the Link header with rel="next" instead of building it by hand |
The answer is { "object": "list", "data": [...], "hasMore": false }, newest first. Each job has id, bucket, storage_key (the key inside the bucket), reason (why the object is purged, for example file_deleted, space_deleted or run_workspace_deleted), attempts, next_attempt_at, completed_at, last_error (the message of the last failed attempt) and createdAt. The list covers every organization of the instance, and a storage key contains a space id and a file name.
To retry a job without waiting for its next scheduled attempt:
curl -X POST "https://your-instance/api/admin/storage-deletion-jobs/$JOB_ID/retry" \
-b cookies.txtIt answers { "id": "...", "retried": true } and makes the job due immediately, so the next worker pass picks it up. A job that is already completed, or an unknown id, answers 404. Retrying does not reset the attempt count, so a dead letter stays one until it succeeds. The two calls are limited to 60 and 30 per minute. Fix the cause shown in last_error first (the storage backend, its credentials or its permissions), or the retry fails the same way.
Both calls answer 403 Platform admin access required unless all of these hold:
- You are signed in with a dashboard session cookie. An API key is refused, and so is an OIDC token, whatever its scopes.
- The session belongs to the
platformaudience. A person who signed in as an end-user of a space is refused. - Your email address is in
AUTH_PLATFORM_ADMIN_EMAILS, a comma-separated list compared without regard to case. With the variable empty, nobody qualifies. The account for a listed address is only created with proof that you own it, see AUTH_MODES.md.
These two routes need no X-Org-Id, so an operator who belongs to no organization can call them. Before 1.0.0-beta.65 they answered 400 without that header.
Rate limit hits (429)
Check Retry-After and the RateLimit headers. Per-endpoint limits and the organization-wide run limits are in Rate Limits. An organization at its concurrent run cap gets org_run_concurrency_exceeded.
Port 3000 is already in use
Change the host side of the port mapping, or set PORT in .env (the Compose files read ${PORT:-3000}). appstrate install --port 3100 sets it for you, and with --yes the installer picks the next free port by itself.
minisign: command not found during install
curl -fsSL https://get.appstrate.dev | bash verifies the CLI binary with minisign. On a TTY it offers to install minisign through your package manager, and with --yes, in CI or without a TTY it installs it automatically when it can. If that fails, install it yourself (brew install minisign, apt install minisign, apk add minisign) and re-run. Setting APPSTRATE_NO_INSTALL_MINISIGN=1 disables the automatic install.
Verbose logging
LOG_LEVEL=debugDebug logs include one access line per request (method, route pattern, status, duration and Request-Id). Return to info once the incident is understood.
Orphaned containers
Containers created for runs carry the label appstrate.managed=true, and the platform reconciles them at startup. If strays remain after a crash:
docker ps -a --filter "label=appstrate.managed=true" -q | xargs docker rm -fPer-run networks (appstrate-exec-*) are cleaned up at boot. The shared appstrate-egress network is never removed, by design.