Skip to content

self-hosted build indexer cannot fetch project env vars with runtime environment key and aborts local deployment build #3258

Description

@flowq-C

Summary

In self-hosted Trigger.dev local deployments, the build image indexer (managed-index-controller) fetches project environment variables during the Docker build stage. In our setup this call fails and aborts the entire deployment build.

Proven behavior

  • Docker build fails inside:
    • managed-index-controller.mjs
  • Error:
    • Failed to fetch environment variables: 403 Forbidden
  • Direct API probe against:
    • GET /api/v1/projects/:projectRef/envvars
  • Response with production runtime key (tr_prod_...):
    • 401 Invalid or Missing API key

Why this is a Trigger.dev self-hosted bug

  • The build pipeline passes TRIGGER_SECRET_KEY into the image build.
  • The indexer depends on fetching env vars during the build.
  • The supplied key path does not authenticate successfully against the envvars endpoint in self-hosted mode.
  • Result: deployment image build hard-fails before publish.

Repro

  1. Self-host Trigger.dev v4 with local builds enabled
  2. Run trigger.dev deploy
  3. Let build reach the managed-index-controller step
  4. Observe failure on env var fetch

Expected

  • Build-stage env var fetch succeeds with the provided deployment/build auth
  • or the indexer degrades gracefully and continues with an empty env set

Actual

  • Build aborts with:
    • Failed to fetch environment variables: 403 Forbidden

Evidence

  • Build log excerpt:
    • Failed to index deployment
    • message: 'Failed to fetch environment variables: 403 Forbidden'
  • Direct endpoint probe:
    • GET /api/v1/projects/proj_ncohokyumnepswndlhei/envvars
    • with tr_prod_...
    • returns 401 Invalid or Missing API key

Suggested fix

  • Use the correct auth token type for build-time envvar resolution
  • or make managed-index-controller non-fatal when envvar fetch fails

Activity

  1. flowq-C commented on Mar 24, 2026

    @flowq-C
    Author

    Validated follow-up from the same live self-hosted environment:

    The envvar/auth failure was real, but there was an additional important detail:

    • once we routed build-time API traffic directly to the internal webapp endpoint in the host network, background worker creation succeeded during indexing
    • before that, the build-stage indexer path was returning 403 Forbidden on the background-worker creation call as well

    So the observed self-hosted behavior was:

    • envvar lookup path was not reliable with the deployment/build auth path
    • build-time API calls from inside the build container were also sensitive to how the API URL was resolved in self-hosted mode

    After bypassing the external route and using the internal webapp endpoint, we got successful worker creation during indexing and the deployment completed successfully.

    This means the issue is still valid, but the build-stage auth/networking path appears to be part of the same self-hosted failure cluster.

  2. flowq-C commented on Mar 24, 2026

    @flowq-C
    Author

    Complete working fix: envvars + background-workers in self-hosted local build indexer

    Following up with a full working solution validated on self-hosted v4.4.3 after multiple deployment cycles.


    The two API calls that break the indexer RUN step

    During --local-build, the indexer script runs inside a Docker BuildKit container via the generated Containerfile RUN step. From inside that container, it makes two API calls:

    Call 1: GET /api/v1/projects/{ref}/envvars

    Called with the runtime environment key (TRIGGER_SECRET_KEY). On self-hosted, this endpoint returns 401 for runtime keys (it requires a personal access token or a specific scope).

    The indexer treats this as fatal and aborts the build.

    Working fix: Intercept the fetch call and return a fake {variables: {}} response before it ever hits the API:

    if (url.includes('/envvars')) {
      return new Response(
        JSON.stringify({ variables: {} }),
        { status: 200, headers: { 'content-type': 'application/json' } }
      );
    }

    This is safe because env vars for the task are injected through syncEnvVars at build time, not at index time.


    Call 2: POST /api/v1/deployments/{id}/background-workers

    Called with TRIGGER_API_URL which is configured as http://localhost:8030. The auth key is correct (tr_prod_ key works for this endpoint). But localhost:8030 is unreachable from inside the BuildKit container.

    Root cause: BuildKit's buildx_buildkit_* container uses bridge networking even when RUN --network=host is specified for the build step. Inside that container, localhost refers to the container's own loopback — not the host. The Trigger.dev webapp is not listening there.

    Working fix: Rewrite the URL to the Docker bridge gateway IP before the fetch:

    if (url.includes('/background-workers')) {
      const rewritten = url.replace('localhost:8030', '172.24.1.1:8030');
      // reconstruct the Request with the new URL, preserving all headers including Authorization
      return originalFetch(rewritten, init);
    }

    Critical: Do NOT override the Authorization header on this call. The original tr_prod_ key is valid. Replacing it with MANAGED_WORKER_SECRET or any other key causes a 401.


    Complete injected fetch intercept (applied to the Containerfile RUN step)

    (async () => {
      const f = globalThis.fetch?.bind(globalThis);
      if (f) globalThis.fetch = async (i, n) => {
        const u = typeof i === 'string' ? i : i instanceof URL ? i.href : i?.url ?? '';
        if (u.includes('/envvars'))
          return new Response(JSON.stringify({ variables: {} }), {
            status: 200, headers: { 'content-type': 'application/json' }
          });
        if (u.includes('/background-workers')) {
          const ru = u.replace('localhost:8030', '172.24.1.1:8030');
          if (typeof i === 'string') return f(ru, n);
          if (i instanceof URL) return f(new URL(ru), n);
          return f(new Request(ru, i), n);
        }
        return f(i, n);
      };
      await import('${options.indexScript}');
    })()

    This is injected into the generated Containerfile RUN step via a patch to buildImage.js.


    Why this is an upstream problem (not just a self-hosted ops issue)

    Both failures are caused by assumptions that only hold in the Trigger.dev cloud environment:

    • The API URL is assumed to be reachable from inside BuildKit containers at localhost
    • The /envvars endpoint is assumed to accept runtime keys

    A proper upstream fix would:

    1. Make the indexer degrade gracefully when /envvars returns non-200 (treat as empty, log a warning)
    2. Make the TRIGGER_API_URL configurable for the build container network context, or default to the Docker gateway IP in self-hosted mode

    Happy to provide a PR for either or both if the contributor vouch path can be resolved.

  3. matt-aitken commented on Aug 29, 2026

    @matt-aitken
    Member

    Thank you for the detailed write-up. We looked into this.

    Authentication. The environment variables endpoint accepts the environment secret key (tr_prod_…). This is the key that the CLI passes into the build as TRIGGER_SECRET_KEY. It is not a key-type mismatch. The webapp does not return 403 from the envvars route or the background-workers route. A 403 on both calls that disappears when you bypass your external route indicates that a component in front of the webapp (a reverse proxy or an access layer) rejects or rewrites the request. The 401 in your probe matches the response for a missing or unrecognised Authorization header.

    Networking. The CLI already rewrites localhost to host.docker.internal and adds an --add-host entry for the build container. You can also run trigger.dev deploy --local-build --network host. This runs the build steps and the buildx builder on the host network, so the indexer can reach localhost:8030 directly.

    The patch. We cannot accept an empty environment variables response as a fallback. Indexing executes your task files with those variables, so an empty set would silently break projects that read process.env at import time. The error must stay fatal.

    Possible PR. If you still cannot reach the webapp from the builder, we would welcome a PR that:

    1. Adds an explicit build-time API URL override (for example --build-api-url or TRIGGER_BUILD_API_URL) that buildImage.ts uses instead of the automatic host.docker.internal rewrite.
    2. Documents the --network host option in the self-hosting docs.
    3. Includes the URL and the HTTP status in the indexer's envvars error message.

    Please see the contributing guide for the vouch process before you open the PR.

  4. SachinD6 commented on Sep 26, 2026

    @SachinD6

    Picking this up. I'll implement the three items @matt-aitken listed: the build-time API URL override in buildImage.ts, docs for --network host and the override, and the URL plus HTTP status in the indexer's environment-variable error.

    @flowq-C if you're already working on it, say so and I'll step back. Otherwise I'll open a PR once my vouch request is approved.

  5. SachinD6 commented on Sep 26, 2026

    @SachinD6

    Ready on my fork: SachinD6/trigger.dev@main...fix/local-build-api-url

    It adds --build-api-url / TRIGGER_BUILD_API_URL for the build container, documents that plus --network host, and puts the API URL and HTTP status in the indexer's envvar error. Vouch request: #4981, and I'll open the draft PR once that is approved.

    Verified with the packages/cli-v3 suite (61 tests) and a stub docker that records the generated build args. There is no Docker on this machine, so a real self-hosted local build is still unverified.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions