Skip to content

Limit pending poll tasks on the shared queue #653

Description

@dahlia

Background

#636 moved import, cleanup, remote replies scraping, and poll notifications onto a dedicated PostgreSQL-backed Fedify task queue. Each worker process runs at most four application tasks, with at most two replies scrapes. Federation delivery uses a separate queue, but both queues share database and process resources.

#639, implemented in #649, replaced the poll notification polling loop with delayed tasks. Poll creation, successful votes, and remote expiry changes schedule wakeups. Handlers reload the current expiry and recipients; if a message arrives before expiry, it schedules another wakeup. A bounded recovery scan finds expired polls with missing notifications after failed dispatch or downtime.

Queue growth

The concurrency cap limits running handlers, not stored messages. Recovery pauses when the shared queue has 200 ready messages, but direct poll scheduling does not check queue depth. Delayed messages do not count toward this threshold.

For example, extending a poll's expiry schedules a new wakeup without removing the old one. When the old message runs, it sees the later expiry and schedules a replacement. Repeated changes can leave several delayed messages for one poll, each scheduling another when it runs before expiry. Database idempotency prevents duplicate notifications, but does not bound queue growth.

These messages may increase PostgreSQL storage use and delay imports, cleanup, and replies scraping. No resource-exhaustion incident has been reproduced.

Measure how the queue grows under repeated scheduling and set a limit on pending poll work. Possible approaches include tracking pending wakeups per poll or limiting enqueues and letting recovery find polls whose wakeups were not queued. A static TTL deduplication key alone can suppress a needed wakeup after expiry changes or failed dispatch.

Acceptance criteria

  • Stress tests cover repeated expiry changes, votes, and concurrent workers, recording queue depth and whether imports, cleanup, and replies scraping continue to make progress.
  • Pending poll work stays within a documented bound without losing notifications when expiry moves earlier or later.
  • Poll work remains recoverable after missed enqueues, exhausted retries, or worker restarts. Duplicate delivery does not change notification counts.
  • Document the chosen limits and how operators can observe queue pressure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Fields

Priority

None yet

Effort

None yet

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions