Email Delivery: Knowing an Email Was Sent (Architecture)
This document describes how the portal tells a maintainer, in one place,
whether each email it was asked to send actually went out, what happens to
the ones that did not, and how a silent worker stops being silent. It is the
reference for any work that touches common/send_emails.py,
common/tasks.py, the Celery settings in portal/settings.py, the
EmailDispatch model, or the Maintenance > Delivery page.
Status: partly built. Sections marked build say what lands and where.
Built on 1 October 2026: the audit and retrigger page (Maintenance >
Invitations, speakers/delivery.py), the LOGGING setting, and log lines
for a queued task, the invitation task and a retrigger. Built on 2 October
2026: the Workers panel at the top of Maintenance > Invitations
(common/workers.py, and the portal_code_version control command in
common/tasks.py), which lists every worker attached to the broker beside
the web process's email-code fingerprint, after a stale worker pod left
over from an earlier release took most of the tasks and sent email without a
record. Everything else below is still design.
Why
On 1 October 2026 the presenter page showed an invitation as sent, and Maintenance > Emails had no record of it. Other invitations to the same session, sent in the same minute, did have records, and at least four more speakers, invited at different times, were in the same state. Nothing in Sentry. A synchronous run of the same task for one of them succeeded, so the invitation, the address, the template and the mail provider were all fine. The email was lost between the web request and the worker.
Three properties of the current code made that possible and invisible:
- "Sent" is a claim made by the web process.
Invitation.issue_token()stampssent_atwhen the task is queued. The presenter page, the session page and the checklist rules all read that stamp. Nothing writes anything when the worker actually delivers, so the stamp and the email can disagree for good. - A lost task raises nothing. The email tasks are acknowledged when the
worker receives them (Celery's default). If the worker is killed or
restarted while it holds one, the task is gone: no exception, no
FAILEDrecord, no Sentry event.enqueue()also swallows a broker outage, logging it and letting the request succeed. - The only record is written by the thing that failed.
SentEmailis written by the worker, after rendering and delivery. Absence of a row is the only evidence, and nobody looks for an absence.
The fix is not a better error handler. It is to record the intent to send in the same transaction that makes the claim, to have the worker close that record, and to have something that notices when it is not closed.
Goals
- Every email the portal is asked to send has a row saying what was asked, for whom, and where it stands: queued, sending, sent, failed or stalled.
- The presenter page says "Queued", "Sent" or "Failed", from that row, never from the web process's own stamp.
- A task lost in a restart is delivered again without anyone acting.
- A task that cannot be delivered, or that stays unfinished, is visible on one page and in Sentry, with enough in the log line to find the worker's side of the story.
- A maintainer can resend from that page, singly or in bulk.
- Silence is detectable: if a worker stops consuming, or stops reporting to Sentry, something says so.
Non-goals
- Delivery receipts, opens and bounces. "Sent" means the mail provider accepted the message. What the recipient's server did with it needs the provider's webhooks, which is a separate piece of work (see Later).
- Replaying a stored body. The record holds the rendered Markdown with the
sign-in link withheld (design
speaker-portal.mdsection 13.1), so a resend always runs the original action again and builds a fresh email. - Changing who receives what. No recipient is ever added, and a resend goes to the address on the dispatch, not a new one.
- Account emails (sign-in codes, password reset). They go through the account adapter and stay out of every record, as today.
The model
EmailDispatch (common app)
One row per request to send an email. It holds no message body, so it is not
personal data in the way SentEmail is, and it is not shown to recipients.
| Field | Meaning |
|---|---|
id |
UUID. Also the Celery task_id, so a worker log line carries it. |
kind |
The template or task family, e.g. speakers/invitation. |
task |
Dotted task name, so the reconciler can enqueue it again. |
args |
JSON list of the task's arguments (ids only). |
to |
The address it was addressed to when requested. |
conference, presenter, invitation |
Nullable links, for filtering and for the presenter page. SET_NULL. |
requested_by, requested_at |
Who asked, and when. |
status |
QUEUED, SENDING, SENT, FAILED, STALLED. |
attempts |
How many times a worker started it. |
started_at, finished_at |
Of the latest attempt. |
worker |
Celery hostname of the latest attempt. |
error |
Exception class and message, truncated, for FAILED. |
sent_email |
The SentEmail it produced, nullable. |
States and moves:
QUEUED --worker starts--> SENDING --delivered--> SENT
| |
| +--raises, retryable--> QUEUED (attempts + 1)
| +--raises, permanent---> FAILED
+--no worker within the stall window--> STALLED --reconciler re-queues--> QUEUED
(after the attempt cap) --> FAILED
SENT and FAILED are final until a person resends, which creates a new
dispatch for the same target and leaves the old one as history.
Why not put this on SentEmail
SentEmail answers "what did the portal tell this person", is read by the
recipient on their own page, and is deliberately written only when a send
happened. A QUEUED row with no body would appear in the speaker's trail
for an email they never got. Two tables, two questions: the dispatch is
what was asked, the sent email is what was sent, and sent_email joins
them.
The request path
build
enqueue_email(task, *args, kind, to, **links) in common/tasks.py
replaces enqueue() for email tasks:
- Inside the caller's transaction, create the
EmailDispatchasQUEUED. If the action rolls back, so does the row. If it commits, the row exists whatever happens next. - On commit, send the task with
task_id=str(dispatch.id). - If sending raises (the broker is down), set
FAILEDwith the error and log at error level. The row exists, so the page shows it, and the request still succeeds, asenqueue()does today.
send_invitation then stops treating sent_at as the delivery fact. It
keeps sent_at as "an organizer asked" (the token and its expiry still hang
off it), and the page reads the latest dispatch for the label:
| Dispatch | Label on the presenter page |
|---|---|
QUEUED |
Queued, not sent yet |
SENDING |
Sending |
SENT |
Sent, and when |
FAILED |
Not sent: the error, and a Resend button |
STALLED |
Not sent: no worker picked it up, and a Resend button |
Existing invitations have no dispatch. The page falls back to today's wording for them, until the backfill below has run.
The worker path
build
Email tasks share one decorator, @email_task, instead of a bare
@shared_task:
bind=True,acks_late=True,reject_on_worker_lost=True. A task a killed worker was holding goes back on the queue. This is already howbuild_export_zip_taskis declared; the email tasks were never given it.autoretry_forthe transient errors (smtplib.SMTPExceptionsubclasses that are not recipient refusals,OSError), with backoff and jitter,max_retries=5. Permanent errors (SMTPRecipientsRefused, an address Django cannot encode) go straight toFAILED.- The first thing the task does is load its dispatch and return if it is
already
SENT. With late acknowledgement a task can be delivered twice (a redelivery after the visibility timeout on Redis), and this is what keeps a presenter from getting two invitations. The dispatch is read withselect_for_updateso two workers cannot both pass the check. - It then sets
SENDING,started_at,attempts + 1andworker, runs the existing body, and on success setsSENT,finished_atandsent_email. The existingSentEmailwrite is unchanged.
Tasks that have no dispatch (queued by code that has not adopted
enqueue_email, or by a deploy mid-flight) still run: the decorator creates
the row at start if it is missing.
The reconciler
build
A beat task, reconcile_email_dispatches_task, runs every five minutes on
the default queue, seeded as a periodic task by the migration like the
prune job.
QUEUEDfor longer thanEMAIL_STALL_MINUTES(default 10), orSENDINGfor longer than the task's time limit: setSTALLED, log at warning, and enqueue it again under the same id withattemptsunchanged.- A dispatch that has been stalled and re-queued
EMAIL_STALL_RETRIEStimes (default 2) becomesFAILEDwith "no worker completed this", and one error is logged for it. That log line is what reaches Sentry. - It never touches
SENTorFAILED.
This is the part that would have caught the 1 October invitations within
ten minutes: the task sat QUEUED, then STALLED, then was delivered by
the next pass.
Logs
Partly built. LOGGING is in place (below), enqueue logs the id of the
task it queued, the invitation task logs when it starts, when it ends and
why it failed (never the address), and a retrigger logs who sent what again.
The dispatch transition lines in the shape below wait for EmailDispatch.
Until then the project had no LOGGING setting. Module loggers (speakers.tasks,
common.tasks) have no handler in the web process, so anything below
warning never reaches the platform's logs there. The Celery worker sets up
its own root handler, so its INFO lines do appear, which is why the worker
log is the place to look.
Add a minimal LOGGING to portal/settings.py: one console handler to
stdout, root at WARNING, and portal.email, speakers, common at
INFO. Every dispatch transition logs one line on the portal.email
logger in a fixed shape:
email.dispatch id=8b1f... kind=speakers/invitation status=SENT attempt=1 worker=celery@abc to_hash=ab12 target=invitation:482
The id is the Celery task id, so the worker's own received and
succeeded lines carry it too, and one search finds the whole story. The
address is logged as a short hash, not in full.
Sentry
build
What the code already does
sentry_sdk.init runs in portal/settings.py, behind SENTRY_SDK_DSN.
Every process imports those settings (the Celery app reads them), so web,
worker, worker-media and worker-beat initialise Sentry from the same
environment variable, and the SDK enables its Celery integration on its own
(it is in the SDK's auto-enabled list, and sentry-sdk[django] is
installed). A task that raises in a worker is reported without extra
configuration.
So the platform sharing one config is expected and not the problem. Two other things explain "alerts only from web":
- A lost task raises nothing. There is no event to alert on. This is the 1 October failure, and no Sentry setting changes it. The dispatch and the reconciler are what turn it into an error.
- Alert rules can be scoped to web (an issue alert filtered by transaction, URL or a tag only web events carry). That is configured in Sentry, not in the repo, and is the first thing to check there.
What to add
- Tag every event with its process. At init, set
processtoweb,worker,worker-mediaorbeat, read fromsys.argv(gunicorn, orcelery ... worker -Q media). Events from a worker are then findable and alert rules can name them. Also setenvironmentfrom an env var. - Report each worker start. On Celery's
worker_ready, log at INFO "worker ready: sentry=queues=<...>" and send one Sentry message event tagged with the process. A worker that never says this is not reporting, and a worker that says it every few minutes is crash-looping (an out-of-memory kill looks exactly like this). - Use Sentry Crons for the silent case. Set
monitor_beat_tasks=TrueonCeleryIntegration, so a beat task that does not run, or does not finish, raises a missed check-in. The reconciler is the one that matters: if it stops running, nothing else here works. Verify at build time that this works withdjango_celery_beat's database scheduler on SDK 2.43; if it does not, wrap the reconciler body withsentry_sdk.monitor(...)instead. - Log, do not swallow.
enqueue_emaillogs a failed send at error level. The default logging integration turns that into an event.
A test that proves the path
A management command, check_workers, and a button on the Delivery page,
queue a task on each queue (default, media) that records a heartbeat and
can be told to raise on purpose. The heartbeat row (WorkerHeartbeat:
queue, hostname, seen_at, sentry enabled) is what the Delivery page shows.
Raising on purpose is how a maintainer confirms, once, that an error from
each worker really arrives in Sentry tagged with its process.
Maintenance > Delivery
build. The worker list below is already on Maintenance > Invitations
(common.workers.inspect_workers); this page would take it over.
A page beside Accounts, Emails and Exports, gated on is_maintainer like
them (MaintainerRequiredMixin). Three parts:
- Workers. One line per queue: last heartbeat, hostname, whether
Sentry was on (the live list of attached workers, their uptime and email
code is already built). Red if no heartbeat in ten minutes, with the text of what
that usually means ("the worker is not running or not consuming this
queue"). This is the same idea as
fail_stale_jobs, which already tells a maintainer thatworker-mediais at zero replicas. - Dispatches. Newest first, filterable by status, kind, edition and presenter, searchable by address. Stalled and failed are on top by default. A row shows the error, attempts, worker, timestamps and a link to the sent email when there is one. The Celery task id is shown and copyable, because it is what to search the logs for.
- Resend. A POST button per row. A bulk "resend all" is left out until the one-at-a-time page has been used for a while.
What resend does
Resend never replays a stored body. It runs the original action again for the target, through the same service function an organizer would use:
- Invitation:
send_invitation(invitation, actor=request.user). This issues a fresh token, so the earlier link stops working, and the page says so before the button is pressed. An accepted invitation is refused, as today. - Other kinds: enqueue the same task again with the same arguments, as a new dispatch. These tasks are idempotent by construction (they read current state), so a second run is safe.
Every resend is recorded in the activity log with the maintainer as actor, and the new dispatch links to the one it replaces. A resend goes to the address on the target now, which is the same address an organizer's resend uses.
The earlier design note on the Emails page ("there is no resend, which is a different decision with its own consent questions") stands for that page, which is the recipient-facing trail. Resend lives here, on the page about the portal's own machinery, and resends only emails the portal itself decided to send.
Backfill
build
Dispatches only exist for sends after this ships. The speakers affected now
have none. A management command, backfill_email_dispatches, finds every
invitation with sent_at set and no SentEmail for it (matched on the
template, context_digest.invitation and a time at or after sent_at),
creates a STALLED dispatch for each, and with --resend queues them.
--dry-run is the default and prints the list. This is also the one-off
audit of how many speakers are affected, and it can run before anything else
here is built (see Rollout).
Rollout
Each step is one pull request, and only step 3 has a migration.
- Find and fix the known cases. Built, as a page rather than a
command: Maintenance > Invitations lists every invitation of the active
edition marked sent with no successful
SentEmailfor that send (a failed record is shown with its error), and sends them again one at a time, each with its own button (there is deliberately no bulk send: it emails a real person and ends their current link). It comparesInvitation.sent_atwith the records, so it needs no new table. It lists nothing sent before thecommonrecord table existed (the time migration0001_sent_emailwas applied, not the oldest surviving row, which the nightly prune moves forward and which a deployment whose worker was down at first does not have), skips invitations that were opened, accepted, declined or cancelled, and disables a row sent in the last five minutes, which may still be queued. The invitation's row is locked while it is sent, so two clicks cannot both send. A retrigger goes throughsend_invitation(a fresh link and expiry, as an organizer's resend does), writes aninvitation.retriggeredactivity entry with the maintainer as actor, and logs a line. The page lists those entries. It cannot see an email the provider accepted before the process died and the row was saved, so sending again there is a duplicate, never a miss. - Stop losing tasks.
@email_taskwithacks_late,reject_on_worker_lostand retries,LOGGING, the Sentryprocesstag and theworker_readymessage, and a pinned--concurrencyforworkerin theProcfile. No migration. This alone ends the silent loss of a task a restarting worker held. - Make it visible.
EmailDispatch,enqueue_email, the presenter-page labels, the reconciler,WorkerHeartbeat, the Delivery page and resend. One migration for the app, with the periodic task seeded in it. - Adopt elsewhere. Volunteer, sponsorship and the reminder digests move
from
enqueue()toenqueue_email(), one app per pull request. Until then they behave as today.
Settings
| Setting | Default | Meaning |
|---|---|---|
EMAIL_STALL_MINUTES |
10 | QUEUED this long counts as stalled. |
EMAIL_STALL_RETRIES |
2 | Re-queues before a stalled dispatch is FAILED. |
EMAIL_DISPATCH_RETENTION_DAYS |
90 | SENT dispatches older than this are pruned with the nightly job. FAILED ones are kept as long as the email records. |
WORKER_HEARTBEAT_MINUTES |
5 | How often beat queues a heartbeat per queue. |
Checking by hand today
For the maintainer who needs the answer before any of this exists:
- Which invitations have no email? In the Django shell, for each
invitation with
sent_atset, look for aSentEmailwithtemplate="emails/speakers/invitation.md",context_digest__invitation=<pk>andsent_at__gte=<invitation.sent_at>. None means no email was recorded. - Why? In the worker log around the invitation's
sent_at, search forTask speakers.tasks.send_invitation_email_task[: areceivedwith nosucceededmeans the worker died holding it; search forWorkerLostErrororWorker exited prematurely, which are how a killed child process appears. Noreceivedat all means it never arrived: search the web log forFailed to enqueue Celery task. - Is the worker restarting? The platform's events for the worker process: restarts and out-of-memory kills at the same times.
- Which process's logs? Cabotage runs each
Procfileline as its own deployment (docs/developer/deployment.md, "Processes"), soworker,worker-beatandworker-mediaeach have their own logs and replica count, separate fromweb. The web log holds only gunicorn's request and error lines, since the project sets noLOGGINGand the app's own INFO messages are dropped there. Open theworkerprocess, notweb, and confirm it is scaled to at least one replica. - Is the worker too big for its container?
celery workerdefaults to one child process per CPU it can see, and in a container that is often the node's core count, not the container's share. Each child imports Django, so a small memory limit is exceeded and the kernel kills a child holding a task. With the default early acknowledgement that task is lost. TheProcfileshould pin--concurrencyforworker, as it already does forworker-media.
Later
- The mail provider's webhooks (delivered, bounced, complained) feeding a
DELIVERED/BOUNCEDstate on the dispatch, so "Sent" can become "Delivered". - Telling the organizers, not only the maintainer, when a speaker's invitation could not be sent, on the session page where they are already working.
- Rate-limiting a bulk resend against the provider's sending limits.