Files
willhaben-tracker/docs/phase-0/task-graceful-shutdown.md
T
hermes f540cbe7ef feat(phase-0): stability hardening — search path fix, auto-migrations, graceful shutdown, health endpoint
- Remove SET search_path from db.py and migration SQL (Supabase uses public schema)
- Add migrate.py with tracking table for forward-only SQL migrations
- Add entrypoint.sh: waits for DB, runs migrations, then starts app
- Copy 01-schema.sql + zz-seed.sql to worker/src/migrations/
- Add health.py: /health endpoint (200/503) with DB connectivity + scheduler staleness checks
  - /stats endpoint with keyword/ad/notification counts
- Rewrite main.py shutdown sequence: signal handler, 5s grace for scheduler, ordered cleanup
- Update Dockerfile: HEALTHCHECK directive, entrypoint, COPY migrations
- Update docker-compose.yml: stop_grace_period=15s, healthcheck config, env vars
- Add aiohttp>=3.9 to requirements.txt for health server
2026-07-04 11:30:45 -04:00

4.4 KiB

Task: Graceful shutdown fix

Description

The current main.py has a broken shutdown sequence. When the application receives a termination signal, calling app.shutdown() while the updater/scheduler are still active raises:

RuntimeError: This Application is still running!

This means:

  • DB connections may not be properly returned to the pool
  • In-flight Telegram messages may be lost
  • The container does not exit cleanly on SIGTERM, causing Kubernetes/Docker health issues during restarts

This task reorders the shutdown sequence and adds proper signal handling so the worker exits gracefully within 10 seconds.

Architecture

Current (broken):
  SIGTERM → cleanup() 
           ├─ app.shutdown()       ← FAILS: RuntimeError "still running"
           └─ close_pool()         ← never reached or runs in bad state

Target (fixed):
  SIGTERM → cleanup()
           ├─ scheduler_task.cancel()   ← stop scraping loop
           ├─ await asyncio.gather(scheduler_task)  ← let it finish current iteration
           ├─ app.updater.stop_polling()  ← stop Telegram long-poll
           ├─ app.shutdown()             ← safe now, nothing running
           └─ close_pool()               ← return all connections to pool

Key design decisions

  • Drain period: After signal received, allow up to 5 seconds for the current scrape cycle and pending notifications to complete.
  • Force kill fallback: If cleanup does not finish in 10 seconds total, force-exit via os._exit(0).
  • asyncpg pool close uses timeout (default 30s) — acceptable since we already cancelled tasks.

Implementation Details

Current main.py shutdown code (problematic):

# Simplified current pattern
try:
    # ... app.run() or equivalent
except KeyboardInterrupt:
    app.shutdown()   # ← RuntimeError here
    loop.run_until_complete(close_pool())

Target implementation:

Replace the signal handling block in main.py with:

import signal
import sys

_scheduler_task: asyncio.Task | None = None

async def cleanup(app: Application) -> None:
    """Gracefully shut down all components."""
    logger.info("Shutting down...")

    # 1. Cancel scheduler task (with grace period)
    if _scheduler_task and not _scheduler_task.done():
        _scheduler_task.cancel()
        try:
            await asyncio.wait_for(_scheduler_task, timeout=5.0)
        except (asyncio.CancelledError, asyncio.TimeoutError):
            pass

    # 2. Stop polling
    if app.updater and app.updater.running:
        app.updater.stop_polling()
    
    # 3. Now safe to shutdown the application
    await app.shutdown()

    # 4. Close DB pool
    await close_pool()
    
    logger.info("Shutdown complete.")


async def main():
    global _scheduler_task

    app = Application.builder().token(TOKEN).build()
    register_handlers(app)
    setup_global_commands(app)

    # Start polling
    app.updater.start_polling(drop_pending_updates=True)

    # Start scheduler
    _scheduler_task = asyncio.create_task(run_scheduler())

    # Set up signal handlers for graceful shutdown
    loop = asyncio.get_running_loop()
    for sig in (signal.SIGTERM, signal.SIGINT):
        loop.add_signal_handler(
            sig, lambda: asyncio.create_task(cleanup(app))
        )

    # Wait until the app is done or cleanup finishes
    try:
        await _scheduler_task  # runs forever until cancelled
    except asyncio.CancelledError:
        pass

    # Final cleanup (in case signal handler already ran)
    await cleanup(app)

What changes in main.py:

  1. Remove the broken try/except KeyboardInterrupt pattern
  2. Add the cleanup() async function with proper ordering
  3. Add signal handlers for SIGTERM and SIGINT via loop.add_signal_handler
  4. Track the scheduler task as a global so cleanup can reference it
  5. Remove any direct calls to app.shutdown() elsewhere

Acceptance Criteria

  • Sending docker stop willhaben-tracker-worker-1 results in clean exit (exit code 0) within 10 seconds
  • No RuntimeError: This Application is still running! appears in logs during shutdown
  • The asyncpg pool closes cleanly — no "connection closed unexpectedly" warnings after shutdown
  • In-flight notifications are allowed to complete (not dropped mid-send)
  • Restarting the container via docker restart works without errors
  • Logs show clear shutdown sequence: "Shutting down..." → "Shutdown complete."