5.2 KiB
5.2 KiB
Task: httpx singleton with connection pool
Description
The current scraper.fetch_ads() creates a new httpx.AsyncClient on every call:
async def fetch_ads(keyword: str):
async with httpx.AsyncClient(timeout=30.0) as client:
resp = await client.get(_API_URL, ...)
This means each scrape cycle incurs the full cost of TCP handshake + TLS negotiation (≈100-300ms per call on a cold connection). For keywords scraped every 5 minutes with multiple pages, this overhead adds up to seconds of unnecessary latency per cycle.
This task replaces the per-call client with a module-level singleton that reuses connections via keepalive.
Architecture
Current:
Cycle 1: create AsyncClient → fetch → close → ~300ms overhead
Cycle 2: create AsyncClient → fetch → close → ~300ms overhead
Cycle N: ... (repeated forever)
Target:
Module load: create AsyncClient (singleton, keepalive pool)
Cycle 1: use client → fetch → ~50ms (warm connection)
Cycle 2: use client → fetch → ~50ms (warm connection)
Cycle N: ...
Shutdown: close client gracefully
Key design decisions
- Module-level singleton (
_client = None, lazy init). Simpler than dependency injection and works with the existing async context. - Keepalive connections: Default
max_keepalive_connections=5handles concurrent keyword scrapes efficiently. - Client recreation on error: If the client is closed or encounters a fatal transport error, it's recreated on the next call. This prevents stale connection issues.
Implementation Details
1. Add singleton getter to scraper.py
import os
_client: httpx.AsyncClient | None = None
async def get_client() -> httpx.AsyncClient:
"""Return a shared AsyncClient with keepalive connection pool."""
global _client
if _client is None or _client.is_closed:
max_conns = int(os.getenv("HTTP_MAX_CONNECTIONS", "10"))
max_keepalive = int(os.getenv("HTTP_KEEPALIVE_CONNECTIONS", "5"))
_client = httpx.AsyncClient(
timeout=float(os.getenv("HTTP_TIMEOUT_S", "30.0")),
limits=httpx.Limits(
max_connections=max_conns,
max_keepalive_connections=max_keepalive,
keepalive_expiry=60, # seconds
),
)
logger.info(
"Created httpx client: max_conns=%d, keepalive=%d",
max_conns, max_keepalive,
)
return _client
async def close_client() -> None:
"""Close the shared AsyncClient. Call during shutdown."""
global _client
if _client and not _client.is_closed:
await _client.aclose()
logger.info("Closed httpx client")
_client = None
2. Update fetch_ads() to use the singleton
Replace:
async def fetch_ads(keyword: str):
params = {...}
async with httpx.AsyncClient(timeout=30.0) as client:
for attempt in range(1, 4):
try:
resp = await client.get(_API_URL, headers=_HEADERS, params=params)
With:
async def fetch_ads(keyword: str):
params = {...}
client = await get_client()
for attempt in range(1, 4):
try:
resp = await client.get(_API_URL, headers=_HEADERS, params=params)
resp.raise_for_status()
data = resp.json()
break
except httpx.ConnectError as exc:
# Transport error — recreate client on next attempt
logger.warning("Transport error on attempt %d: %s", attempt, exc)
await close_client() # force recreation
if attempt < 3:
await asyncio.sleep(2 ** attempt)
continue
raise
except Exception as exc:
logger.warning("fetch_ads attempt %d failed: %s", attempt, exc)
if attempt < 3:
await asyncio.sleep(2 ** attempt)
continue
raise
# ... rest unchanged (extract ads_raw, total_hits)
3. Call close_client() during shutdown in main.py
Add to the cleanup function (from Phase 0 task-graceful-shutdown):
async def cleanup(app: Application) -> None:
logger.info("Shutting down...")
# ... existing cleanup steps ...
# Close HTTP client
from scraper import close_client
await close_client()
# ... rest of cleanup (close DB pool, etc.)
4. Update .env.example with new config options
# HTTP Client Configuration
HTTP_MAX_CONNECTIONS=10 # Max concurrent connections to willhaben API
HTTP_KEEPALIVE_CONNECTIONS=5 # Connections kept alive in the pool
HTTP_TIMEOUT_S=30.0 # Request timeout in seconds
Acceptance Criteria
- Only one
httpx.AsyncClientis created per process lifetime (logged once at startup) - Subsequent calls to
fetch_ads()reuse the existing client (no "Created httpx client" log) - After calling
close_client(), a new call toget_client()creates a fresh client - Connection keepalive reduces latency for sequential API calls (verifiable via timing in logs)
- Fatal transport errors trigger client recreation without crashing the scheduler
- The client is properly closed during graceful shutdown (no resource warnings)