Stream writes enqueue rows in webhook_outbox inside the same database transaction as the stream update. The live dispatcher in src/webhooks/service.ts polls that table and sends each event to the configured consumer endpoint.
Required configuration:
WEBHOOK_URL: HTTPS endpoint that receives webhookPOSTrequests.WEBHOOK_SECRET: HMAC signing secret used forx-fluxora-signature.WEBHOOK_POLL_INTERVAL_MS: polling interval in milliseconds. Defaults to10000.WEBHOOK_BATCH_SIZE: rows claimed per poll. Defaults to10.WEBHOOK_RETRY_RPS: maximum outbound retry attempts per second per consumer URL. Defaults to10. Set lower (e.g.2) for consumers known to be slow or fragile.WEBHOOK_CIRCUIT_BREAKER_THRESHOLD: consecutive retryable failures before the circuit opens. Defaults to0(disabled). Set e.g.10to enable cross-instance protection.WEBHOOK_CIRCUIT_BREAKER_RESET_MS: how long the circuit stays open before a single half-open probe. Defaults to300000(5 minutes).WEBHOOK_DNS_TIMEOUT_MS: DNS lookup resolution timeout in milliseconds (fail-closed). Defaults to2000(2 seconds).
The service startup path starts the dispatcher after migrations are checked. Shutdown registers the dispatcher as a drainable service, so SIGTERM/SIGINT stops future polls and waits for the in-flight batch before closing database connections.
The dispatcher claims rows with:
SELECT ...
FROM webhook_outbox
WHERE processed = false
AND created_at <= NOW()
ORDER BY created_at ASC, id ASC
LIMIT $1
FOR UPDATE SKIP LOCKEDFOR UPDATE SKIP LOCKED lets multiple API instances run dispatchers concurrently without claiming the same row at the same time. A row is marked processed = true only after the HTTP attempt is complete. If the process exits before commit, PostgreSQL releases the lock and the row remains unprocessed for another worker to deliver, which provides at-least-once delivery.
Failed retryable deliveries are delegated to src/webhooks/retry.ts. The original row is marked processed and a new unprocessed row is inserted with created_at set to the next retry time. The dispatcher only claims rows whose created_at is due, so retries remain durable in PostgreSQL without holding process memory.
To prevent a slow or error-prone consumer from being bombarded with retries, attemptWebhookDeliveryWithRateLimit in src/webhooks/retry.ts enforces a per-consumer-URL sliding-window rate limit before each outbound attempt.
Per-consumer circuit breaker state is persisted in Redis (src/redis/webhookCircuitBreakerStore.ts) so multiple dispatcher instances and process restarts share the same open / half-open / closed view of a struggling consumer.
| State | Behaviour |
|---|---|
closed |
Deliveries allowed; consecutive failures increment toward the threshold. |
open |
Deliveries blocked until circuitBreakerResetMs elapses. |
half-open |
After reset expiry, one probe delivery is allowed across all instances. Success closes the circuit; failure re-opens it. |
- Before firing a retry, the dispatcher calls
checkWebhookDeliveryGate/attemptWebhookDeliveryWithRateLimitwith the consumer endpoint URL. - The circuit breaker store reads/writes JSON state at
webhook_cb:{sha256(url)}. Half-open probe ownership is tracked withwebhook_cb_probe:{sha256(url)}via RedisSET NX. - When the circuit is open, the outbox row is re-enqueued with
created_at = resetAt— no HTTP call is made. - Successful deliveries reset the breaker; retryable failures increment the shared failure counter.
- State transitions increment
fluxora_webhook_circuit_breaker_transitions_total{from_state,to_state}.
- Consumer URLs are SHA-256-hashed before use as Redis key segments (same approach as the rate limiter) to prevent key injection and to avoid storing raw URLs in Redis keys.
- A crafted URL cannot trip a breaker for a different consumer because keys are derived from the full URL digest.
| Condition | Behaviour |
|---|---|
| Circuit closed | Delivery proceeds (subject to rate limit). |
| Circuit open | Delivery deferred to resetAt; no consumer traffic. |
| Half-open probe succeeds | Circuit resets to closed. |
| Half-open probe fails | Circuit re-opens for another circuitBreakerResetMs. |
| Redis unavailable | Fail-open for gate checks; deliveries proceed. Failure recording is best-effort. (Rule 2 of docs/security/redis-outage-policy.md — availability-only, no deny/abuse gate.) |
- Before firing a retry, the dispatcher calls
attemptWebhookDeliveryWithRateLimitwith the consumer's endpoint URL and the configuredRateLimitConfig({ limit, windowMs }). - The rate limiter (
src/redis/webhookRateLimit.ts) maintains a Redis sorted set keyed by a SHA-256 hash of the consumer URL. Each recorded attempt is a member with score = timestamp (ms). - Entries older than
windowMsare pruned on every check. If the remaining count is at or abovelimit, the attempt is deferred rather than dropped. - A deferred attempt returns
{ shouldRetry: true, rateLimited: true, retryAt: now + windowMs }. The dispatcher re-inserts the outbox row withcreated_at = retryAt, so the deferral is durable in PostgreSQL. WEBHOOK_RETRY_RPS(default10) controlslimit;windowMsis1000 ms(one second).
To allow momentary spikes in retry traffic above the steady-state
WEBHOOK_RETRY_RPS limit (e.g. a batch of stream creations), opt in to
the token-bucket layer with WEBHOOK_RETRY_BURST. When the burst is
exhausted the limiter reverts to the steady-state rate configured above.
| Env var | Default | Description |
|---|---|---|
WEBHOOK_RETRY_BURST |
0 |
Token-bucket capacity per consumer. 0 keeps the legacy sliding-window behaviour (backward compat). |
When WEBHOOK_RETRY_BURST > 0:
- The bucket starts full with
bursttokens. Each successful delivery decrements the bucket by1.0. - Tokens refill at the steady-state rate
WEBHOOK_RETRY_RPStokens perwindowMs(default10 / 1000 ms=0.01tokens/ms), clamped toburst. - When the bucket has fewer than
1.0tokens, the attempt returns{ canAttempt: false, retryAfterMs: ceil((1.0 - tokens) / refillRateMs) }and the dispatcher defers it identically to the sliding-window path (outbox row re-enqueued withcreated_at = retryAt).
When WEBHOOK_RETRY_BURST = 0 (default), the limiter is exactly the
WEBHOOK_RETRY_RPS sliding-window described above — behaviour is
unchanged for existing deployments.
Observability — the bucket fill level per consumer is exported as
the Prometheus gauge fluxora_webhook_rate_limiter_bucket_fill{consumer_hash="…"}
(the consumer_hash label is the same SHA-256 prefix used as the
Redis sliding-window key, so dashboards that already join by consumer
continue to work). The gauge is updated on every
TokenBucketRateLimiter.checkLimit call. See
src/metrics/requestProtectionMetrics.ts for the label cardinality
guarantees (one time-series per currently-tracked consumer, not per
historical attempt).
Security — the bucket only refills at the steady-state
WEBHOOK_RETRY_RPS, so a configured burst cannot be abused to sustain
an effective outbound rate above the configured limit. Bursts absorb
instantaneous spikes; over a windowMs window the average rate is at
most WEBHOOK_RETRY_RPS. Bucket entries from inactive consumers are
cleaned up after 30 s of idleness, so a long-burst-then-disconnect
consumer does not pin a stale gauge series. See
src/webhooks/rate-limiter.ts and tests/webhooks/rate-limiter.test.ts.
| Condition | Behaviour |
|---|---|
| Within rate limit | Attempt proceeds; attempt recorded in Redis. |
| Limit exceeded | Attempt deferred; outbox row re-enqueued with retryAt = now + windowMs. No delivery is dropped. |
| Redis unavailable | Fail-open: attempt proceeds normally. A Redis outage does not halt deliveries. (Rule 2 of docs/security/redis-outage-policy.md — the only harm of a false default is lost availability, and there is no authorisation/abuse gate.) |
maxAttempts reached |
shouldRetry = false; row moves to dead-letter queue regardless of rate limit. |
Both the retry rate limiter (src/redis/webhookRateLimit.ts) and the circuit
breaker (src/redis/webhookCircuitBreakerStore.ts) fail open when Redis is
unavailable — the attempt is allowed and recorded best-effort. This is a
deliberate, rule-2 classification of the governing outage policy, not an
accidental catch: the cost of a false "allow" is only availability (extra
deliveries), while a false "deny" would stall all webhook deliveries. The fail-open
is observable via fluxora_webhook_rate_limiter_fail_open_total and error logs.
This is deliberate opposite of the fail-closed stores (JWT revocation, WS ban),
which must never admit a denied subject. See
docs/security/redis-outage-policy.md.
- Consumer URLs are SHA-256-hashed before use as Redis key segments to prevent key-injection via crafted URLs and to bound key length.
- The rate limiter counts all outbound attempts (not just failures) to protect consumers from burst traffic regardless of outcome.
- Redis credentials are consumed from environment variables only and are never logged.
Webhook requests are signed with the configured secret and include delivery metadata headers. Production endpoints must use HTTPS unless they target loopback for local deployments. URLs with embedded credentials are rejected.
Consumers must treat webhook delivery as at-least-once: verify the signature, deduplicate by x-fluxora-delivery-id, and make handlers idempotent.
Webhook consumers verify incoming requests by recomputing the HMAC-SHA256 signature using the shared signing secret. The verification path lives in src/webhooks/signature.ts.
| Header | Description |
|---|---|
x-fluxora-delivery-id |
Unique identifier for the delivery (used for deduplication). |
x-fluxora-timestamp |
Unix timestamp (seconds) at which the request was signed. |
x-fluxora-signature |
HMAC-SHA256 hex digest of {timestamp}.{rawBody}. |
x-fluxora-event |
Event type (e.g. stream.updated). |
- Reject if the payload exceeds
DEFAULT_MAX_WEBHOOK_BODY_BYTES(256 KiB). - Reject if the timestamp is not a positive integer.
- Reject if the timestamp is outside
DEFAULT_WEBHOOK_TOLERANCE_SECONDS(300s) of the current time. - Compute the expected signature and compare using a constant-time comparison (
timingSafeEqualover HMAC-hashed inputs) to prevent timing attacks. - If
isDuplicateDelivery(deliveryId)returns true, reject with409 duplicate_delivery.
When a webhook consumer rotates its signing secret via the admin API, there is a transition period during which some producers may still be signing with the old secret. To avoid spurious verification failures, the verification path supports a bounded dual-secret grace window:
- During the grace window, both the previous and current secret are accepted.
- After the grace window expires, the previous secret is rejected with code
previous_secret_expired(HTTP 401). - The rotation timestamp and grace-window expiry are persisted in the
webhook_secretstable (not held in memory), so a process restart cannot silently extend or shrink the window. - The default grace window is
DEFAULT_WEBHOOK_SECRET_GRACE_WINDOW_SECONDS(86 400 seconds / 24 hours).
verifyWebhookSignature accepts the following optional fields for rotation support:
| Field | Type | Description |
|---|---|---|
secretPrevious |
string |
The previous signing secret, valid only during the grace window. |
previousSecretRotatedAt |
number |
Unix timestamp (seconds) when the previous secret was rotated out. When omitted, the previous secret is accepted unconditionally (backward compatibility). |
graceWindowSeconds |
number |
Bounded grace window in seconds. Defaults to 86 400. Only consulted when previousSecretRotatedAt is also provided. |
- Set initial secret:
webhookSecretRepository.setSecret(id, secret)inserts a row with no previous secret. - Rotate:
webhookSecretRepository.rotateSecret(id, { newSecret, graceWindowSeconds })atomically moves the current secret toprevious_secret, setsprevious_secret_rotated_atandprevious_secret_expires_at, and activates the new secret ascurrent_secret. - Verify: The verification path checks both secrets. The previous secret is only accepted if
now < previous_secret_expires_at. - Cleanup:
webhookSecretRepository.clearExpiredPreviousSecret(id, now)nulls out the previous secret once the grace window has expired, providing defense-in-depth so the stale secret cannot be used even if the verification path is misconfigured.
- Bounded acceptance: The previous secret is rejected after
graceWindowSeconds— no indefinite acceptance of a stale secret. - Constant-time comparison: Both secrets are verified using the same
constantTimeComparepath, preventing timing-based secret enumeration. - Persistence: Rotation state survives process restarts because it is stored in PostgreSQL, not in-memory.
- Defense-in-depth cleanup: The
clearExpiredPreviousSecretmethod provides a second layer of protection by physically removing the previous secret after expiry. - Backward compatibility: When
previousSecretRotatedAtis not provided, the previous secret is accepted unconditionally, preserving existing behavior for callers that have not yet adopted the grace window.
| Code | Status | Description |
|---|---|---|
ok |
200 | Signature verified successfully. |
previous_secret_expired |
401 | The previous secret was provided but has exceeded its grace window. |
signature_mismatch |
401 | No provided secret matched the signature. |
missing_secret |
401 | No signing secret configured. |
missing_delivery_id |
401 | x-fluxora-delivery-id header missing. |
missing_timestamp |
401 | x-fluxora-timestamp header missing. |
missing_signature |
401 | x-fluxora-signature header missing. |
invalid_timestamp |
400 | Timestamp is not a positive integer. |
timestamp_outside_tolerance |
401 | Timestamp is outside the allowed tolerance window. |
payload_too_large |
413 | Request body exceeds the maximum allowed size. |
duplicate_delivery |
409 | Delivery ID has already been processed. |
All webhook target URLs are validated before any network call to prevent Server-Side Request Forgery (SSRF) attacks. This protection is applied in both the WebhookDispatcher class and the dispatchWebhook helper function.
The SSRF guard blocks the following IP address ranges:
- Loopback addresses:
127.0.0.0/8,::1(includinglocalhost) - Link-local addresses:
169.254.0.0/16(includes AWS metadata endpoint169.254.169.254),fe80::/10 - Private networks:
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,fc00::/7(IPv6 unique local) - Reserved ranges:
0.0.0.0/8,240.0.0.0/4,224.0.0.0/4(multicast) - IPv4-mapped IPv6 loopback:
::ffff:127.0.0.0/8
- HTTPS required by default: All webhook URLs must use HTTPS unless explicitly configured otherwise
- HTTP/HTTPS only: Other protocols (FTP, etc.) are rejected
The guard resolves hostnames to IP addresses and validates each resolved IP against the blocked ranges. This prevents DNS rebinding attacks where an attacker might initially point a hostname to a public IP, then change it to a private IP after validation.
The WEBHOOK_ALLOWED_HOSTS environment variable can be set to restrict webhook delivery to specific hosts:
WEBHOOK_ALLOWED_HOSTS=api.example.com,*.trusted.com- Supports exact hostnames:
api.example.com - Supports wildcard subdomains:
*.trusted.commatchessub.trusted.comandtrusted.com - When not configured, all non-blocked hosts are allowed
- Blocked IP ranges are always rejected, even if in the allowlist
All webhook fetches enforce a timeout (default 30 seconds) to prevent slow-loris attacks and hanging requests. The timeout is applied via AbortController in both the class-based dispatcher and the helper function.
Add to your environment configuration:
# Optional: Restrict webhook delivery to specific hosts
WEBHOOK_ALLOWED_HOSTS=api.example.com,*.trusted.comSSRF validation failures are logged without exposing the full URL for security. The validation fails closed: any ambiguous or unresolvable target is rejected with a WebhookTargetValidationError.
- Validation function:
validateWebhookTarget(url, options)insrc/webhooks/ssrfGuard.ts - Applied in:
WebhookDispatcher.dispatch()anddispatchWebhook()insrc/webhooks/dispatcher.ts - Timeout: Uses
DEFAULT_RETRY_POLICY.timeoutMs(30 seconds) - DNS resolution: Uses Node.js
dns.promises.lookup()