# Webhook Best Practices

> **Operating principle:** A webhook is a real-time push integration. Design the receiving endpoint to acknowledge quickly, process asynchronously and recover safely when either side experiences a problem.


## 1. Design and set up

### Keep each subscription purposeful

1. Create a clear business owner, technical owner, downstream system and criticality for every subscription.
2. Use separate subscriptions for materially different workloads or service levels. This makes troubleshooting and recovery more controlled.
3. Subscribe only to required entities and event types. Avoid catch-all subscriptions unless the business use case explicitly needs them.


### Build a resilient receiving endpoint

1. Use HTTPS and an appropriate authentication mechanism. Protect secrets, rotate credentials and monitor TLS certificate expiry.
2. Return HTTP 200 promptly after safely accepting the event, then perform downstream processing asynchronously. Sprinklr recommends responding within 10 seconds.
3. Make processing idempotent by using a stable event or business identifier so duplicate deliveries or replays do not create duplicate outcomes.
4. Size the endpoint, proxy and downstream queue for expected peak and replay traffic. Apply controlled back-pressure rather than silently dropping events.


## 2. Apply filters at the source

| Filter decision | Recommended approach | Example validation |
|  --- | --- | --- |
| Entity and event type | Select only events the consumer uses. | Confirm create/update/delete or message event coverage against the use case. |
| Business attributes | Filter by relevant queues, brands, statuses or business units where supported. | Generate one matching and one non-matching event; only the matching event should arrive. |
| Change control | Treat filter changes as production changes. | Record the previous rule, expected volume impact and rollback approach. |


## 3. Test before production

| Test | Exit criteria |
|  --- | --- |
| Happy path | Authentication works; expected payload and headers arrive; endpoint returns 200; downstream result is correct. |
| Filter coverage | Matching events are delivered and non-matching events are excluded. |
| Failure handling | Simulate timeout, non-200 response and endpoint unavailability; verify retries and failed-event retrieval. |
| Replay and duplicates | Replay a bounded test window; verify idempotent processing, reconciliation and no duplicate business action. |
| Load and latency | Validate peak traffic, proxy capacity and prompt acknowledgement before production cutover. |


*SPRINKLR PLATFORM • OPERATIONS & RECOVERY*

## 4. Monitor the integration end to end

| Within Sprinklr | Customer endpoint and downstream |
|  --- | --- |
| Review subscription status and configuration. | Track request rate, 2xx and non-2xx responses, timeouts and response latency. |
| Review available success/failure indicators and retrieve failed events for the affected subscription and time range. | Monitor TLS certificates, DNS, proxies, load balancers, queues and consumer health. |
| Compare expected business activity with delivered/retrieved events. | Maintain processed-event counts and a reconciliation checkpoint. |
| Escalate sustained failures with subscription ID, timeframe and endpoint evidence. | Alert on abnormal failure rate, latency, backlog and capacity saturation. |


## 5. Understand retries and the failure queue

**Documented delivery behavior:** Sprinklr waits up to 10 seconds for HTTP 200. If delivery is unsuccessful, the event is retried through three retry queues. Events that fail all three retries are stored in the Deleted Queue. The 10-second setting is not configurable.

```
Send --10s--> Retry 1 --10s--> Retry 2 --> Retry 3 --> Failed Queue
```

## 6. Recover and replay safely

1. **Stabilize** — Resolve endpoint, network, certificate or downstream failure. Confirm the endpoint can accept live and replay traffic.
2. **Define scope** — Identify the subscription ID and precise start/end time. Keep the replay window as narrow as practical.
3. **Retrieve** — Use Retrieve Failed Webhook Events for deliveries recorded as FAILED. Failed webhooks are available for retrieval for seven days.
4. **Replay** — Use the Webhook Replay API for the selected subscription and time duration. Continue with the returned cursor when more data is available.
5. **Reconcile** — Compare retrieved/replayed event counts with successfully processed outcomes. Investigate discrepancies before closing the incident.
6. **Prevent recurrence** — Capture the cause and update monitoring, filtering, capacity, certificate controls or runbooks.


> **Important:** If your endpoint returned 200 but the event was lost later in your downstream process, it may not be recorded as FAILED. Use a bounded time-range replay and rely on idempotent processing and reconciliation.


## Go-live and incident checklist

- [ ] Owner and use case documented
- [ ] Filters validated
- [ ] Endpoint returns 200 promptly
- [ ] Alerts configured
- [ ] Failure retrieval tested
- [ ] Replay tested
- [ ] Idempotency verified
- [ ] Runbook and escalation path agreed