The Hidden Cost of DIY Webhooks: Why Your Retry Loop Is Losing Money
Every API company builds webhooks in-house first. The reasoning is airtight: the initial version takes a sprint, the infrastructure is boring, and paying a vendor for something you can hand-write feels wasteful. Two years later, the same company is spending more on webhook maintenance than the vendor would have cost, and every DX engineer knows the retry code by heart from grepping it during incidents.
This is a breakdown of where the money actually goes.
Why does the two-week estimate always become a two-year estimate?
The first webhook build hits a working state fast. You have a job queue, a POST call, a retry-on-failure wrapper, and a database table logging attempts. It ships to production, customers integrate, everyone claps.
Then the operational load starts. A customer's endpoint goes down over the weekend and your queue backs up 40,000 events behind their slow retries. You add per-endpoint isolation. A customer complains they cannot verify signatures after you rotated your secret. You add dual-key rotation. Support cannot answer "did you send the event" so engineering writes a query for the on-call runbook. That query gets run 400 times a quarter.
None of these tasks were in the original scope. All of them are non-negotiable once they surface. The two-week build is real; the two-year drip of maintenance is what nobody planned for.
Where does the DX team's time actually go?
Interviews with platform teams at API-first companies show a consistent pattern. Roughly one in three integration support tickets traces back to a webhook that either did not arrive, arrived late, or failed signature verification on the customer's side. That translates to a specific breakdown of DX time.
| Activity | Share of DX time on webhook issues |
|---|---|
| Answering "did you send it" tickets | 35 to 45% |
| Debugging signature or rotation issues | 15 to 20% |
| Investigating retries and delivery gaps | 15 to 20% |
| Manual replays and backfills | 10 to 15% |
| Migrating customers between event schema versions | 10 to 15% |
Every one of those categories has a technical solution. None of them are in the initial two-week build.
What is the actual dollar cost of that load?
Take a Series B API company with a five-person DX team, one platform engineer partially allocated to webhooks, and 300 customers consuming events. The load looks like this:
- DX support hours. 20% of five engineers at fully loaded $220K, so $220K per year is going to webhook-related tickets, alone.
- Platform engineering. One engineer at 25% permanent allocation to webhook maintenance, another $55K per year.
- On-call pages. Roughly two webhook-related pages per month, each consuming three hours across triage and postmortem, so around $30K per year in interrupted engineering time.
- Migration and reprocessing. Once per quarter someone spends a week on a payload schema migration or a backfill run, so around $30K per year.
That totals $335K per year, low end $180K if the numbers are smaller, high end $420K if you have a bigger customer base or higher event volume. A vendor solving the same problem lists at $30K to $60K per year at the same scale. The gap is not close.
How does webhook unreliability show up as churn?
The number one reason developer buyers churn from an API product, according to internal exit interviews at three companies who share this data anonymously, is "the integration was flaky." Webhooks are almost always the flaky part.
Here is the mechanism. A customer's product depends on your webhook to update state. Your webhook fails silently one time in a hundred. The customer notices their side is out of sync, files a support ticket with the customer of theirs who complained, and starts logging your reliability. Six months later, when a competitor pitches them at a conference, the conversation is not about price. It is about reliability.
You never see this on your delivery dashboard because your delivery dashboard reports a 99.9% success rate, which sounds fine until you multiply by the customer's event volume and realize you are losing three events a day for their biggest use case.
What breaks at scale that does not break in the first year?
Three failure modes hit in year two, not year one.
- Signature rotation without dual-key support. Your first rotation breaks fourteen customers because they hardcoded the secret in a config file that nobody updated. Support runs for a week. The engineering fix is a two-key rolling window that should have been in v1.
- Slow endpoint contagion. One customer's endpoint takes 30 seconds to respond. Your worker pool has 50 threads. Their events consume all 50 threads. Every other customer's webhook is now delayed. Fix requires per-endpoint concurrency isolation, which is a nontrivial refactor of your queue.
- Schema evolution. Two years in, you want to change the shape of
invoice.paid. Half your customers integrated against v1 and cannot handle v2. You have no per-endpoint transformation, so you have to run two event streams in parallel and email everyone about a migration. That migration takes a quarter and burns political capital across every account.
Each of these is solvable. None of them are cheap. All of them are predictable to anyone who has done this before.
When should you actually stop building this yourself?
Two signals, whichever comes first.
- Customer count. At 200 or more customers with webhook subscriptions, per-endpoint issues start becoming per-endpoint incidents. The support cost curve steepens.
- Ticket share. When webhook-related tickets exceed 10% of your total support volume, the operational cost has crossed the vendor cost. Every quarter past this point is a quarter of writing a check to yourself for the privilege of maintaining internal infrastructure.
If either is true and you are not already evaluating a vendor, you are making the buy-versus-build decision by default rather than by decision. Default-build is the most expensive path.
What actually matters
The mistake to avoid is treating webhook infrastructure as a build-once decision. It is a compounding operational cost, and the compounding is invisible because it hides inside DX support hours and churn analytics rather than a line item on your budget. Look at the numbers directly: what percent of your support volume is webhook-related, what percent of your DX team's week goes to signature debugging and replay requests, and how many customers cite reliability in exit interviews. If those numbers are non-trivial, you are already paying for webhook infrastructure. The only choice left is whether you keep paying yourself or start paying someone whose entire roadmap is webhooks.
Frequently asked questions
How much engineering time does a homegrown webhook system consume per year?
For a company with 100 or more webhook-consuming customers, expect one engineer at 20 to 40% capacity permanently allocated to webhook maintenance. That is $80K to $160K in loaded engineering cost, before you count the DX and support team hours that route through the same problem. Most CFOs never see this line because it lives inside general platform engineering.
What is the customer-churn cost of unreliable webhooks?
For an API-first product, unreliable webhooks are the number one cited reason for churn among developer buyers, ahead of price. A single lost payment notification is often enough to trigger a switch. Companies that have measured this internally typically find webhook reliability accounts for 8 to 15% of gross churn, or roughly $250K to $600K in ARR per year for a $10M ARR company.
Why does a homegrown retry loop lose money silently?
Because failures without observability show up as customer-side issues, not vendor-side issues. The customer thinks their integration is buggy, files a ticket, escalates internally, and eventually decides your product is unreliable. You never see the root cause on your dashboard because you did not ship the dashboard. The revenue leaks through NPS and renewal, not through your webhook logs.
Isn't a background job queue enough for webhook delivery?
It is enough to get to a first version, but a job queue does not do per-endpoint isolation, HMAC rotation, customer-facing debug UI, or signature verification for inbound webhooks. You end up building three more systems around the queue and calling the collection your webhook stack. That is the point at which it starts consuming a permanent engineer.
When is the right time to move off a homegrown webhook stack?
When you cross 200 customers with webhook subscriptions, or when webhook tickets exceed 10% of your support volume, whichever comes first. Both signals mean the operational cost has surpassed the migration cost. Waiting past that point is a decision to keep paying the higher of the two indefinitely.
Send webhooks. Prove they arrived.
Yaranex handles retries, signing, rate limits, and payload search behind one API, so your customers stop opening tickets and your on-call sleeps.
Request early access