Building for the happy path only
Most integrations start the same way. You call the third-party API, it returns the expected response, you parse it, done. That code gets tested against a handful of real requests, it works every time, and it ships with confidence.
The problem shows up later, when the API doesn't behave the way it did in every test run. It times out. It returns a 500 error. It sends back a field that's suddenly null, an empty array instead of an object, or a response shape that's slightly different from what the docs described. An integration that only handles the "everything worked" case has no plan for any of that, and it will eventually fail in production, usually at the worst possible time, like during a launch, a sale, or a demo.
This is easy to miss because the failure modes rarely show up in a development environment. Your test API key doesn't get rate limited. The sandbox environment doesn't time out under load. The vendor's staging environment is usually more stable than their production one. None of that is representative of what happens once real users and real traffic are involved.
Handling the unhappy path isn't optional polish. It's the difference between a failed API call being a minor blip your system absorbs and retries quietly, and it being an outage your users notice and complain about.
- Decide what happens on a timeout, not just a clean error response.
- Validate the shape of the response before you trust it, instead of assuming every field will be there.
- Have a defined fallback, even if that fallback is just a clear error message instead of a broken page.
- Write at least one test that simulates a bad response, not just a good one.
Rate limits: fine in testing, a real problem at scale
Most third-party APIs cap how many requests you can make in a given time window, whether that's per second, per minute, or per day. During development, with a handful of test calls a day, you will never come close to that limit. It simply doesn't show up as a problem until real traffic hits.
Then it does. The moment usage scales up, whether that's more users, a scheduled batch job, or a sudden traffic spike, requests start getting throttled or rejected outright. If there's no retry or backoff logic built in, this doesn't fail loudly with an obvious crash. It often fails silently, with requests just quietly not completing while everything else in the system looks normal from the outside.
A common version of this: a nightly sync job that pulls data for every customer in one long loop, calling the same API hundreds or thousands of times back to back. It works fine with ten test accounts. It starts failing partway through with ten thousand real ones, and nobody notices until someone asks why last week's data is missing.
Code that has never hit a rate limit in testing is not proof the rate limit doesn't matter. It's proof you haven't sent enough traffic yet.
The fix is straightforward but has to be built in deliberately. Read the rate-limit headers most APIs return so you know how much headroom you have left. Back off and retry with increasing delay when you're throttled, rather than immediately hammering the endpoint again. Queue or batch requests instead of firing them all at once, and spread bulk jobs out over time when the vendor allows it.
- Check for rate-limit headers (remaining requests, reset time) on every response, not just failed ones.
- Add jitter to retry delays so many clients don't all retry at the exact same moment.
- Cache responses where the data doesn't need to be fetched fresh every time.
Authentication tokens that expire without a refresh plan
API keys and OAuth tokens often expire by design. That's a security feature, not an oversight on the vendor's part, and it's meant to limit how much damage a leaked credential can do. But an integration that stores a token once and assumes it stays valid forever will work fine for days or weeks, and then mysteriously stop working with no code change on your end to explain why.
The tricky part is how this usually gets discovered. It's rarely caught by monitoring, because from the system's point of view nothing crashed, nothing threw an unhandled exception, a request just started returning 401 errors that may or may not be logged anywhere useful. Instead it's reported by a confused user who can't figure out why a feature that worked yesterday doesn't work today, days after the token actually expired.
This is especially common with OAuth integrations where a user connected their account once, months ago, and the refresh token itself has since been revoked or expired without anyone noticing. The integration looks connected in your database. It just silently stopped working.
- Refresh tokens automatically before they expire, not after a request fails.
- Handle the expired-token error explicitly, with a retry after refresh rather than surfacing a raw error.
- Alert on repeated auth failures instead of letting them fail silently in a log nobody reads.
- Give users a clear way to reconnect an account if a refresh token has been revoked entirely.
Treating a third-party outage as your own bug
When an external API goes down or starts degrading, every feature in your system that depends on it fails too. That part is unavoidable, no amount of good engineering on your end prevents a vendor's own infrastructure from having a bad day. What's avoidable is what happens next: without clear error handling and status visibility, an outage on the vendor's side gets misread as a bug in your own code.
That misdiagnosis is expensive. Engineers spend hours digging through logs and recent deploys looking for a regression that doesn't exist, comparing this week's code against last week's, when the real answer was a status page update from the vendor that nobody thought to check first.
The pattern usually looks like this: a feature starts failing intermittently, someone assumes it's related to a recent release, and a chunk of the team spends the afternoon reviewing a deploy that had nothing to do with it. Meanwhile the actual cause resolves itself an hour later when the vendor fixes their outage, and the postmortem never quite explains what really happened.
Integrations should make it obvious, at the point of failure, whether the problem is external. A clearly labeled error like "payment provider unavailable" saves far more time than a generic exception that looks identical to an internal crash. It's also worth subscribing to the status pages of your critical vendors, so an outage shows up as a known event rather than a mystery.
No monitoring on integration health specifically
General application monitoring is good at telling you when your own server is unhealthy: CPU spikes, memory leaks, a crashed process. It's often much worse at tracking the health of the third-party APIs you depend on, things like failure rates, response latency, or how close you are to a rate limit. Those signals live in a different place, and if nobody builds a dashboard for them, nobody sees them.
Without that specific visibility, integration problems don't get caught by an alert. They get caught by users complaining, which is a much slower and much more painful way to find out something is wrong. By the time a support ticket lands, the problem may have been quietly happening for hours or days.
The fix doesn't need to be elaborate. Even a simple dashboard tracking success rate and average response time per external API, checked weekly, catches slow degradation long before it becomes a full outage. The goal is to know about a problem before a customer does, not after.
- Track error rates and latency per external API, separately from overall app health.
- Alert on a rising trend, not just a hard failure threshold.
- Log rate-limit warnings before they turn into hard failures.
- Review integration health on a regular cadence, not only when something breaks.
Building integrations that assume the third party will never change
Third-party APIs are not static. Versions get deprecated, response formats change, fields get renamed or removed, and default behavior shifts between releases. Vendors ship these changes on their own schedule, often with nothing more than a changelog entry as warning, and sometimes with a deprecation window measured in weeks rather than years.
An integration built with no buffer for this, no pinned API version, no tested upgrade path, breaks unexpectedly the day the vendor ships a change you didn't know was coming. And because it often worked fine for months or even years beforehand, it's rarely the first place anyone looks when something goes wrong. Teams end up debugging their own recent code before they think to check whether the vendor changed anything.
This is worse when an integration was built once, by someone no longer on the team, with no documentation of which API version it targets or what assumptions it makes about the response format. Nobody owns keeping it current, so it just runs until it doesn't.
Pin the API version you integrate against where the vendor supports it, keep an eye on their changelog or deprecation notices, and treat an upgrade as a tested change rather than something that happens automatically underneath you. Build in a little slack for the fact that the ground will shift eventually. This is exactly the kind of resilience we build into every integration on our integrations and APIs work, so it's handled up front instead of discovered in production.