How to Debug Intermittent API Failures in Production
- Author
- Vishal Maurya
- Published on
- Reading time
- 5 min read
Overview
An API can work locally and still fail intermittently in production. A third-party service may slow down, a database pool may run out of connections, or traffic may expose a race condition that a local test never hits. Repeating the same request may succeed, which makes guesswork especially expensive.
The first goal is not to change code. It is to collect enough evidence to distinguish an application bug from a dependency, capacity, or network problem.
1. Classify the symptom
| Symptom | First places to inspect |
|---|---|
| 500 from the application | Exception logs, input-specific code paths, failed dependencies |
| 502 or 504 from a gateway | Upstream health, proxy timeout, connection resets |
| 429 from a dependency | Quotas, request rate, retry behavior |
| Requests hang | Missing timeouts, blocked event loop, connection pools |
| Failures rise with traffic | CPU/memory pressure, database capacity, concurrency |
These are leads, not diagnoses. Record the route, status, latency, timestamp, request ID, and dependency timings for both successful and failed requests.
2. Set explicit outbound timeouts
A request to another service should have a timeout that fits inside the overall request budget. With HTTPX:
import httpx
async def fetch_customer(customer_id: str) -> dict:
timeout = httpx.Timeout(10.0, connect=3.0)
async with httpx.AsyncClient(timeout=timeout) as client:
response = await client.get(
f"https://api.example.com/customers/{customer_id}"
)
response.raise_for_status()
return response.json()
This example shows timeout configuration; it is not a complete client-lifecycle pattern for every application. In a real service, reuse a configured AsyncClient where appropriate so connection pooling works as intended. The timeout is not necessarily a hard end-to-end deadline for the whole application operation.
3. Retry only recoverable failures
Retries can help with transient connection errors and some server failures. They can also amplify an outage if every caller immediately repeats requests against an unhealthy dependency.
A retry policy should define eligible errors, maximum attempts, backoff and jitter, the total time budget, and what happens after the last attempt. Respect Retry-After when the provider supplies it. Do not retry every 4xx response; invalid input and permission failures usually require a different action.
Be careful with state-changing operations. If a payment or order was committed but its response was lost, repeating the request can duplicate the operation. Use the provider's idempotency mechanism when available and make internal writes idempotent where practical.
4. Correlate logs across services
A request may pass through a gateway, an API, a worker, and several dependencies. A correlation ID makes the path searchable. In FastAPI, middleware can attach an ID and record duration:
import logging
import time
import uuid
from fastapi import FastAPI, Request
app = FastAPI()
logger = logging.getLogger(__name__)
@app.middleware("http")
async def log_request(request: Request, call_next):
request_id = str(uuid.uuid4())
started = time.perf_counter()
try:
response = await call_next(request)
except Exception:
logger.exception("request_failed request_id=%s", request_id)
raise
response.headers["X-Request-ID"] = request_id
logger.info(
"request_completed request_id=%s method=%s path=%s status=%s duration_ms=%.2f",
request_id,
request.method,
request.url.path,
response.status_code,
(time.perf_counter() - started) * 1000,
)
return response
This is a minimal example. If clients are allowed to provide their own request IDs, validate them or create a separate trusted internal ID. In a multi-service system, propagate the ID to internal calls. Never log authorization headers, API keys, or sensitive request bodies.
5. Check database and pool behavior
If failures appear after the service has been running under load, inspect connection-pool usage, slow-query logs, query plans, and transaction duration. Common causes include too many connections across multiple application workers, long-running transactions, missing indexes, or holding a database transaction open while waiting on an external API.
Increasing the pool size without checking the database's connection capacity can make an overload worse. Compare application and database metrics before changing the limit.
6. Measure latency by dependency
Useful signals include request counts by route and status, p50/p95/p99 latency, outbound dependency duration, database pool usage, queue depth, CPU, and memory. Keep high-cardinality identifiers such as request IDs in logs and traces rather than metric labels.
Look at the timeline. If dependency latency rises before application errors, investigate that dependency first. If the application becomes slow while dependencies remain healthy, inspect local CPU, blocking operations, locks, and resource pools.
7. Reproduce one failure mode
Once the evidence points to a cause, write a focused test: simulate a slow dependency, a 429 response, a dropped connection, or malformed input. A targeted test is more useful than generating random traffic and hoping the incident reappears.
Deploy the smallest safe change, watch the same metrics that exposed the problem, and keep a rollback path. An error rate returning to normal is evidence of improvement, but continue monitoring for a recurrence under similar conditions.
Conclusion
Intermittent failures are easier to fix when requests are measurable, dependencies have explicit timeouts, retries are bounded, and logs can be correlated across services. If your application is suffering unexplained timeouts or unstable integrations, I can help trace the request path and fix the underlying backend issue rather than masking the symptom.
Contact me with the failing endpoint, observed status codes, and what you have already checked.