vM.

A Practical Observability Checklist for Python APIs in Production

Author
Vishal Maurya
Published on
Reading time
4 min read

Overview

When a production API slows down or returns errors, a log line saying “request failed” is rarely enough to explain why. You need to connect the request to its timing, database calls, outbound HTTP requests, and any background work it triggered.

Observability is the ability to infer what is happening inside a system from the signals it emits. For a Python API, start with a small set of useful logs, metrics, and traces rather than collecting every possible field.

1. Use Structured Logs

Structured logs store fields such as request ID, route, status code, and duration separately from the message. This makes them easier to filter and aggregate than arbitrary text.

import logging
import time
from fastapi import FastAPI, Request

app = FastAPI()
logger = logging.getLogger(__name__)


@app.middleware("http")
async def request_metrics(request: Request, call_next):
    started = time.perf_counter()
    try:
        response = await call_next(request)
    except Exception:
        logger.exception("request_failed", extra={"path": request.url.path})
        raise

    logger.info(
        "request_completed",
        extra={
            "method": request.method,
            "path": request.url.path,
            "status_code": response.status_code,
            "duration_ms": round((time.perf_counter() - started) * 1000, 2),
        },
    )
    return response

This example assumes your logging configuration serializes the extra fields. In a real service, normalize route templates instead of recording arbitrary path values when those paths contain IDs. Add a validated correlation ID if requests need to be traced across services.

2. Choose Metrics That Support Decisions

Useful API metrics include request rate, error rate, and latency percentiles. Add dependency latency, database pool usage, queue depth, and worker task duration when those components are part of the service.

Avoid metric labels with unbounded values such as user IDs, request IDs, or raw URLs. High-cardinality labels can overwhelm monitoring storage. Put per-request identifiers in logs and traces instead.

3. Trace Slow Requests Across Dependencies

A distributed trace can show that a request spent most of its time waiting for a database query or third-party API. It can also reveal repeated calls that are invisible in a single aggregate duration.

OpenTelemetry provides instrumentation and APIs for traces, metrics, and logs. Instrument the framework and important dependencies, then propagate context across service boundaries where supported.

Tracing is not a substitute for query plans or provider logs. It helps identify where to investigate next.

4. Alert on User Impact

An alert that fires whenever one request fails may be noisy. Prefer signals that reflect sustained problems, such as an elevated error ratio, rising tail latency, a growing queue backlog, or a dependency outage.

Set thresholds based on the service's expected behavior and user-facing objectives. A background report generator and a checkout API should not necessarily have the same latency target.

5. Keep Sensitive Data Out of Telemetry

Do not log API keys, authorization headers, passwords, full payment details, or complete sensitive document contents. Review exception messages because libraries sometimes include URLs, query parameters, or input values.

Apply access controls and retention policies to logs and traces. Observability data can contain sensitive information even when it was collected for debugging.

6. Make the Signals Actionable

Every important alert should point to a next step: inspect a dashboard, check a dependency, review a recent deployment, or follow a runbook. Include service ownership and escalation expectations so an alert does not become an unexplained notification.

Before relying on a dashboard, test it by generating a known error or latency spike in a controlled environment. Confirm that the signal appears and the alert behaves as expected.

7. Review Observability During Releases

After deploying a change, compare error rate, latency, resource use, and dependency behavior with the previous revision. If a regression appears, use the traces and logs to narrow the cause before making unrelated changes.

Conclusion

Useful observability starts with a few consistent signals tied to actual operational questions. Structured logs explain individual requests, metrics show trends, and traces connect work across dependencies.

If your Python API is difficult to debug in production, I can help instrument the service, improve correlation across components, and build actionable dashboards and alerts. Contact me.

Additional Resources

  • OpenTelemetry documentation
  • Prometheus documentation
  • FastAPI documentation