Picture a typical morning for an on-call engineer. A message lands in the chat: "Everything is slow." What exactly is slow? The whole pool, or one slice? A proxy in a specific country, or with a specific carrier? Is the target site responding slowly, or is the retry queue growing? Without numbers, that's not an incident, it's reading tea leaves. And while the team guesses, time passes and money bleeds.

This article is about turning a vague feeling of "it's slow" into a precise diagnosis in one minute. We'll cover what metrics to collect around a proxy pool, what to log per request and what you absolutely must never log, why percentiles matter more than averages, how to slice data by dimensions, and how to set up alerts that don't wake you up for no reason. At the end, you'll find a quick-response table and a practical FAQ.

An important note on scope. Here we only talk about measurement and alerting. Choosing a specific IP for a request, node health checks, and quarantine logic inside the pool are a separate big topic covered in a dedicated article. Our job here is narrower: detect degradation, localize it, and raise the alarm in time. Examples use Proxeon infrastructure, but the principles are universal.

Fundamentals: what observability is and why proxies need it so acutely

Let's start with the basics. Observability is a property of a system where its external outputs let you understand its internal state without poking around inside with a debugger. The classic observability triad is metrics, logs, and traces. Metrics answer "what's happening" in aggregate, logs answer "what exactly happened to this specific request," and traces answer "how did the request travel through the whole chain."

How are proxies different from a regular web service? You have a third party you don't fully control: the proxy node itself, the channel to it, and the target resource behind it. A regular app can be profiled down to the last function. A proxy adds a layer of network uncertainty where degradation can come from anywhere: the carrier, the routing, an overloaded node, or a changed behavior of the target site.

That's exactly why observability around a proxy pool isn't a luxury, it's hygiene. Without it, you're flying blind. With it, you see the structure of the problem: not "everything is bad," but "the mobile proxy slice for one carrier in one region degraded, everything else is fine." That's the difference between panic and surgical precision.

Three levels where problems live

It's useful to keep three levels in mind from the start, where degradation originates:

  • Transport level: the channel to the proxy, packet loss, connection setup time. This is where timeouts and slow TTFB live.
  • Proxy node level: overload of a specific IP, limit exhaustion, carrier issues. This is where a rising error rate on a specific slice lives.
  • Target resource level: the site started responding slower, returned non-standard codes, changed its limits. Here it's crucial not to confuse a site problem with a pool problem.

A good observability system lets you see at a glance which of the three levels the problem is on. That's the very "one minute to diagnosis" we're after.

Four signals for proxies: why these specifically

There's a temptation to collect everything. Hundreds of metrics, dozens of dashboards, miles of charts. That's a trap. Too many metrics means noise, and noise means that during an incident you won't find what you need. Experienced engineering practice says the opposite: start with a minimal set of signals that cover most problems. For a proxy pool, there are four such signals.

Signal one: success rate

Success rate is the percentage of requests that completed as expected out of the total. It's the main health indicator. If success rate drops, something broke right now, right at the user.

The key question: what counts as success? The naive answer "code 200" is wrong. It's more correct to define success through the contract of your usage. Success is often counted as all 2xx and 3xx codes, as well as meaningful 4xx codes that are a valid response from the target resource rather than a proxy problem. But timeouts, connection drops, proxy-level errors, and mass 5xx are failures.

Formally, success rate belongs to the "availability" family of metrics in the SLI (Service Level Indicator) model. It's the very indicator around which SLOs (Service Level Objectives) and error budgets are later built.

Signal two: latency by percentiles

Latency is the time from sending a request to receiving a response. But a single latency number is meaningless. You need percentiles: p50, p95, p99. Why percentiles and not the average we'll cover in detail in a separate section, because it's one of the most underrated topics in all of monitoring.

For proxies, TTFB (Time To First Byte) is especially valuable. It separates network latency and server reaction time from the time it takes to transmit the response body. If TTFB grows, it's a network or node problem. If the total time grows but TTFB is stable, perhaps the responses just got bigger or the channel throughput dropped.

Signal three: retry rate

Retry rate is the percentage of requests that required a repeated attempt. It's an early harbinger of trouble. Often success rate is still normal because retries are pulling the situation through, but retry rate has already started creeping up. It's like a temperature of 37.2: formally you're still working, but the body is already fighting.

Retries mask degradation from the end user, but they devour resources: time, traffic, pool capacity. Ignoring this signal is doubly dangerous, because rising retries can avalanche the system when repeated requests add load to already overloaded nodes.

Signal four: bandwidth usage

Bandwidth usage is the volume of data transferred. Why is it in the top four? First, it's direct money, because traffic is billed. Second, anomalous usage is a signal: a sudden spike can mean someone is pulling extra, responses have bloated, or retries are looping the same data around. A sudden drop to zero on a slice that usually has activity means the slice simply stopped working.

These four signals aren't random. They echo the "golden signals" observability methodology popularized by reliability engineers: latency, traffic, errors, saturation. We've adapted it to the specifics of proxies, where retries deserve their own place as a domain-unique harbinger.

Deep dive: what to log per request

Metrics show trends. But when you need to understand what happened to a specific request, logs save you. A well-designed request log is your black box that you turn to during incident analysis. Let's break down what you must log and what you must never log.

What you must log

  • Proxy identifier: not the raw IP in the clear, but a stable node or pool identifier. This lets you link a request to a specific resource and see which nodes are causing problems.
  • Response code: HTTP status or transport error code (timeout, connection refused, drop). This is the foundation for calculating success rate.
  • Time to first byte (TTFB): in milliseconds. One of the most informative indicators for localizing network problems.
  • Total request time: from start to finish, also in milliseconds.
  • Response size: in bytes. Feeds the traffic metric and helps spot anomalously large or empty responses.
  • Attempt number: is this the first attempt or already a retry, and which one. Without this field, you can't calculate retry rate.
  • Dimensions for slicing: country, carrier, proxy type. A separate section covers them, but they must be in the log.
  • Timestamp and trace identifier: to link records to each other and to external systems.

Here's what a structured log entry in JSON format might look like. Note: it's machine-readable, which is critical for later analysis.

{"ts":"2026-02-14T08:12:33Z","trace_id":"a1b2c3","proxy_id":"pool7-node042","country":"RU","carrier":"op-a","proxy_type":"mobile","status":200,"ttfb_ms":187,"total_ms":342,"bytes":20481,"attempt":1,"outcome":"success"}

What you must NEVER log

This is no less important than what you must log. Logs have a way of leaking, getting copied into analytics systems, ending up in backups. Everything you put there lives long and in unexpected places.

  • Credentials: logins, passwords, proxy authorization tokens, API keys. Never. Not even partially. Not even "temporarily for debugging."
  • The full response body: first, it's a huge volume; second, it may contain personal and sensitive data. Log only the size and, if needed, a hash or short signature.
  • Headers with secrets: Authorization, Cookie, Set-Cookie and the like. They must be stripped before logging.
  • Full URLs with sensitive parameters: if the query string has tokens or personal identifiers, they must be masked.
  • Users' personal data: anything covered by personal data legislation must either not end up in the log or be anonymized.

A practical masking technique at log formation time:

def sanitize(entry):  secret_keys = {"authorization", "cookie", "set-cookie", "proxy-authorization"}  headers = {k: ("***" if k.lower() in secret_keys else v) for k, v in entry.get("headers", {}).items()}  entry["headers"] = headers  entry.pop("body", None)  entry.pop("proxy_credentials", None)  return entry

The golden rule: a log should let you diagnose a problem but must never turn into a database of leaked secrets. If you're in doubt about whether to log a field, don't log it. Diagnostic value can almost always be obtained through safe surrogates: hashes, sizes, flags, categories.

Percentiles over averages: why the average hides the problem

This is the section worth reading twice. Because here lies the most common and most insidious mistake in performance monitoring.

Why the average lies

Imagine: you have a hundred requests. Ninety-nine of them completed in 100 milliseconds, and one in 10 seconds. The average time will be around 199 milliseconds. Looks great, almost nothing changed. Meanwhile, one of your users waited ten seconds and, most likely, already left, cursing.

The average is a device that smears outliers across the whole sample. It's sensitive to extremes but insensitive to the structure of the distribution. And the performance of network systems almost always has a long-tail distribution: most requests are fast, but a minority are very slow. And it's this tail that determines real user experience and the presence of problems.

What percentiles are and how to read them

A percentile is the value below which a given percentage of observations falls. Let's break down the three main ones:

  • p50 (median): half the requests are faster than this value, half are slower. It's the "typical" experience.
  • p95: 95 percent of requests finished within this time. It's the "almost worst case" experience that affects a noticeable share of users.
  • p99: 99 percent of requests are faster. It's the very long tail where timeouts, retries, and angry users live.

In our example with a hundred requests, p50 and p95 stay around 100 milliseconds, but p99 jumps to 10 seconds. The percentile honestly showed the problem the average hid. That's why experienced engineers look at p95 and p99 first and almost never use the average for latency assessment.

How to calculate percentiles correctly

The naive way is to collect all values, sort them, and take the needed position. It's accurate but doesn't scale: with millions of requests, storing the whole array is impossible. In practice, approximate counting structures are used: fixed-bucket histograms or special algorithms like t-digest and HDR histograms.

The histogram idea is simple: you predefine time ranges (buckets) and just count how many requests fell into each. From the accumulated counters you can easily reconstruct any percentile with acceptable accuracy, while memory is fixed.

import bisectclass PercentileTracker:  def __init__(self):    self.buckets = [10, 25, 50, 100, 200, 500, 1000, 2000, 5000]    self.counts = [0] * (len(self.buckets) + 1)  def add(self, ms):    i = bisect.bisect_left(self.buckets, ms)    self.counts[i] += 1  def percentile(self, p):    total = sum(self.counts)    if total == 0:      return None    target = total * p / 100    acc = 0    for i, c in enumerate(self.counts):      acc += c      if acc >= target:        return self.buckets[min(i, len(self.buckets) - 1)]    return self.buckets[-1]

One important warning about aggregation. Percentiles cannot be averaged. If you have p95 across ten nodes, you can't average those p95s and call the result the overall p95. That's mathematically incorrect. For correct aggregation, you need to add up histograms and then compute the percentile from the combined histogram. That's exactly why modern monitoring systems store histograms, not ready-made percentiles.

Slicing by dimensions: seeing a slice, not the whole pool

Here's the moment where observability transforms from a chart into a diagnostic tool. A single success rate number for the whole pool tells you little. It might be 97 percent and look normal, hiding that one slice fell to 40 percent while the rest pull the average up.

Three key dimensions for proxies

  • Country: the geography of the proxy. Degradation is often geographically localized: a routing problem in one region, changes on the target resource side for certain countries.
  • Carrier: for Proxeon mobile proxies, this is a critically important dimension. A problem with a specific carrier will show up right here, and you'll immediately understand its scale.
  • Proxy type: mobile, server, residential. Different types behave differently, and degradation of one type shouldn't get lost in the overall mass.

Cardinality: where to stop

There's a temptation to slice data by everything: by each IP, by each target domain, by each user. That leads to a cardinality explosion, the number of unique label combinations. High cardinality kills monitoring systems: storage grows, queries slow down, infrastructure gets more expensive.

A practical rule: slice by dimensions with a limited and stable set of values. Countries number in the dozens, carriers in the units or dozens, proxy types in the units. That's safe. But a single IP or full URL as a metric label can't be used, they're unique in huge quantities. Such details belong in logs, where they're stored line by line, not in metrics, where they multiply data series.

from prometheus_client import Counter, Histogramrequests_total = Counter(  "proxy_requests_total",  "Total proxy requests",  ["country", "carrier", "proxy_type", "outcome"])latency_ms = Histogram(  "proxy_latency_ms",  "Request latency",  ["country", "carrier", "proxy_type"],  buckets=[10, 25, 50, 100, 200, 500, 1000, 2000, 5000])def record(country, carrier, ptype, outcome, ms):  requests_total.labels(country, carrier, ptype, outcome).inc()  latency_ms.labels(country, carrier, ptype).observe(ms)

With this slicing, you can build a query in seconds: show success rate by carrier for the last hour. And immediately see that it's not the whole pool that degraded, but one slice. That's localization. The beauty of this approach is that it turns the panic of "everything is broken" into a calm "slice X of carrier Y needs attention."

Alerts that don't make noise

The most common reason teams stop trusting monitoring is noisy alerts. When the system wakes you five times a night for false alarms, you'll very quickly start ignoring it. And then you'll miss a real incident. This is called alert fatigue, and it kills observability more effectively than a complete lack of monitoring.

Three principles of quiet alerts

Principle one: thresholds on symptoms, not causes. Alert on what the user feels: falling success rate, rising p99 latency. Not on intermediate technical fluctuations that don't mean a problem on their own.

Principle two: observation windows. Don't react to a single spike. One slow request is noise. A sustained deviation over a time window is a signal. Configure the alert to fire when the condition holds, say, for five minutes straight, not at a single moment.

Principle three: hysteresis. This is a term borrowed from engineering, meaning different thresholds for firing and clearing. The alert fires when success rate drops below 90 percent, but clears only when it rises above 95. The gap between thresholds prevents "flapping," when the metric oscillates around a single value and the alert blinks on and off.

Example alert configuration

Here's an example rule in a style understood by most monitoring systems. It fires if the success rate on any carrier slice stays below the threshold for the observation window.

groups:- name: proxy-health  rules:  - alert: LowSuccessRateByCarrier    expr: |      sum by (carrier) (rate(proxy_requests_total{outcome="success"}[5m]))      /      sum by (carrier) (rate(proxy_requests_total[5m]))      < 0.90    for: 5m    labels:      severity: warning    annotations:      summary: "Success rate below 90 percent for carrier"

Severity levels and routing

Not all alerts are equal. Split them by severity:

  • Warning: something deviated, worth a look during business hours. Doesn't wake you at night.
  • Critical: users are suffering right now, immediate reaction needed. Wakes the on-call.

A sensible starting set includes just a few alerts: a critical drop in overall success rate, success rate degradation on a slice, a sharp rise in p99 latency, an anomalous spike in retry rate. No more. Every new alert is a promise that someone will react to it. Don't make promises you can't keep.

Error budget as a frame

An advanced technique is to use an error budget instead of hard thresholds on instantaneous values. If your goal is 99 percent successful requests per month, the error budget is that very one percent you can "spend." An alert on the burn rate reacts not to a single failure but to the fact that you're spending the allowable error limit too fast. Such alerts are much calmer and more accurately reflect the real threat to your SLO.

Quick diagnostics by dashboard: three typical pictures

Now for the most interesting part. How do you understand in a minute where the degradation is? The answer is to train your eye on a few typical patterns. A good dashboard is not a dump of charts but a pattern-recognition tool. Let's break down three classic pictures.

Picture one: one slice fell, everything else is normal

You look at success rate broken down by carrier. The overall number dipped slightly, but looking at the breakdown you see: one carrier collapsed to 50 percent, the rest hold at 98. Latency on the problem slice grew, on the rest it's stable.

What it means: localized problem at the node or carrier level. The cause isn't in your system or the target resource as a whole, but in a specific slice of the pool. It's a transport or node problem.

First action: pull the problem slice out of active rotation (that's already quarantine territory, a separate topic) and keep watching. Check whether the degradation is tied to a specific region within the carrier.

Picture two: latency grew everywhere, but there are no errors

Success rate is stable, close to one hundred percent. But p95 and p99 latency grew across all slices simultaneously and evenly. Retry rate rose slightly.

What it means: when everything degrades at once and evenly, look for a common factor. Most often it's either your own infrastructure (overload, resource shortage, a bottleneck in your code) or the target resource started responding slower for everyone. Proxy nodes aren't the issue here, otherwise the degradation would be uneven.

First action: look at TTFB separately. If TTFB grew, it's the network or the server. If TTFB is stable but total time grows, the problem is in body transmission or processing. Check the load on your side and the target resource's metrics.

Picture three: retry rate grows while success rate is stable

Success rate looks fine, around 97 percent. But retry rate started creeping up: it was 3 percent, now it's 15. Latency also grew, because retries add time.

What it means: this is the most insidious pattern, because the final result is still normal. But the system is running at its limit: it's spending more and more repeated attempts to hold the success rate. It's a precursor to collapse. If the trend continues, retries will stop saving it, and success rate will crash.

First action: don't wait for success rate to collapse. Find which slice has rising retries (again, slicing by dimensions) and dig into the root cause before it's too late. Check whether the retries themselves are creating extra load that spins the spiral.

Dashboard layout for one-minute diagnostics

To read these patterns in a minute, put just four panels on the main screen, corresponding to the four signals, each with quick slicing by dimensions:

  1. Success rate: overall and broken down by carrier and country.
  2. Latency: p50, p95, p99 on one chart, so you can see tail divergence.
  3. Retry rate: trend over the last hours.
  4. Bandwidth usage: by slice, to catch anomalies.

Everything else is secondary and lives on separate screens. The main screen should answer one question: is everything fine, and if not, where exactly. Nothing extra.

Quick response table: metric, its rise, and first action

This table is worth printing and hanging next to the on-call engineer's desk. It turns observation into action without extra deliberation.

Metric: falling success rate (whole pool)

What it means: a mass failure affecting most requests. A system-level problem.
First action: check your own infrastructure and the target resource, since an even drop rarely comes from individual nodes.

Metric: falling success rate (one slice)

What it means: localized degradation of a carrier, country, or proxy type.
First action: localize the slice via the breakdown and pull it from rotation while watching the dynamics.

Metric: rising p99 latency with stable p50

What it means: the tail lengthened, some requests became very slow, while the typical request is normal.
First action: find the slice with the growing tail, check timeouts and nodes producing spikes.

Metric: rising p50 and p95 simultaneously and evenly

What it means: overall performance degradation, likely infrastructure or the target resource.
First action: separate TTFB from body transmission time, check the load on your side.

Metric: rising retry rate with stable success rate

What it means: the system is masking degradation with repeats, a precursor to collapse.
First action: find the slice with rising retries and eliminate the root cause before success rate collapses.

Metric: anomalous bandwidth usage growth

What it means: bloated responses, extra repeats, or unplanned activity.
First action: correlate traffic growth with request count and response size, find the source.

Metric: bandwidth usage drops to zero on an active slice

What it means: the slice stopped serving requests entirely.
First action: check the availability of the slice's nodes and connectivity, escalate if confirmed.

Typical proxy observability mistakes

The experience of analyzing dozens of incidents lets us collect a catalog of rakes people step on most often. Knowing these mistakes saves months of pain.

Mistake 1: looking at the average instead of percentiles

We've covered this, but let's repeat it, because the mistake is so widespread. Average response time doesn't show long-tail problems. The team sees a stable average and is sure everything is fine, while users complain about freezes. Always p95 and p99.

Mistake 2: logging secrets "temporarily for debugging"

Temporary has a way of becoming permanent. A token logged "for five minutes to check" settles in the log storage system for months and ends up in backups. Masking secrets must be a hard rule at the logging library level, not each developer's moment-to-moment decision.

Mistake 3: label cardinality explosion

Slicing metrics by each IP or URL seems convenient until the monitoring system starts choking and demanding more and more resources. Labels only for dimensions with a limited set of values. Details go in logs.

Mistake 4: too many alerts

A team proud of its monitoring sets up forty alerts. A month later half of them are noisy, the on-calls turn them off in notifications, and one day they miss a real incident because it got lost in the stream. Fewer alerts, but more precise.

Mistake 5: alerts without a window and hysteresis

An alert fires on an instantaneous value and immediately clears, then fires again. Notification flapping is annoying and devalues the system. Observation windows and hysteresis are mandatory.

Mistake 6: counting only code 200 as success

This way you either understate success rate by treating valid responses as failures, or conversely miss problems. Define success through your usage contract meaningfully, not mechanically by a single code.

Mistake 7: no slicing by dimensions

A single overall success rate chart hides local problems. Without slicing you see that "it's fine overall" and miss a fallen slice. Slicing isn't an option, it's a necessity.

Mistake 8: ignoring retries

Many don't even count retry rate, relying only on the final success rate. And they lose the earliest harbinger. By the time success rate drops, retries have been screaming about the problem for a while.

Mistake 9: storing percentiles instead of histograms

If you save ready-made p95s by slice, you can't correctly compute the overall p95, because percentiles don't add up. Store histograms, compute percentiles at query time.

Mistake 10: dashboard as a dump

Fifty panels on one screen isn't observability, it's information noise. During an incident, the eye gets lost. The main screen is minimalist, details on click.

Tools and resources

The good news: quality observability around a proxy pool doesn't need an expensive, complex stack. Let's cover the minimally sufficient set and the logic of choice.

Metrics collection and storage

For metrics, time-series-model systems with label and histogram support work great. The key requirement is histogram support for correct percentile calculation and slicing by dimensions without cardinality explosion. Such systems let you run queries like "success rate by carrier for an hour" on the fly.

Visualization

For dashboards you need a tool that can build time-series charts, overlay several percentiles on one chart, and quickly switch the slice by dimensions. The ability to have variable filters matters: chose a carrier and all panels rebuild for it. This speeds up diagnostics many times over.

Log collection and storage

Request logs must be structured (JSON) and land in a system that allows filtering by fields: by proxy identifier, by response code, by slice. A mandatory requirement is a retention policy with automatic deletion after expiry, so sensitive data doesn't accumulate forever.

Alerting

The alerting system must support observation windows (condition holds for N minutes), severity levels, and channel routing. Especially valuable is support for error budget burn rate alerts for calm but precise firing.

Libraries for instrumenting code

In the proxy client code, use metric libraries that support counters and histograms with labels. Wrap each request in a measurement: record the start time, and on completion record the outcome, latency, and traffic. Instrumentation must be centralized, in one place, so a new developer can't accidentally skip it.

import timedef instrumented_request(client, url, meta):  start = time.monotonic()  attempt = meta["attempt"]  try:    resp = client.get(url)    elapsed = (time.monotonic() - start) * 1000    outcome = "success" if resp.status_code < 500 else "server_error"    record(meta["country"], meta["carrier"], meta["type"], outcome, elapsed)    log_request(meta, resp.status_code, elapsed, len(resp.content), attempt, outcome)    return resp  except TimeoutError:    elapsed = (time.monotonic() - start) * 1000    record(meta["country"], meta["carrier"], meta["type"], "timeout", elapsed)    log_request(meta, None, elapsed, 0, attempt, "timeout")    raise

2026 trends

The observability industry in 2026 is moving in several notable directions. First, telemetry standardization based on open protocols, which simplifies integrating metrics, logs, and traces into a unified picture. Second, growing interest in exponential histograms, which give accurate percentiles with minimal memory use. Third, the use of automatic anomaly detection based on statistical models, which complements threshold alerts by noticing unusual patterns for which you can't preset a threshold. And fourth, a shift in focus from the quantity of data collected to its meaningfulness: fewer metrics, but the right ones. That's exactly what we're talking about in this article.

Case studies and results

To give the principles flesh, let's go through several generalized scenarios built on typical production experience with proxy pools. The numbers are illustrative, but the patterns are real.

Case 1: the invisible degradation of one carrier

A team worked with a Proxeon mobile proxy pool and relied on overall success rate. The figure held around 96 percent, no alarm. Meanwhile users on one direction complained about failures. After adding carrier breakdown, the picture became instantly clear: one carrier gave a 62 percent success rate, the rest around 99. The overall figure masked the failure of a whole slice.

Result: after adding carrier slicing and an alert on slice degradation, the time to detect such problems dropped from several hours (via complaints) to several minutes (via alert). The problem slice started being pulled from rotation in time, and overall success rate on the direction rose to 98 percent.

Case 2: the long tail hidden behind the average

Another team monitored average latency, which held at a comfortable 240 milliseconds. Occasional complaints about "freezes" were written off as quirks. The switch to percentiles opened their eyes: p50 was indeed around 190 milliseconds, but p99 reached 8 seconds. Every hundredth request was painfully slow.

Result: after the team started watching p99 and set an alert on its growth, they found the tail was generated by requests to a certain group of nodes during peak load hours. They localized the problem by slicing. p99 was reduced to 1.2 seconds, and the number of freeze complaints dropped to practically zero.

Case 3: the retry spiral

The third scenario is instructive in its danger. A system with an aggressive retry policy held success rate around 97 percent, and everything seemed stable. But nobody watched retry rate. One day, with a small load spike, retry rate grew from 5 to 40 percent in half an hour. Repeated requests added load, nodes overloaded more, retries grew even more, a classic spiral. An hour later success rate collapsed to 60 percent.

Result: the incident postmortem led to introducing a separate retry rate metric and an early alert on its growth. Now when retry rate reaches the threshold, the team gets a warning well before success rate collapses. Similar situations started being caught at the harbinger stage, never reaching an outage. This is a clear illustration of why retries deserve a place among the four main signals.

The common takeaway from the cases

Three stories, three different problems, but one pattern. In every case the data for detection physically existed in the system but wasn't presented so the problem became visible. Slicing by dimensions, percentiles over averages, and attention to retries aren't abstract recommendations. They're specific lenses, each making its own class of problems visible.

FAQ: common questions about proxy observability

Where to start if you currently have nothing at all?

Start with logging every request in structured form with mandatory fields: proxy identifier, response code, TTFB, size, attempt number, dimensions. Even without metrics and dashboards, this already gives you the ability to analyze incidents. As a next step, add the four metrics and one or two critical alerts. Don't try to build everything at once, a minimal working set is worth more than a perfect unfinished one.

How often should metrics be sampled?