Picture this. It's Friday evening, your parser or account management system is finally live. The first few minutes go smoothly. Then things start getting weird: some requests fail, a CAPTCHA pops up here and there, an account suddenly asks for identity verification, and your logs turn into a mess of connection errors. Sound familiar? Nine times out of ten, the root cause isn't the target site or your business logic. It's how your application manages its IP addresses.

Most projects start simple: a text file with a list of addresses, reading line by line, random selection. And it works—right up until the load becomes real. Then the naive list falls apart, and you get floating failures that can't be reproduced with a single console request.

This article is a comprehensive guide to designing an in-app proxy pool. We'll cover how to pick an address for a specific task, how to check address health, how to quarantine problematic IPs and bring them back, and how to maintain a sticky session for a single account. It's an engineering topic, but we'll discuss it in accessible language, with code, checklists, and real-world mistake analysis.

Let's set the boundaries upfront. We're not discussing timeouts and retry policies for a single request—that's a whole other topic. We're also not covering how the address physically changes on a modem or hardware—that's infrastructure, not application. Our focus is on pool logic in your code.

Why a Text File of Addresses Breaks Under Real Load

Let's be honest about the typical starting implementation. There's a file with a hundred lines of addresses. The app reads the file, stuffs the lines into an array, and picks a random element for each request. Simple, understandable, works in a demo. So why does it fall apart?

Problem One: No Memory of Address State

Random selection from an array knows nothing about what just happened to that address a second ago. If address 47 just returned five errors in a row and is clearly unhealthy, the random algorithm will likely pick it again. You'll keep banging on a dead address, wasting time, retries, and—more importantly—the trust of the target platform.

Problem Two: No Task-to-Address Binding

When you're working with accounts, it's critical that one account always goes out through the same IP. Antifraud systems notice when a user's session jumps across a dozen different subnets from different geographic zones in just a few minutes. That's unnatural. Real people don't behave that way. Random selection from a file guarantees that chaos.

Problem Three: Race Conditions with Multiple Workers

As soon as you run multiple parallel processes or threads, a simple in-memory array becomes a source of race conditions. Two workers read the same index, both grab the same address, both load it twice as hard. There's no coordination. Counters, if you have any, get lost between processes.

Problem Four: No Observability

A text file is silent. It won't tell you which address is returning eighty percent CAPTCHAs, which has become slow, or which has been dead for a day. You're working blind and only learn about problems through indirect signs in business metrics—when it's too late.

The bottom line is simple. A list of addresses is just data. Managing those addresses under load is a system. The difference between the two is what this article is all about. Next, we'll build that system piece by piece.

Fundamentals: Pool, Address Lease for a Task, and What Counts as a Session

Let's start with the basics. If you're an experienced engineer, I still recommend reading this—we're establishing the terminology we'll use throughout the article.

What is an Address Pool

A pool isn't just a list of addresses; it's an object with behavior. It stores a set of addresses along with their state and provides two main operations: give out an address for a task and take it back. Classic analogy: a library. You have shelves of books, but between you and the shelves stands a librarian. They know which books are checked out, which are damaged and in repair, and which are available. You don't browse the shelves yourself—you ask the librarian, and they hand you a suitable book.

Leasing an Address for a Task

A lease is a temporary assignment of an address to a specific task. The pool hands out the address, marks it as busy (or accounts for new load on it), and when the task finishes, the address is returned. That's why the pool interface has two methods that will be our workhorses: acquire—to get an address, and release—to return it with a report of the outcome.

The report on return is crucial. When a task returns the address, it tells the pool how things went: success, connection error, CAPTCHA, block. That information feeds all further health and quarantine logic.

Sticky Binding vs. Random Selection

Here lies a critical boundary. Random selection means the pool gives out a random address for each request. This mode is good for tasks that don't have a concept of a continuous session: mass collection of public data from independent pages where each request stands alone.

Sticky binding means that a specific key—say, an account identifier—always gets the same address. This is critical where continuity of identity over time matters. Working with an account is the typical example. The account should appear to the platform as a single stable user, not a swarm of entities hopping across the network.

What Counts as a Session at the Business Level

The word session is overloaded. At the HTTP level it's one thing, at the TCP level another. But we're interested in the business session. It's a logically connected sequence of actions that, from the platform's perspective, should come from the same user from the same address.

Examples of business sessions:

  • Logging into an account, performing a series of actions inside it, logging out—all from the same IP.
  • A multi-step scenario: opening a product page, adding to cart, checking out—all steps are connected.
  • Working under one profile throughout a workday, where a change of IP would look like a suspicious move.

Defining session boundaries is your design decision, not a technical given. You decide whether a session lasts fifteen minutes, an hour, or until explicit logout. That definition determines how long the pool holds the key-to-address binding. Remember this—we'll come back to sticky sessions several times.

Deep Dive: The Lifecycle of an Address in the Pool

Before we dive into individual strategies, it's helpful to see the big picture. Every address in a well-built pool has a lifecycle, a set of states it transitions between.

Address States

  • Healthy – the address is available for use, metrics are normal.
  • Degraded – the address is still handed out, but with reduced weight because performance has declined.
  • Quarantined – the address is temporarily excluded from selection; it's on timeout.
  • Probing – the address is undergoing an active check before returning to service.
  • Dead – the address is considered non-functional for the long term or permanently.

The transitions between these states are the system we're building. A healthy address degrades as errors accumulate. When a degraded address hits a threshold, it goes into quarantine. After the backoff time, it goes to probing. If probing succeeds, it comes back healthy. If not, it returns to quarantine with an increased backoff. Several consecutive failures, and the address is declared dead.

Why an Abstraction Layer Is Needed

Key insight: the application should not know about address states. Business code simply asks for an address and returns it with a result. All the complexity of the lifecycle is hidden inside the pool. This is the principle of encapsulation, and it pays off a hundredfold. When you want to change the quarantine strategy, you modify one module, not dozens of places in business logic.

Mental Model of Load

It's helpful to keep three axes in mind for each address:

  • Freshness – how recently we last checked its health.
  • Load – how many tasks are currently using it.
  • Reputation – the accumulated history of successes and failures.

Any selection strategy essentially combines these three axes into a single score. Next, we'll look at specific strategies.

Address Selection Strategies: From Round-Robin to Key-Based Hashing

The heart of the pool is the algorithm that decides which address to hand out on each acquire call. There's no single right answer. The strategy choice is driven by the task. Let's go through four basic approaches and explain when each is appropriate.

Round-Robin: Going in Circles

Round-robin is the simplest fair algorithm. Addresses are arranged in a ring, and a pointer moves one step per request. Once you've gone through the whole circle, you start over. Pro: perfectly even load distribution with identical addresses. Con: the algorithm is blind to differences between addresses. Fast and slow ones get an equal share of traffic.

When to use: a homogeneous pool, tasks without sessions, where all addresses are roughly equal in quality and capacity.

Weighted Selection

Weighted selection is an evolution of round-robin where each address is assigned a weight. An address with a higher weight gets more traffic. The weight can be static, based on known capacity, or dynamic, recalculated from live metrics like success rate and latency.

Dynamic weight is a powerful tool. An address starts returning more CAPTCHAs? Lower its weight—it gets less traffic but isn't removed entirely. Metrics recover? The weight goes back up. This is smooth self-regulation, gentler than an abrupt quarantine.

A simple formula for weight: weight = success rate divided by normalized latency. Higher success and lower latency mean a higher weight. Recalculate it over a sliding window, say the last fifty requests.

Least-Connections: The Least Loaded

Least-connections hands out the address currently handling the fewest active tasks. This works brilliantly when tasks vary greatly in duration. Round-robin in such a scenario could load one address with long tasks while another sits idle. Least-connections naturally balances actual load, not just the number of handed-out requests.

To implement this, you need a counter of active leases per address. It goes up on acquire and down on release. The address with the smallest counter is selected. Note: this counter is shared state, and with multiple workers it must live in a common store. We'll come back to that in the section on state storage.

Key-Based Hashing: One Account, Always One Address

Now for the main event when working with accounts. Key-based hashing is a strategy where the address is chosen deterministically based on the task's key, not randomly. Take the account ID, compute a hash, take the modulo of the number of addresses, and you get an index. The same account always yields the same index, and thus the same address.

Why is this not just convenient but fundamentally important? Here's the key insight of the entire article.

Why Deterministic Hashing Beats Randomness for Antifraud

Antifraud systems build behavioral profiles. One of the strongest signals of trust is the stability of the network environment. A real person day after day goes online from roughly the same set of IPs. Their ISP changes rarely, their geography is stable.

Now imagine an account showing a new IP from a different subnet with every action. To the system, that looks like impossible behavior: a person can't physically be in ten places in a minute. Random address selection guarantees exactly that red flag.

Deterministic hashing solves the problem at its root. The account is bound to the address mathematically, without storing state. Even if the app restarts and loses all memory, the same account will compute the same address again. It's self-healing binding. Neat, isn't it?

The Pitfall: Changing Pool Size

Naive modulo hashing has a sneaky weakness. If the number of addresses changes—one is added or taken to quarantine—the modulo changes for almost all keys. Nearly all accounts suddenly move to new addresses. That's precisely the disaster we were trying to avoid.

The solution is consistent hashing. This technique ensures that adding or removing one address reassigns only a small fraction of keys, not all of them. Addresses and keys are placed on a virtual ring; a key goes around the ring to the nearest address. Remove an address—only its keys move, the rest stay put. Consistent hashing is the correct foundation for sticky binding in a live pool where addresses come and go.

Combining Strategies

In a real application, strategies are often combined. A typical advanced scenario: for account-keyed tasks, use consistent hashing; within the group of addresses responsible for a reservation, use least-connections. For sessionless data collection, use weighted round-robin. The pool can hold multiple strategies and select the right one based on task type.

Strategy Selection Checklist

  • Has session or account binding? Use key-based hashing, preferably consistent.
  • Tasks are independent and homogeneous? Round-robin.
  • Addresses vary in quality? Weighted selection with dynamic weight.
  • Tasks have very different durations? Least-connections.
  • Mixed load? Combination with pool subdivision into subgroups.

Health Checks: Passive and Active Probes

A pool is only as good as its knowledge of address health. That brings us to the health-check mechanism. There are two complementary approaches: passive and active. You need both.

Passive Health Check from Traffic Errors

Passive checking doesn't make separate requests. It observes the real traffic that's already going through the address. Each release call brings a result, and the pool updates the address's reputation based on it. It's free—you already made the request for business reasons; you just also note its outcome.

What to consider as a deterioration signal in passive observation:

  • Connection errors: the address is not responding, connection drops.
  • Responses typical of network-level blocking.
  • A spike in CAPTCHAs—indirect but important signal of declining address reputation.
  • A sharp increase in latency relative to the address's historical norm.

The advantage of passive is that it's current and adds zero extra load. The downside is that it only reacts after traffic has already flowed—so the first few affected requests are inevitable. And it knows nothing about idle addresses.

Active Health Check via Probes

Active checking means the pool makes separate probes, independent of working traffic. Typically a lightweight request to a known control resource that responds reliably and allows judgment of address health.

Active probes solve what passive can't: they check idle addresses and addresses in quarantine before returning them to service. The active probe is the gate that decides whether an address gets released from quarantine back into the pool.

What Counts as a Probe Failure

Defining a failure is a delicate matter. Too strict, and you'll discard normal addresses due to random spikes. Too lenient, and dead addresses will hang around. A reasonable approach is a multi-factor threshold.

  • A single error is not a failure. It's noise. Networks are inherently unreliable.
  • A failure is an accumulated signal: for example, three errors out of the last five requests, or the success rate on a window falling below seventy percent.
  • CAPTCHA rate should have its own, lower threshold because CAPTCHAs signal address reputation, not random glitches.

What Intervals to Use

Active probe intervals are a trade-off between data freshness and extra load. General guidelines:

  • Healthy addresses that are under load don't need active probing—passive traffic covers them.
  • Idle healthy addresses—probe every thirty to sixty seconds to keep them ready.
  • Quarantined addresses—probe on the backoff schedule discussed below.
  • Don't probe all addresses at once. Stagger probes in time, add random jitter to avoid synchronous bursts.

Flapping and How Hysteresis Suppresses It

Now one of the most underestimated phenomena. Flapping is when an address rapidly bounces between healthy and sick states. A probe passes—back to service. Immediately an error—quarantined. A second later another probe passes—back in service. And so on. This exhausts the system, creates jitter in metrics, and prevents the address from either working or resting properly.

The cure is hysteresis. A term from electronics meaning different thresholds for entering and leaving a state. The idea is simple: you need fewer signals to declare an address sick than to declare it healthy again. For example, three consecutive errors trigger quarantine, but to return, you need five consecutive successful probes. The asymmetry creates a zone of stability where small fluctuations don't flip the state.

A second tool against flapping is a minimum hold time in a state. An address that enters quarantine must stay there for a minimum time, even if a probe succeeds earlier. That dampens rapid oscillations. The combination of hysteresis and minimum hold turns a nervous system into a calm, predictable one.

Health-Check Checklist

  • Both types of checks work: passive via traffic and active via probes.
  • Failure is defined as an accumulated signal, not a single error.
  • CAPTCHA rate has its own, more sensitive threshold.
  • Active probes are staggered in time with random jitter.
  • Hysteresis is configured: entry threshold lower than exit threshold.
  • Minimum hold time per state is set to prevent flapping.

Quarantine and Return to Service: Backoff, Limits, and Pool Protection

When an address is recognized as problematic, you can't just throw it away. Often the problem is temporary. The job of quarantine is to give the address a break and then check if it has recovered. There's a lot of fine engineering here.

Exponential Backoff

Naive quarantine holds an address for a fixed time, say a minute, and then returns it. But if the address is persistently problematic, you'll keep returning it, each time getting another round of failures. The solution is exponential backoff.

The principle: with each consecutive trip to quarantine, the backoff time grows. First time—one minute. Second consecutive time—two minutes. Then four, eight, sixteen. Up to a reasonable ceiling, say an hour. Once the address has worked successfully for long enough, the backoff counter resets to the initial value.

This elegantly solves two problems at once. A temporarily stumbling address comes back quickly. A persistently sick one will bother the system less and less frequently, effectively becoming dead without cluttering the pool with frequent useless probes.

Add random jitter to the backoff. Without it, all addresses that entered quarantine at the same moment would return in a simultaneous burst. Jitter spreads the returns out in time.

Limit on Simultaneous Quarantined Addresses

Here's a critical safeguard that is most often forgotten. What if an incident affects half the addresses at once? For example, the target platform tightens its checks. The pool will honestly start moving addresses into quarantine one by one. And if you don't set a limit, you'll end up with almost no one left in service.

Hence the rule: a hard limit on the proportion of addresses in quarantine at the same time. For example, no more than thirty percent of the pool. If the limit is reached, new candidates are not quarantined but only reduced in weight. The logic is: better to work through slightly degraded addresses than to be left with no resources at all.

Protection Against the Entire Pool Going into Quarantine

This is an extension of the previous thought, taken to the extreme. Imagine: the probe to the control resource itself breaks. The resource goes down or changes its response. The pool decides all addresses are dead and puts the entire pool in quarantine. The application grinds to a halt even though the addresses are actually fine.

Protection mechanisms:

  • Guaranteed minimum in service. Always keep at least one or two addresses available, even if they technically failed a check. Better to let them work under suspicion than to stop the system.
  • Correlation analysis. If failures hit all addresses simultaneously, that's suspicious. More likely the common factor broke: the probe, your network, the target resource. The pool should recognize a mass failure and not panic by removing everyone.
  • Separate control resources. Don't tie your active probe to a single control point. If it becomes your single point of failure, false positives will crash the entire pool.

Return to Service Through a Half-Open State

Coming out of quarantine is not instantaneous. Good practice is a half-open state, an idea from the circuit breaker pattern. An address leaving quarantine doesn't immediately get full traffic. First, it gets a tiny fraction, a test trickle. If the test requests succeed, the fraction gradually increases until the address regains full weight. If the tests fail, back to quarantine with an increased backoff.

This smooth ramp-up protects against a premature burst of traffic hitting an address that hasn't fully recovered.

Quarantine Checklist

  • Backoff grows exponentially with repeated quarantine entries.
  • Jitter is added to backoff to prevent synchronous returns.
  • A limit is set on the proportion of addresses in quarantine at once.
  • A minimum number of addresses are guaranteed to be in service under all conditions.
  • Mass failure detection is in place to recognize common problems.
  • Return to service uses a half-open state with gradual load increase.

State Storage: Process Memory vs. Shared Store

All the logic we've discussed—load counters, reputation, quarantine states—is data that lives somewhere. Where it lives is a defining architectural decision. It depends on whether you have one process or many.

Process Memory: Simple but Lonely

If your application runs in a single process, the simplest place to hold pool state is in memory. Regular data structures: a dictionary of addresses, their counters and timers. Fast, no dependencies, no network latency.

The downsides are obvious. A restart wipes all history; the pool starts from scratch, forgetting who was in quarantine. And crucially, it doesn't work as soon as you have more than one process. Each process will have its own isolated view of the world. One process puts an address in quarantine, but another doesn't know and keeps loading it.

Shared Store: Redis and Memcached

As soon as you have multiple workers—and under real load you almost always do—the state must become shared. That's where fast stores like Redis or Memcached come in. They run as separate services that all workers talk to, giving a single consistent view of the pool state.

Redis is preferable to Memcached in most cases because it offers atomic operations, data structures like sorted sets and hashes, atomic increment counters, and the ability to run small atomic scripts. All of that will be useful.

Race Conditions with Multiple Workers

A shared store solves visibility, but creates a new problem: race conditions. Classic least-connections scenario: two workers simultaneously read counters, both see that address X is the least loaded, both pick it, both increment the counter. Result: the address gets double the load, even though the algorithm was supposed to avoid that.

Solutions:

  • Atomic operations. Incrementing a counter in Redis is atomic by nature. Use that instead of read, increment, write separately.
  • Scripts. For complex selection logic that needs to read multiple values and make a decision, package it as a single atomic script on the store side. Then no one can interject between reading and writing.
  • Distributed locks. For critical sections, you can take a short lock. But be careful—locks hurt performance and can themselves become a source of problems. Atomic operations are almost always better.

TTL for Records

A crucial tool in a shared store is TTL, time-to-live, after which a record automatically disappears. It saves you from accumulating garbage and from states getting stuck.

Where to apply TTL:

  • Sticky key-to-address binding. Remember the business session? Its duration is the TTL of the binding record. Set a TTL of fifteen minutes, and after fifteen minutes of inactivity, the binding automatically dissolves—the account can get a fresh binding on the next access. That's a clean, elegant way to express session lifetime.
  • Active lease counter. If a worker crashes without calling release, the counter will forever be inflated. TTL or periodic reconciliation protects against that. Often a lease is stored as a record with a TTL slightly longer than the maximum task duration, so that a crashed worker doesn't hold the address forever.
  • Quarantine state. The backoff time itself is naturally expressed through TTL: the quarantine record lives exactly as long as the backoff, and when it expires, the path to return is opened.

Hybrid Approach

In practice, a hybrid is often used. Hot, frequently read data is cached locally in the worker's memory for a short period, while the source of truth remains the shared store. This reduces the number of calls to Redis, but requires care with desynchronization. A good compromise is a local cache with a very short TTL, say one or two seconds, for metrics that don't require instant accuracy.

Observability: Per-Address Metrics and How to Tell a Bad IP from a Bad Site

You can't manage what you don't measure. Observability turns the pool from a black box into a transparent system where you can see the health of each address and understand the causes of problems.

Per-Address Metrics

At a minimum, track the following for each address over a sliding window:

  • Success rate. The ratio of successful completions to total tasks. The main integral health indicator.
  • Latency. Track not just the average, but percentiles. Median and, say, the 95th percentile tell you much more about tails than an average, which is easily skewed by outliers.
  • CAPTCHA rate. A separate, very telling metric. A rise in CAPTCHA rate on an address is an early signal of its reputation decline, often before direct errors increase.
  • Number of active leases. Current load, needed for least-connections and for understanding distribution.
  • Quarantine history. How many times and for how long has the address been quarantined? Chronic offenders are immediately visible.

Pool-Wide Metrics

  • Proportion of addresses in each state: healthy, degraded, quarantined, dead.
  • Overall throughput and aggregate success rate.
  • Frequency of quarantine entries over time—a spike indicates an incident.
  • Quarantine limit utilization—approaching the limit is a warning sign.

How to Tell a Bad IP from a Bad Site

Here's a question that separates a mature system from a naive one. Errors are pouring in—but who is to blame? A problematic address, or is the target site itself temporarily unavailable to everyone? If you confuse them, you'll start quarantining healthy addresses for the sins of another server.

The method for differential diagnosis is correlation analysis:

  • Problem on a single address. If errors and CAPTCHAs are concentrated on one or a few addresses while the rest work fine—the address is at fault. Quarantine it.
  • Problem across all addresses to a single site. If all addresses suddenly experience a drop in success rate to one specific target site, but other sites are fine—the pool is not to blame; it's that site. Quarantining addresses is pointless and harmful.
  • Problem across all addresses to all sites. Likely something on your side: network, infrastructure, the probe system itself. Again, don't blame the addresses.

Practical takeaway: metrics should be sliced not only by address but also by the address-site pair. Then a matrix of successes immediately shows where a whole row is red—the address is at fault—or where a whole column is red—the site is at fault. This simple two-dimensional view saves hours of debugging and prevents the pool from self-destructing over nothing.

Logs and Tracing

Beyond aggregated metrics, it's useful to log pool decisions: why this address was chosen, why that one was sent to quarantine. When something goes wrong, these decision traces become your lifeline. Don't log every request in detail—you'll drown. Log state transitions and atypical decisions.

Practice: A Pool Skeleton in Python and Node.js

Let's move from theory to code. We'll look at key implementation pieces on two popular platforms. These are skeletons, frameworks onto which you'll hang the specifics of your project. We'll omit details for clarity in the main ideas: the acquire and release interface, address selection, and result accounting.

Interface: acquire and release

Let's agree on the contract. The acquire method takes an optional key—for example, an account ID for sticky binding—and returns an address. The release method takes the address and the task result. The result is minimally described by two facts: was it a success, and was there a sign of CAPTCHA or block.

Skeleton in Python

Let's consider a simplified Python pool. We'll keep logic in memory for clarity, but note where a shared store would connect.

Key structural elements: a dictionary keyed by address, with a value object having fields for reputation, active lease count, quarantine exit timestamp, and current backoff duration. Separately, we store a sticky key-address binding table.

Pseudo-code for the acquire method in a Python-like language looks like this. First, check if there's a key. If a key is given and there's an active binding and the bound address is healthy—return it, extending the binding's TTL. If no binding exists—choose an address using consistent hashing of the key among healthy ones, save the binding with a TTL, return the address. If there is no key at all—apply the strategy for sessionless tasks, say weighted selection among healthy addresses. In all cases, increment the active lease count of the selected address.

Pseudo-code for release: decrement the active lease count. Update the sliding reputation window based on the result—add success or failure, separately note CAPTCHA flag. If the accumulated signals cross the failure threshold, taking hysteresis into account—move the address to quarantine: set the state, compute the backoff time as base times two to the power of repeated failures, capped at the ceiling, add jitter, set TTL or return time marker. If the result is good and the address has been stable for long—reset the quarantine failure counter.

A separate background loop periodically goes through addresses in quarantine whose backoff has expired, and runs an active probe. If the probe succeeds, with hysteresis requiring a certain number of consecutive successes—move the address to a half-open state with low weight, then gradually to healthy. If it fails—extend quarantine with an increased backoff.

Important practical Python details: use asynchronicity if you have many parallel tasks; wrap access to shared structures with appropriate synchronization; when moving to multiple processes, replace internal dictionaries with calls to Redis through atomic commands and scripts. The active lease counter fits nicely with atomic increment; key-address binding fits a record with TTL; sets of addresses by state can be conveniently held in sorted sets where the address's score is its rating.

Skeleton in Node.js

In Node.js, the asynchronous model is natural, and the pool is usually implemented as a class with async acquire and release methods returning promises. The idea is the same; the idiom changes.

State structure: a dictionary object of addresses, where values contain reputation as a ring buffer of recent results, active lease count, and quarantine fields. For multiple workers—and in Node this is often a cluster of processes—state is also pushed to Redis because each process is isolated.

The acquire method in Node: asynchronously check for a sticky binding via the store. If there is a live and healthy binding—return it and extend the TTL. Otherwise, select an address—for a key use consistent hashing, for a sessionless task use weighted selection. Atomically increment the lease counter. Return the address.

The release method in Node: atomically decrease the counter. Write the result to the reputation window. Check thresholds with hysteresis, and on failure set quarantine with exponential backoff and jitter, while respecting the overall limit on the number of quarantined addresses—if the limit is reached, reduce weight instead of quarantining.

Background checks in Node are easily implemented via a periodic timer that picks addresses with expired backoff and runs active probes, spreading them out in time. Successful probes smoothly return the address, increasing its weight in steps.

General advice for both platforms: don't try to make it perfect on the first try. Start with round-robin plus passive health check plus simple in-memory quarantine. Make sure the acquire and release interface fits your business logic. Then sequentially add sticky via hashing, active probes, shared store, limits, and observability. Let each layer mature under real load.

Mini Implementation Checklist

  • Interface reduced to acquire (with optional key) and release (with result).
  • All state complexity hidden inside the pool; business code is unaware.
  • Sticky implemented via deterministic, preferably consistent, hashing.
  • Counters and bindings are atomic under multiple workers.
  • A background cycle for active probes of quarantined and idle addresses.
  • Metrics written on every release.

Common Mistakes: What Not to Do

We've covered the theory and built the skeleton. Now let's go over the rakes people step on again and again. Knowing these errors will save you weeks of debugging.

Shared Pool for Different Sites

It's tempting to have one big pool and route traffic through it to all target sites at once. Don't do that mindlessly. An address's reputation differs across sites. An address that works great for one may be blocked on another. If you mix everything into one pool and one set of statistics, you'll get a blurred picture where a good address for one task is unfairly punished for problems with another.

Correct: maintain reputation per address-site pair, and logically separate pools or subgroups for different directions. Then quarantine decisions are made precisely and fairly.

No Limit on Tasks per Address

Even a healthy address has limits. If the pool happily piles hundreds of simultaneous tasks onto a popular address, you create an anomaly: unnaturally high concentration of activity from one point. This both overloads the address and looks suspicious to the platform.

Introduce a ceiling on the number of simultaneous leases per address. When the ceiling is reached, the address is temporarily removed from consideration; traffic goes elsewhere. This simple constraint protects you from many troubles.

Rotation in the Middle of a Session

A cardinal sin when working with accounts. An account started an action on one address, then due to quarantine firing or pool size change, suddenly moves to another in the middle of the session. From the platform's perspective, the user teleported. That's one of the most obvious signs of automation.

Protection: while an active session is ongoing, the key-to-address binding must be inviolable, even if the address is slightly degraded. Change the binding only at session boundaries—upon natural completion or TTL expiry. And if the address dies completely and there's no alternative, it's better to gracefully end the session than to move it mid-stream.

Other Common Mistakes

  • Single error as a death sentence. Sending an address to quarantine due to one random glitch. Networks are unreliable; an isolated miss is normal. Only accumulated signals matter.
  • No hysteresis. Leads to flapping and jitter, as we discussed.
  • Fixed backoff instead of exponential. A persistently sick address keeps returning and ruining statistics.
  • No quarantine limit. An incident can put the entire pool into quarantine and stop the system.
  • Synchronous probes. All addresses probed at the same moment, creating load spikes.
  • State only in memory with multiple workers. Each process lives in its own reality, no coordination.
  • Stuck leases. Release never called due to crash; counter stays inflated forever, address considered permanently busy. Cured with TTL on lease.
  • Blind faith in a single control resource. It goes down, and the entire pool thinks it's dead.

Tools and Resources for Implementation

Let's gather a practical arsenal. What actually helps when building a pool.

State Stores

  • Redis. The primary choice for shared state. Atomic counters, sorted sets for weights and states, hashes for address metadata, built-in TTL, scripts for atomic complex logic. Almost ideal for our task.
  • Memcached. Simpler and lighter; suitable if you only need basic caching with counters, but loses to Redis in richness of data structures and atomicity of complex operations.

Ecosystem Libraries and Approaches

  • Python. Async stack for parallel tasks, Redis client with async and script support, ready-made consistent hashing implementations, and the circuit breaker pattern for half-open state.
  • Node.js. Idiomatic async classes, mature Redis clients, cluster model where shared state is mandatory, libraries for consistent hashing and circuit breaker implementations.

Observability

  • Time-series metrics system. For storing success rate, latency, and CAPTCHA per address and per address-site pair.
  • Dashboards. Visualize pool state distribution, address-site matrix, quarantine frequency. A dashboard lets you tell a bad address from a bad site at a glance.
  • Alerts. Notify when approaching quarantine limit, when mass failure occurs, when overall success rate drops below a threshold.

What to Build Yourself

A ready-made universal pool library covering all your nuances likely doesn't exist—the business logic of sessions is too specific. So the pool core is usually written from scratch, relying on the building blocks listed: store, hashing, metrics, circuit breaker pattern. The good news is that with a clean acquire/release interface, this core becomes compact and reusable across projects.

Use Cases and Results

Let's go through a few generalized scenarios showing how the principles from this article work in practice. Numbers are illustrative, but the proportions reflect typical dynamics.

Case One: Public Data Collection Without Sessions

Task: mass collection of publicly available information from independent pages. Started with naive random selection from a file. Success rate hovered around seventy percent because dead addresses got as much traffic as live ones.

Introduced a pool with weighted selection and passive health check. Address weight recalculated on a sliding window of success rate. Dead addresses quickly lost weight and almost stopped getting tasks. Added active probes for idle addresses and exponential quarantine.

Result: success rate rose from around seventy to over ninety percent. The number of futile attempts on dead addresses dropped dramatically. Main takeaway: even without sessions, adaptive reputation-based routing gives a noticeable efficiency boost.

Case Two: Working with Accounts and Sticky Sessions

Task: stable operation of many accounts where the stability of the network environment for each is critical. Initially used random address selection, and accounts constantly jumped between subnets. Outcome: a flood of verification requests and rising CAPTCHA rates.

Switched to deterministic binding via consistent hashing of the account ID, with bindings stored in Redis and TTL equal to the business session duration. Introduced a hard rule: no rotation in the middle of a session. Binding changed only at session boundaries.

Result: CAPTCHA rates for accounts dropped significantly, verification requests fell, account behavior looked stable and natural to the platforms. Takeaway: for antifraud, predictability of the network environment is worth more than any sophisticated rotation.

Case Three: Incident and Protection Against a Quarantine Avalanche

During a running system, the target platform suddenly tightened its checks. CAPTCHAs rained across all addresses at once. The pool, which had no quarantine limit, in the first version of this scenario put almost the entire pool into quarantine within minutes—the system effectively stopped.

After rework, three things were added: a limit on the proportion of addresses in quarantine, recognition of mass failure as a sign of a common problem, and a guaranteed minimum number of addresses in service. When the incident repeated, the pool recognized that failures correlated across all addresses to a single site, understood it wasn't the pool's fault, and didn't create an avalanche. It simply lowered weights and continued working through a minimum of resources until things normalized.

Result: instead of a complete stop, the system experienced the incident with degradation rather than failure. Takeaway: resilience is measured by behavior on the worst day, not the best.

Case Four: Diagnosis—Bad Address vs. Bad Site

A team spent weeks periodically quarantining healthy addresses and couldn't understand why. Metrics were kept only per address, not sliced by site. When a two-dimensional address-site matrix was added, the picture cleared instantly: the problem wasn't the addresses, but one finicky site that gave error spikes for all addresses simultaneously.

Logic was set so that failures correlating across a site column didn't penalize the addresses. Result: the number of false quarantines dropped to nearly zero, and the useful capacity of the pool grew without adding a single new address. Takeaway: proper metric slicing can be more valuable than any algorithm.

Frequently Asked Questions

Do I need a pool if I only have a few addresses?

Even with a handful of addresses, a pool pays off as soon as you have any load or any notion of a session. Passive health check and simple quarantine protect you from hammering a dead address. Sticky binding preserves account stability. Full consistent hashing and a shared store might be overkill for two addresses, but a basic in-memory pool with acquire/release interface is almost always worth having.

How do I choose the sticky session duration?

Base it on business meaning, not technical convenience. Ask: how long should a sequence of actions be perceived as a single continuous visit? For short scenarios, it could be minutes; for working under a profile over a period of activity, it could be tens of minutes or hours. Set that duration as the TTL of the binding, and extend it on each action within the session so active work doesn't get cut off mid-sentence.

What if the bound address dies in the middle of a session?

This is the most unpleasant case. The general rule is not to rotate in the middle of a session. If the address is degraded but still alive, continue through it until the session ends. If it's completely dead, it's preferable to gracefully end or pause the session rather than move it to another address and create a teleportation effect. Design your business logic so that ending a session is less painful than migrating it.

How often should active probes be run?

Don't probe healthy busy addresses—live traffic speaks for them via passive checks. Probe idle healthy addresses every tens of seconds to keep them ready. Probe quarantined addresses on the exponential backoff schedule. Always spread probes in time with jitter to avoid synchronous activity bursts.

Is Redis mandatory, or can I rely on memory?

If you have strictly one process and you're okay with losing state on restart, in-memory is acceptable at the start. As soon as a second worker appears, a shared store becomes practically mandatory—otherwise processes live in different realities without coordination. Redis is the most convenient choice due to atomic operations, data structures, and TTL. You can start with memory, but abstract the storage layer so that moving to Redis doesn't rewrite half your code.

How do I tell the difference between an address problem and a target site problem?

Maintain metrics per address-site pair, not just per address. If errors are concentrated on one address across all sites—the address is at fault. If errors occur across all addresses to one site—the site is at fault, and you shouldn't punish addresses. If errors occur across all addresses to all sites—look for the problem on your side: network, infrastructure, the probe system itself. This two-dimensional slicing is the most reliable way to get the correct diagnosis.

What if almost the entire pool goes into quarantine?

This should be impossible with proper safeguards. Set a limit on the proportion of addresses in quarantine at once. Guarantee a minimum number of addresses in service under all circumstances. Teach the pool to recognize a simultaneous mass failure as a sign of a common problem, not a fault of the addresses. When the limit is reached, switch from quarantining to simply reducing weights—better to work through slightly degraded resources than to stop completely.

What selection strategy should I use by default?

If you have no sessions and addresses are homogeneous—start with round-robin. If addresses vary in quality—weighted selection with dynamic weight based on metrics. If tasks have very different durations—least-connections. If there is account or session binding—consistent hashing by key, and that's non-negotiable because it defines stability in account work. In complex systems, combine strategies by task type.

Do I need to track CAPTCHAs separately in metrics?

Yes, definitely, and with a more sensitive threshold than for regular errors. CAPTCHAs are an early and precise signal of address reputation decline, often preceding outright failures. If you mix CAPTCHAs with general errors, you lose this leading indicator. A separate CAPTCHA rate per address allows you to reduce the weight of a problematic address even before it starts failing openly.

How do I avoid stuck leases when a worker crashes?

Don't rely solely on an explicit release call. Represent the lease as a record with a TTL slightly exceeding the maximum allowable task duration. If the worker crashes and doesn't return the address, the record expires automatically, and the active load counter recovers. Additionally, you can periodically reconcile counters with reality. This protects the pool from a slow capacity leak due to stuck leases.

Conclusion and Next Steps

We've come a long way. From the ugly truth about why a text file list of addresses breaks under real load, to a full-fledged engineering system for managing addresses inside an application. Let's sum up the key points.

A pool is not a list, but an object with behavior, hidden behind a simple acquire/release interface. The address selection strategy is driven by the task: round-robin for homogeneous, weighted for variable quality, least-connections for tasks of varying duration, and—critically important for account work—deterministic, preferably consistent, hashing by key. It is the predictability of the network environment, not sophisticated randomness, that earns trust from antifraud systems.

Address health rests on two legs: passive control via real traffic and active probes. Failure is defined by accumulated signal, not a single error, and flapping is dampened with hysteresis and minimum hold times. Quarantine is built on exponential backoff with jitter, must be limited by a cap on the proportion of addresses removed, and protected from the catastrophe of the entire pool ending up on the bench. State in multi-worker environments lives in a shared store with atomic operations and TTL. And observability sliced by address-site pairs lets you distinguish a bad address from a bad site—a skill that saves weeks.

What to do next? Here's a concrete step-by-step implementation plan.

  1. Wrap your current address handling in an acquire/release interface, changing nothing inside yet. This prepares the ground.
  2. Add passive health check and simple in-memory quarantine. You'll already feel a stability boost.
  3. Introduce the needed selection strategy. If you work with accounts, immediately lay the foundation for sticky via key-based hashing.
  4. Add active probes and exponential backoff with protective limits.
  5. Move state to a shared store as soon as a second worker appears. Build in atomicity and TTL.
  6. Set up metrics and dashboards, always with the address-site slice.
  7. Fine-tune hysteresis and thresholds based on your real data—there are no universal numbers here, only your observations.

Don't try to build everything at once. Let each layer mature under load, gather metrics, draw conclusions, and move on. A mature proxy pool is not a one-time build but a living system you tune as your tasks grow. But even the first steps from this guide will turn a fragile list of addresses into a reliable backbone for your application. And that, you'll agree, is worth the effort.