How to Export Large Data Volumes Through a Proxy Without Starting Over After a Disconnect
Table of contents
- Introduction: why long exports almost always break, and that's normal
- Preliminary preparation: tools, access, and environment
- Basic concepts: a resilient export glossary in plain words
- Step 1: resumable downloads via http range header
- Step 2: checkpoints for paginated exports
- Step 3: idempotency so retries don't create duplicates
- Step 4: deduplicating results without blowing up memory
- Step 5: parallelism without losses
- Step 6: resuming after a long pause
- Step 7: a ready-made resilient downloader skeleton in python
- Verifying the result: a resilient export checklist
- Common mistakes and their solutions
- Additional capabilities and optimization
- Faq: common questions about resilient export
- Conclusion: what you now know and where to go next
Introduction: Why Long Exports Almost Always Break, and That's Normal
If you've ever launched a big data export, you know the feeling. The process ran for hours, hit ninety percent, and then broke. The connection dropped, the server returned an error, the laptop went to sleep. And you have to start all over again. This guide is written so that this never happens again.
What you'll get out of it. You'll learn to build a downloader that survives disconnects. It resumes files from the middle, remembers which page it stopped on, doesn't create duplicates on retry, and can resume even after a long pause. You'll get a ready-made Python skeleton you can adapt to your task.
Who this guide is for. For engineers, analysts, and developers who export data from APIs, download large files, or collect paginated results through a proxy. Intermediate level. You should understand the basics of HTTP and be able to read Python code. No deep knowledge of network programming is required.
What you need to know beforehand. Basic Python, the concept of an HTTP request and response, what headers and status codes are. If you've worked with the requests library, that's enough.
How much time it'll take. Reading and understanding the concepts will take about forty minutes. Building a working downloader from our skeleton will take one to three hours, depending on your data source.
An important clarification on the topic. We won't cover status codes like 429 or retry strategies with delays (backoff). There's a separate piece on that. Here the focus is on just one thing: process state and its resumption. How to save progress, how not to lose or duplicate data, how to continue from where you stopped.
Tip: Keep a notebook or a separate file handy where you'll jot down your source's parameters: whether it supports resumption, whether it has pagination, what its cursor format is. These notes will come in handy at every step.
Preliminary Preparation: Tools, Access, and Environment
Before writing code, let's set up the working environment. It'll take ten minutes but save you hours of debugging.
What to install
- Install Python version 3.10 or newer. Check the version with the command
python --versionin the terminal. - Create a virtual environment with the command
python -m venv venvso your project dependencies don't mix with system ones. - Activate the environment. On Windows it's
venv\Scripts\activate, on macOS and Linux it'ssource venv/bin/activate. - Install the HTTP request library with the command
pip install requests. - For faster work with the state database, nothing extra needs to be installed: the
sqlite3module is already part of the Python standard library.
What access you need
- Access to your data source: URL, token or API key if required.
- A Proxeon proxy with an address, port, and credentials. Without a stable proxy, resilient export loses its meaning, because it's the proxy that distributes the load and makes connections predictable.
- Disk space for the state file and the exported data itself.
Checking the Proxeon Proxy
- Take a connection string like
http://login:password@address:port. - Test it with a simple request. In the terminal, run
curl -x http://login:password@address:port https://api.ipify.organd make sure it returned the proxy's IP address, not your own.
Tip: Store the proxy connection string in an environment variable, not in the code. That way you won't accidentally push the password to version control. Read it in code via os.environ.
⚠️ Warning: Always work only with data sources you have legitimate access to. Follow the terms of service and the limits set by their owner. The Proxeon proxy is intended for legitimate engineering work: load distribution, connection stability, and correct data export.
✅ Check: The environment is ready if the command python -c "import requests, sqlite3" ran without errors and the request through the proxy returned the proxy server's address.
Basic Concepts: A Resilient Export Glossary in Plain Words
Before writing code, let's go over key terms. Without them, the following steps will sound like magic spells.
Resumable download
Resumable download is continuing a file download from the byte where it was interrupted. Instead of downloading the file all over again, you ask the server to send only the missing chunk. It works via the HTTP Range header.
Checkpoint
A checkpoint is a saved point of progress. Think of saving in a video game. If something goes wrong, you return to the last save, not the beginning of the game. In an export, a checkpoint stores which page or record you stopped at.
Cursor
A cursor is a marker that an API gives you so you can request the next batch of data. Often it's a string like eyJvZmZzZXQiOjEwMH0. You send it back, and the server knows where to continue from.
Idempotency
Idempotency is a property of an operation where repeating it doesn't change the result. If you wrote the same row twice with the same key, in the end there's one row, not two. This is protection against duplicates on retries.
Deduplication
Deduplication is filtering out repeating records. Even with careful work, the same object can come twice. Deduplication ensures that in your final dataset it remains only once.
The core principle
A resilient downloader is built on one idea: progress must be saved constantly, not just at the end. Any step could be the last before a disconnect. That means after each successful chunk of work, the state must be written to disk. Then resumption is just reading the state and continuing.
Tip: Remember the rule of three questions for any export. First: where did I stop? Second: how do I avoid duplicating what I've already got? Third: what will go stale while I'm away? The answers to these make up resilience.
Step 1: Resumable Downloads via HTTP Range Header
Goal of this stage. Learn to download a large file so that after a disconnect you continue from the un-downloaded byte, not from scratch.
How it works
HTTP lets you request not the whole file, but a part of it. To do this, a Range header is added to the request. For example, Range: bytes=1048576- means: give me everything starting from byte number 1048576. But first you need to make sure the server supports this.
- Send a HEAD request or a regular GET to the file and look at the response headers.
- Find the
Accept-Rangesheader. If its value isbytes, the server supports resumption. - If the header is absent or says
none, resumption isn't possible. In that case, you'll have to download the file in one go or look for an alternative source.
Checking resumption support
Here's code that checks whether the server can serve file parts.
import requests
def supports_resume(url, proxies):
resp = requests.head(url, proxies=proxies, timeout=30, allow_redirects=True)
accept = resp.headers.get("Accept-Ranges", "none")
total = resp.headers.get("Content-Length")
return accept.lower() == "bytes", totalResuming a file from the middle
Now the main code. It looks at how many bytes have already been downloaded locally and asks the server for only the remainder.
import os
import requests
def download_resumable(url, dest, proxies):
already = 0
if os.path.exists(dest):
already = os.path.getsize(dest)
headers = {}
if already > 0:
headers["Range"] = f"bytes={already}-"
mode = "ab" if already > 0 else "wb"
with requests.get(url, headers=headers, proxies=proxies,
stream=True, timeout=60) as r:
if already > 0 and r.status_code == 200:
mode = "wb"
already = 0
with open(dest, mode) as f:
for chunk in r.iter_content(chunk_size=65536):
if chunk:
f.write(chunk)
return os.path.getsize(dest)Let's break down the important moments. If the server returned status 206, it honestly served part of the file, and appending will be correct. If the server returned 200 despite the Range header, it ignored resumption and is serving the whole file. In that case, we switch to full overwrite mode so we don't glue the old chunk to the new one and corrupt the file.
⚠️ Warning: Never append data in ab mode unless you're sure the server responded with code 206. Otherwise you'll get a corrupted file where the beginning is the remainder from a previous attempt and the continuation is the new full file. Such a file will fail to open, and you'll waste time finding the cause.
Tip: Download not directly to the target file, but to a temporary file with a .part extension. When the download fully completes, rename it to the final name. That way you'll never confuse a finished file with a partially downloaded one.
Integrity check
After a full download, it's good to verify the file wasn't corrupted. If the server sent a Content-Length header, compare it with the actual file size on disk. If the sizes match, the file came through whole.
def verify_size(dest, expected):
if expected is None:
return True
return os.path.getsize(dest) == int(expected)✅ Check: Interrupt the download halfway by closing the program. Run it again. In the logs you should see that the request went out with a Range header and the file was appended to, not started over. The final size matches the expected one.
Step 2: Checkpoints for Paginated Exports
Goal of this stage. Set up progress saving for APIs that serve data in pages, so after a disconnect you continue from the right page.
What exactly to save
A file is resumed by bytes, but a paginated export works on completely different logic. There are no bytes here, there are pages and records. That means the checkpoint needs to store something else.
- Cursor if the API works on cursors. This is the most reliable option, because the cursor itself knows where to continue from.
- Page number or offset if the API works on offset and limit. Store the number of the last successfully processed page.
- Last record ID if it can be sorted by ascending ID or date. Then the next request asks for records with an ID greater than the saved one.
- Processed record count for control and reporting.
Where to store state
You have three main options, from simple to reliable.
- JSON file. The simplest. You write a dictionary with the cursor and count to a file after each page. Suitable for single, non-parallel exports.
- SQLite database. More reliable. Provides transactions, so state won't get corrupted if the process crashes mid-write. Good when there's a lot of data and deduplication is needed.
- External database. For large distributed exports, when multiple processes share one job.
Saving a checkpoint to JSON
import json
import os
def save_checkpoint(path, cursor, page, last_id, count):
tmp = path + ".tmp"
data = {
"cursor": cursor,
"page": page,
"last_id": last_id,
"count": count,
}
with open(tmp, "w") as f:
json.dump(data, f)
os.replace(tmp, path)
def load_checkpoint(path):
if not os.path.exists(path):
return {"cursor": None, "page": 0, "last_id": None, "count": 0}
with open(path) as f:
return json.load(f)Note the trick with the temporary file. We write to a file with the .tmp suffix, then atomically rename it via os.replace. This protects against the situation where the program crashes right during the checkpoint write. The old checkpoint stays intact rather than turning into half a JSON that can't be read.
⚠️ Warning: Never write the checkpoint directly to the same file over the old one without a temporary file. A crash in the middle of writing will leave you with a corrupted checkpoint, and resumption becomes impossible. Atomic replacement solves this completely.
Tip: Save the checkpoint only after the page data is actually written to storage. The order is: got the page, wrote the data, then updated the checkpoint. If you swap the order, on a disconnect you'll skip a page and lose data.
The main loop with checkpoints
def paginate(fetch_page, save_data, cp_path, proxies):
cp = load_checkpoint(cp_path)
cursor = cp["cursor"]
count = cp["count"]
while True:
items, next_cursor = fetch_page(cursor, proxies)
if not items:
break
save_data(items)
count += len(items)
last_id = items[-1].get("id")
save_checkpoint(cp_path, next_cursor, cp["page"] + 1,
last_id, count)
cursor = next_cursor
if next_cursor is None:
break
return count✅ Check: Start the export, let it process several pages, interrupt it. Open the checkpoint file and make sure it contains the current cursor and count. Start it again: the export should continue from the saved cursor, not from the first page.
Step 3: Idempotency So Retries Don't Create Duplicates
Goal of this stage. Make sure a re-run or a repeat of a specific request doesn't lead to identical records appearing in your storage.
Why duplicates occur
Imagine: you got a page of data, wrote it to a file, but the program crashed before the checkpoint was updated. On the next run you'll request the same page again. The data comes again and gets written a second time. That's how duplicates are born. It's an inevitable consequence of disconnects, and you need to combat it at the architecture level.
Deduplication key
The main tool of idempotency is the deduplication key. It's a field or combination of fields that uniquely identify a record. Choosing the right key solves half the problem.
- Natural ID. If a record has a unique identifier from the source, use it. This is the ideal key.
- Combination of fields. If there's no single ID, build a key from several stable fields. For example, email plus registration date.
- Content hash. If there are no stable fields at all, compute a hash of the entire record. This is a last resort, because any field change creates a new key.
Writing without duplicates via UPSERT
If you store the result in SQLite or another database, use insert with conflict ignore. Then a repeated write with the same key simply does nothing.
import sqlite3
def init_db(path):
conn = sqlite3.connect(path)
conn.execute(
"CREATE TABLE IF NOT EXISTS records ("
"dedup_key TEXT PRIMARY KEY, payload TEXT)"
)
conn.commit()
return conn
def save_records(conn, items):
rows = [(item["id"], json.dumps(item)) for item in items]
conn.executemany(
"INSERT OR IGNORE INTO records (dedup_key, payload) "
"VALUES (?, ?)", rows
)
conn.commit()The key detail here is the PRIMARY KEY on the dedup_key field. The database itself will reject a repeated insert with the same key, because INSERT OR IGNORE will silently swallow the conflict. You don't need to manually check whether such a record already exists. The database does it for you, and does it fast.
Tip: Choose the deduplication key once at the start of the project and lock it into the documentation. Changing the key mid-export means old and new records stop matching, and duplicates will still appear. The stability of the key matters more than its elegance.
✅ Check: Run the export twice in a row over the same data range. Count the number of rows in the database with the command SELECT COUNT(*) FROM records. The number should be the same after the first and after the second run.
Step 4: Deduplicating Results Without Blowing Up Memory
Goal of this stage. Filter out repeating records across millions of rows without loading the computer's memory with a mass of already-seen keys.
The naive approach and its problem
The simplest deduplication: keep a set of all seen keys in memory. For each new record, check if the key is in the set. Works great on hundreds of thousands of rows. But on millions and tens of millions, the set grows and eats gigabytes of RAM. The program slows down or crashes.
Solution one: rely on the database
The simplest and most reliable way on large volumes is to not store seen items in memory at all, but to trust the check to the database via PRIMARY KEY, as we did in the previous step. The database stores the index on disk, not in your process's memory. It'll handle tens of millions of keys without load on your RAM.
Solution two: record hash
When there's no natural key, compute a compact hash from the record. The hash takes a fixed and small amount of space regardless of the record's size.
import hashlib
import json
def record_hash(item):
raw = json.dumps(item, sort_keys=True, ensure_ascii=False)
return hashlib.sha256(raw.encode("utf-8")).hexdigest()The sort_keys=True parameter is critical here. It ensures that records identical in content produce an identical hash, even if their fields came in a different order. Without this sorting, two identical objects could get different hashes and slip through as different records.
Solution three: Bloom filter for memory savings
If you still need a fast in-memory check on huge volumes, use a Bloom filter. It's a structure that takes little space and quickly answers whether we've seen a key or definitely haven't. It has a quirk: it can occasionally falsely say a key was already seen when it wasn't. That's why a Bloom filter is used as a fast preliminary filter, while the final check is left to the database.
- Check the key with the Bloom filter.
- If the filter says we definitely haven't seen it, write to the database right away.
- If the filter says we possibly have, do an exact check in the database.
⚠️ Warning: Don't try to deduplicate tens of millions of rows with a regular in-memory set. On a typical laptop this will lead to memory exhaustion and the process crashing halfway through the export. Move the load to disk via a database or use a Bloom filter.
Tip: If you export data in batches and duplicates are possible within a single batch, deduplicate the batch in memory with a regular set before writing to the database. The batch is small, memory won't suffer, and fewer redundant inserts will go to the database.
✅ Check: Run deduplication on a large test set with deliberate repeats. Verify that the final number of unique records is correct and that the process's memory consumption stays stable and doesn't grow linearly with the number of rows.
Step 5: Parallelism Without Losses
Goal of this stage. Speed up the export with parallel requests without losing a single task and correctly retrying the failed ones.
Task queue
The basis of safe parallelism is a task queue. You split the work into independent pieces in advance. For example, a list of pages or ranges. You put them in a queue. Several workers take tasks from the queue, execute them, and put the result down. If a worker crashes, its task can be returned to the queue and given to another.
Limiting concurrency
You can't launch an infinite number of parallel requests. That'll overload the source and your proxy. The right approach is to limit the number of simultaneous workers to a reasonable value. Start with a small number and increase it while watching stability.
from concurrent.futures import ThreadPoolExecutor, as_completed
def run_parallel(tasks, worker, proxies, max_workers=5):
results = []
failed = []
with ThreadPoolExecutor(max_workers=max_workers) as pool:
future_map = {
pool.submit(worker, t, proxies): t for t in tasks
}
for future in as_completed(future_map):
task = future_map[future]
try:
results.append(future.result())
except Exception:
failed.append(task)
return results, failedRetrying failed tasks
The collected failed list isn't lost data, it's a list of what needs to be retried. After the first pass, you run the failed tasks again. Usually that's enough to finish off the remainder.
def run_with_retry(tasks, worker, proxies, rounds=3):
remaining = tasks
for _ in range(rounds):
done, remaining = run_parallel(remaining, worker, proxies)
if not remaining:
break
return remainingThe role of the Proxeon proxy in parallelism. In parallel work, the proxy distributes connections, which makes the export more stable and predictable. Each worker runs through its own connection, and the load doesn't concentrate at a single point.
⚠️ Warning: Parallel writing to one file or one checkpoint creates data races. Two workers can overwrite each other's state. Write results only to a database with transactions, or use a separate file per worker and assemble the summary checkpoint in a separate thread.
Tip: Make tasks small and independent. If one task covers too large a range, its failure discards a lot of work. Small tasks are retried cheaply and almost unnoticed.
✅ Check: Launch a parallel export, deliberately crash some workers. After the retry rounds, the remaining list should become empty and the final dataset complete. Compare the number of records received with the expected one.
Step 6: Resuming After a Long Pause
Goal of this stage. Correctly continue the export if a lot of time has passed between attempts, and understand what could have gone stale during that period.
What goes stale over time
A disconnect for a minute and a pause for a day are different situations. During a long pause, part of your state may become invalid.
- Session. Many services keep a session for a limited time. After a long pause, the server will forget it, and requests will start returning an authorization error.
- Access token. API tokens often have a lifetime of minutes or hours. An expired token needs to be refreshed before continuing.
- Cursor. Some cursors don't live long. If the cursor has expired, you'll have to start from the nearest stable point, for example by the last record ID.
- The data itself. During the pause, new records could have appeared in the source or old ones changed. This affects offsets in offset-based pagination.
Safe resumption strategy
- At startup, check the checkpoint's age. If it's old, be prepared for part of the state being outdated.
- Refresh the access token and create a new session before the first request. Don't rely on the old ones.
- Prefer resumption by the last record ID, not by page number. The ID is stable, while the page number shifts if the data changed.
- Make a test request with the saved cursor. If it returned an invalid cursor error, switch to resumption by last_id.
def resume(cp, fetch_by_id, fetch_by_cursor, proxies):
if cp["cursor"]:
try:
return fetch_by_cursor(cp["cursor"], proxies)
except CursorExpired:
pass
return fetch_by_id(cp["last_id"], proxies)Why resumption by ID is more reliable. Imagine you stopped on page 50 when sorting by date. While you were away, new records were added at the beginning. Now page 50 contains completely different data, and you'll skip some records. Resumption by the last record's identifier doesn't suffer from this: you simply ask for everything greater than the saved ID.
Tip: Always save both the cursor and the last record ID in the checkpoint. The cursor is faster, but the ID is your safety rope in case the cursor expires during a long pause.
✅ Check: Stop the export, wait long enough for the token or cursor to expire, and start it again. The downloader should refresh the token, detect the expired cursor, and continue by ID without losing or duplicating records.
Step 7: A Ready-Made Resilient Downloader Skeleton in Python
Goal of this stage. Assemble everything you've learned into a single working framework that you adapt to your data source.
Below is a framework combining checkpoints, deduplication via a database, token refresh, and resumption. You plug in your own page-fetching functions for the specific API.
import os
import json
import sqlite3
import requests
class ResilientLoader:
def __init__(self, cp_path, db_path, proxies):
self.cp_path = cp_path
self.proxies = proxies
self.conn = sqlite3.connect(db_path)
self.conn.execute(
"CREATE TABLE IF NOT EXISTS records ("
"dedup_key TEXT PRIMARY KEY, payload TEXT)"
)
self.conn.commit()
def load_cp(self):
if not os.path.exists(self.cp_path):
return {"cursor": None, "last_id": None, "count": 0}
with open(self.cp_path) as f:
return json.load(f)
def save_cp(self, cp):
tmp = self.cp_path + ".tmp"
with open(tmp, "w") as f:
json.dump(cp, f)
os.replace(tmp, self.cp_path)
def save_records(self, items):
rows = [(str(i["id"]), json.dumps(i)) for i in items]
self.conn.executemany(
"INSERT OR IGNORE INTO records "
"(dedup_key, payload) VALUES (?, ?)", rows
)
self.conn.commit()
def run(self, fetch_page):
cp = self.load_cp()
while True:
items, next_cursor = fetch_page(
cp["cursor"], cp["last_id"], self.proxies
)
if not items:
break
self.save_records(items)
cp["count"] += len(items)
cp["last_id"] = items[-1]["id"]
cp["cursor"] = next_cursor
self.save_cp(cp)
if next_cursor is None:
break
return cp["count"]An example page-fetching function for your source. Here you implement the request logic and response parsing.
def fetch_page(cursor, last_id, proxies):
params = {"limit": 100}
if cursor:
params["cursor"] = cursor
elif last_id:
params["after_id"] = last_id
r = requests.get(
"https://example-source/api/records",
params=params, proxies=proxies, timeout=60
)
r.raise_for_status()
data = r.json()
return data["items"], data.get("next_cursor")Running the whole mechanism looks simple.
proxies = {
"http": os.environ["PROXEON_URL"],
"https": os.environ["PROXEON_URL"],
}
loader = ResilientLoader("state.json", "out.db", proxies)
total = loader.run(fetch_page)
print("Total records:", total)Tip: Add logging every hundred records to the loop: time, counter, current cursor. That way you'll see progress and easily understand if the export is stuck at one spot.
✅ Check: Run the framework on a real source, interrupt it halfway, run it again. The final number of records after resumption will match the full number of records in the source, and a repeat run won't increase the unique row count.
Verifying the Result: A Resilient Export Checklist
Go through this list. If all points are met, your downloader is truly resilient.
- File resumption continues from the un-downloaded byte, not from scratch.
- The downloader correctly handles the case where the server ignores the Range header.
- The checkpoint is saved after each processed page, not just at the end.
- The checkpoint is written atomically via a temporary file and rename.
- A repeat run doesn't create duplicates in the final dataset.
- Deduplication doesn't grow in memory linearly with the number of rows.
- Parallel workers don't lose failed tasks and retry them.
- After a long pause, the token is refreshed, and an expired cursor is replaced by resumption by ID.
How to test
- Run a full export of a small dataset and remember the number of records.
- Run it again on the same dataset and make sure the number hasn't changed.
- Interrupt the export at different points: at the beginning, in the middle, closer to the end.
- After each interruption, restart and verify the result is the same.
Success metrics. The number of unique records is stable across runs. Memory consumption doesn't grow uncontrollably. Resumption always continues from the saved point. There are no corrupted files after resumption.
Common Mistakes and Their Solutions
Problem: the file won't open after resumption. Cause: data was appended in ab mode even though the server responded with code 200 and served the whole file. Solution: check the response status; on 200, switch to fully overwriting the file from scratch.
Problem: after a disconnect, the export starts from the first page. Cause: the checkpoint was saved only at the end of the work or not saved at all. Solution: save the checkpoint after each processed page, right after writing the data.
Problem: the checkpoint can't be read, the JSON is corrupted. Cause: the program crashed during a write directly to the target file. Solution: write to a temporary file and replace atomically via os.replace.
Problem: there are duplicates in the final dataset. Cause: there's no deduplication key, or it's unstable. Solution: set a PRIMARY KEY on a reliable key and use INSERT OR IGNORE.
Problem: the process crashes from out-of-memory on large volumes. Cause: all seen keys are held in an in-memory set. Solution: move the uniqueness check to the database or apply a Bloom filter.
Problem: after a long pause, requests return an authorization error. Cause: the token or session expired during the break. Solution: refresh the token and create a new session on every start.
Problem: after a pause, some records are skipped or duplicated. Cause: resumption went by page number while the data in the source changed. Solution: resume by the last record's identifier, not by offset.
Problem: parallel workers lose some data. Cause: several workers wrote to one checkpoint and overwrote each other. Solution: write results to a database with transactions, not to a shared state file.
Additional Capabilities and Optimization
Batched writing
Don't write to the database one row at a time. Collect a batch of several hundred records and insert them at once via executemany. This speeds up writing many times over on large volumes.
Periodic commits
Call commit not on every write, but once every few hundred rows. Too frequent a commit slows down the database. Too rare risks losing more data on a disconnect. Find the balance for your load.
Progress report
Add an estimate of the remaining time. Knowing the page processing speed and the total number of records, you can estimate how much longer to wait. This is convenient for long exports.
Separate storage for raw and processed data
Store raw responses separately from parsed records. If you later change the parsing logic, you won't have to re-download the data. It's enough to run the raw responses through the new parser.
Tip: Configure the Proxeon proxy so that connections are stable throughout the export. A stable connection reduces the number of disconnects, which means your downloader enters resumption less often and works faster.
FAQ: Common Questions About Resilient Export
How do I know whether the server supports file resumption? Send a HEAD request and look at the Accept-Ranges header. A value of bytes means support. A missing header or none means resumption isn't possible.
What if the API doesn't give a cursor, only pages? Save the page number and, if possible, the last record ID. Resume preferably by ID, because page numbers shift when data changes.
How often should I save the checkpoint? After each successfully processed and written page. That way, on a disconnect you lose at most one page of work, not the entire export.
Can I deduplicate without a database? On small volumes, yes, with a regular in-memory set. On millions of rows it's dangerous due to memory. Better to use a database with PRIMARY KEY or a Bloom filter.
What should I choose as the deduplication key? The source's natural unique ID, if it has one. If not, a combination of stable fields. As a last resort, a hash of the entire record with key sorting.
Why is resumption by ID more reliable than by page number? Because data in the source can change. New records shift pages, and by page number you'll skip or duplicate data. IDs don't depend on this.
How many parallel workers should I set? Start with a small number and increase it while watching stability and the source's limits. Excessive parallelism hurts more than it helps.
How do I store the proxy connection string securely? In an environment variable, not in the code. Read it via os.environ. That way the password won't end up in version control.
What if the cursor expires during a long pause? Catch the invalid cursor error and switch to resumption by the saved last record ID.
Should I verify the integrity of the downloaded file? Yes. Compare the actual file size with the Content-Length header. If the server provides a checksum, verify that too.
Conclusion: What You Now Know and Where to Go Next
You've gone from a fragile export that collapses at the first disconnect to a resilient downloader. Now you have all the tools to make a disconnect stop being a catastrophe and become an ordinary working situation.
What you've mastered. File resumption via the Range header with Accept-Ranges verification. Progress saving via atomic checkpoints. Idempotency at the level of the deduplication key. Deduplication across millions of rows without blowing up memory. Parallelism with retry of failed tasks. Resumption after a long pause with token refresh and replacement of expired cursors. And most importantly, a ready-made Python framework that brings it all together.
What to do next. Take your real data source and adapt the page-fetching function for it. Start with a small volume, debug resumption on interruptions, then scale up. Set up a stable Proxeon proxy so that connections are predictable throughout the export.
Where to grow. Study batched writing and Bloom filters in more depth. Add progress monitoring and time estimation. Separate storage of raw and processed data. Gradually