Technical Blog

A Python scraper often starts as a short request followed by a selector. When DataDome verification interrupts retrieval, adding more conditions to that same function makes the program harder to maintain. A small access adapter gives you a place to integrate a confirmed API workflow without rewriting extraction, scheduling or storage.

Define the boundary as a return contract

The caller should receive either expected content or a reason it cannot continue. That is more useful than returning an empty string on every error. Keep transport diagnostics separate from a parser result: a request timeout is not a missing product description.

Python laptop and API adapter cartridge with a DataDome puzzle badge

The following Python is an internal interface example. It does not call a real solver or describe Scrapingbypass response fields. Implement the adapter using the documented API configuration after target compatibility has been confirmed.

from dataclasses import dataclass
from typing import Literal

@dataclass
class AccessResult:
    kind: Literal["content", "challenge", "denied", "error"]
    body: str = ""
    diagnostic_id: str | None = None

def collect(url, adapter, parser):
    result = adapter.fetch(url)
    if result.kind != "content":
        return {"collected": False, "reason": result.kind}
    record = parser(result.body)
    return {"collected": True, "record": record}

The parser still needs its own field validation before storage. Keeping the example small makes its boundary visible; it is not a complete production collector. Your adapter should reject unexpected redirects or bodies rather than assigning every response the content outcome.

Give the adapter explicit configuration

Load secrets from the server environment, not source files or query strings that enter logs. Validate required settings when the worker starts. Separate the approved target configuration from credentials so that an accidental URL change cannot silently redirect a collection job to another resource.

Set both connection and response time limits as appropriate for the chosen client and API mode. Timeouts should become typed failures that the scheduler can classify. A worker must not wait indefinitely or retry a denied resource because all exceptions were collapsed into one generic error.

Session ownership is an architectural choice

Requests Session can preserve local state, but it does not automatically implement a remote service’s session model. Determine which task owns the approved proxy and cookie context. Keep unrelated tasks isolated; avoid a process-wide mutable cookie container used by concurrent workers.

After a task ends, release its state according to the configured lifecycle. Do not export successful verification state into unrelated targets. This makes cleanup predictable and prevents later jobs from depending on undocumented leftovers.

Test with fixtures before a larger run

Use saved, sanitized examples of expected HTML, a verification page, a denial response and malformed data. Exercise the adapter mapping and the parser independently. Check that a verification fixture never reaches storage, and that a parser mismatch does not erase a valid older observation.

Then run one authorized live target through the confirmed Scrapingbypass DataDome integration. Verify the expected identity and fields, not just the absence of an exception. If the live response differs from your fixture, update the mapping deliberately instead of broadening every success condition.

This approach preserves your investment in the scraper. The access adapter is replaceable, diagnostics remain explicit, and business code no longer needs to know every detail of the verification step.

Trial Offer
+ 200 API Credits
+ Rotating Proxies
Claim Now ›