RepoDaily · 2026-08-18 · Self-hosted app

Scrapling review: adaptive Python scraping you self-host as a library, Docker image, or MCP server

#10 Self-hosted app Python +338 D4Vinci/Scrapling Open repository

Scrapling 0.4.14 pairs an adaptive parser that survives site redesigns with stealth fetchers built to bypass Cloudflare Turnstile, plus a spider engine, an MCP server, and a Docker image you self-host.

Repo typeSelf-hosted app
Best forPython 3.10+ developers who need selectors that survive site redesigns, fetchers that clear Cloudflare Turnstile, and a self-hosted MCP server for AI agents
Risk levelMedium — strict CI and 90–92% test coverage, but version 0.4.14 is Beta and maintained by one author
Time to evaluate1–2 hours: fetch one page with StealthyFetcher, test the auto_save/adaptive selector pair, and register scrapling-mcp with an MCP client

Primary question: Does Scrapling's adaptive relocation keep your extractions alive through a target-site redesign better than your current selector stack?

90/100

RepoDaily adoption score

RepoDaily rates this as 90/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
100Evidence quality

7 source(s) across 4 source category/categories, plus a RepoDaily-specific evidence module when available.

100Installability

6 workflow step(s), 5 next-action step(s), and 3 command/install signal(s) were detected.

62Maintenance confidence

Trending momentum is +338 stars, with maintenance/release/issue signals counted when present.

93Production readiness

Risk is marked medium, with 5 security note(s) and 5 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 5 alternative(s), and 4 type-specific section(s) support differentiation.

82License clarity

License source or license wording is present.

78Agent / AI fit

5 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

Scrapling (D4Vinci/Scrapling) is a Python web scraping framework that, in its own words, handles "everything from a single request to a full-scale crawl." Version 0.4.14 ships as a BSD 3-Clause package for Python 3.10 through 3.13, written and maintained by Karim Shoair. Although it is a library at heart, this page treats it as a self-hosted app because of the deployment artifacts around it: a scrapling CLI entry point, a scrapling-mcp entry point that starts an MCP server for AI agents, and a pyd4vinci/scrapling image on Docker Hub — all of which you run inside your own infrastructure rather than consume as a vendor API.

The framework stacks three layers. The parser is adaptive: select elements once with auto_save=True, and when the site redesigns, re-running the same selector with adaptive=True lets the parser relocate the elements instead of returning empty results. The fetcher layer offers Fetcher for plain requests, StealthyFetcher for anti-bot evasion — the docs example calls StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True) and claims bypasses of systems like Cloudflare Turnstile out of the box — and DynamicFetcher for browser rendering, backed by curl_cffi, playwright, and patchright in the fetchers extra. The spider layer scales selection into concurrent, multi-session crawls with pause/resume and automatic proxy rotation, exposed through a Spider class whose async parse(response) yields dict items and starts with MySpider().start(). The docs also promise real-time stats and streaming during crawls.

Maturity signals cut both ways. On the reassuring side: CONTRIBUTING.md reports test coverage around 90–92%, CI runs tox across supported Python versions plus mypy, pyright, ruff, bandit, and vermin checks, and the project publishes nine translated READMEs (Arabic through Korean) and a Discord with a #help channel. On the caution side: the pyproject classifier still reads "Development Status :: 4 - Beta" with the Production/Stable line commented out, the author and maintainer fields name the same single person, and anti-bot countermeasures can invalidate behavior at any time regardless of code quality. It earned 338 stars in the 2026-08-18 trending window, landing at rank 10.

Problem it solves

  • CSS and XPath selectors break when target sites redesign; scrapers keep running but return empty or wrong fields.
  • Anti-bot systems such as Cloudflare Turnstile block plain HTTP clients and off-the-shelf headless browsers.
  • Growing from one page to a site-wide crawl normally means assembling a fetch library, a parser, concurrency, session handling, proxies, and pause/resume yourself.
  • AI agents that need page content usually lack a structured fetch-and-extract tool they can call under your control.

How it works

  1. Choose a fetcher from scrapling.fetchers: Fetcher for plain requests, StealthyFetcher for anti-bot evasion (the docs example sets StealthyFetcher.adaptive = True and fetches with headless=True, network_idle=True), DynamicFetcher for browser rendering.
  2. Select with CSS or XPath on the response object: response.css('.product') returns element objects, and the spider example extracts item.css('h2::text').get().
  3. Save the match signature once with auto_save=True so the parser remembers the traits of the elements you selected.
  4. After the site changes, re-run the same selector with adaptive=True; the parser relocates the elements instead of returning nothing.
  5. Scale to a crawl: subclass Spider from scrapling.spiders, set name and start_urls, implement async parse(self, response: Response) yielding dicts, and call MySpider().start() — the framework adds concurrent multi-session crawling, pause/resume, and automatic proxy rotation.
  6. Expose it to agents: the scrapling-mcp console script starts an MCP server, with mcp>=2.0.0 provided by the ai extra.

Command surface: two console scripts, four extras, one Docker image

pyproject.toml version 0.4.14 registers two entry points: scrapling = scrapling.cli:main and scrapling-mcp = scrapling.cli:mcp. The second starts the MCP server the README names io.github.D4Vinci/Scrapling, and it is the piece that turns the library into a self-hosted service for agent clients.

Dependency groups keep the base install light: the core requires only lxml>=6.1.1, cssselect>=1.5.0, orjson>=3.11.8, tld>=0.13.2, w3lib>=2.4.1, and typing_extensions. The fetchers extra adds click, curl_cffi>=0.16.0, playwright>=1.61.0, patchright>=1.61.2, browserforge>=1.2.4, apify-fingerprint-datapoints>=0.15.0, msgspec>=0.21.1, anyio>=4.14.0, and protego>=0.6.2. The ai extra layers mcp>=2.0.0 and markdownify on top of fetchers; shell adds IPython>=8.37 (noted as the last release supporting Python 3.10) and markdownify; all combines ai and shell.

For container deployment, the README links a Docker Hub image at pyd4vinci/scrapling, and a pyproject comment explains the version is pinned statically to improve Docker layer caching. That comment is a small but telling detail: the image is maintained deliberately, not bolted on.

Integration surface: Python API, MCP, agent skills, docs

  • Python API: scrapling.fetchers exposes Fetcher, StealthyFetcher, and DynamicFetcher; scrapling.spiders exposes Spider and Response; responses support .css() with auto_save and adaptive flags.
  • MCP: the ai extra pulls mcp>=2.0.0 and the scrapling-mcp console script starts the server, matching the mcp and mcp-server repo topics.
  • Agent skills: the repo carries an agent-skill directory, and a README badge links an OpenClaw skill at clawhub.ai/D4Vinci/scrapling-official.
  • Docs map: readthedocs pages cover selection methods (parsing/selection.html), choosing fetchers (fetching/choosing.html), and spider architecture (spiders/architecture.html), plus a changelog at the GitHub releases URL declared in pyproject.
  • Runtime: requires-python >=3.10 with classifiers for 3.10–3.13 on CPython, and Typing :: Typed signals type hints enforced by mypy and pyright in CI.

Maintenance read: Beta label, one maintainer, unusually strict gates

  • Version 0.4.14 still carries "Development Status :: 4 - Beta"; the Production/Stable classifier exists in pyproject.toml but is commented out.
  • Author and maintainer are both Karim Shoair ([email protected]), so bus factor is one.
  • Offsetting that: CONTRIBUTING.md cites roughly 90–92% test coverage, and every code-related commit runs tox on all supported Python versions through GitHub CI.
  • Code quality is enforced, not encouraged: mypy and pyright gate PRs, pre-commit hooks run ruff, bandit, and vermin, and failing checks mean rejection.
  • Process discipline: PRs must target the dev branch — anything against main is rejected — and AI-assisted contributions must disclose that fact per AI_POLICY.md.
  • Scope control: spider platform templates are accepted only when a platform exposes a uniform structure across many independent domains; single-site scrapers are explicitly excluded from the library.

Alternative matrix: where each option wins

  • Scrapy: the mature pick for scheduled, large-volume crawls with middleware and job persistence, but it does not market stealth fingerprinting; Scrapling lists "selenium-alternative" among its PyPI keywords.
  • BeautifulSoup: pure parsing of HTML you already fetched; no fetchers, no anti-bot handling, no spider scheduling.
  • Playwright (a Scrapling dependency via the fetchers extra) and its stealth-oriented fork patchright: browser automation toolkits you would otherwise wire together yourself; Scrapling packages them with fingerprint generation via browserforge and apify-fingerprint-datapoints behind one response API.
  • Commercial bypass APIs: README sponsor HyperSolutions sells bypass for Akamai, DataDome, Incapsula, and Kasada — the paid route for protections StealthyFetcher may not clear.
  • Managed scraping platforms: a vendor owns proxies and compliance for a usage fee, versus Scrapling's self-host-everything model under BSD 3-Clause.

Who should pay attention?

Good fit if

  • Python 3.10–3.13 codebases that currently glue together requests, lxml, and a headless browser.
  • Data pipelines that lose fields every time a target site ships a redesign and changes selectors.
  • AI agent builders who want a self-hosted MCP fetch-and-extract server instead of a third-party API.
  • Self-hosters who prefer the pyd4vinci/scrapling Docker image or the scrapling-mcp entry point over vendor scraping APIs.
  • Contributors: good first issue and help wanted labels exist, and the rules are written down in CONTRIBUTING.md.

Skip for now if

  • Projects that need a stable 1.0 API contract before depending on a framework.
  • Shops unwilling to audit a fetchers extra that adds ten packages to the dependency tree.
  • Targets whose terms of service or applicable law prohibit automated access — stealth tooling sharpens that risk.
  • Non-Python stacks; there is no binding outside CPython 3.10–3.13.
  • Anyone wanting a vendor to own proxy pools, CAPTCHAs, and compliance end to end.

Risks and cautions

Medium

Engineering signals are strong — 90–92% test coverage, tox across supported Python versions, mypy/pyright/ruff/bandit gates — but the project is versioned 0.4.14 with a Beta classifier and a single author-maintainer, so bus factor is the main exposure.

  • The "Development Status :: 4 - Beta" classifier in pyproject.toml, with Production/Stable still commented out.
  • Author and maintainer fields name the same person, Karim Shoair.
  • The API surface is documented on a latest docs URL rather than frozen majors, so names like adaptive flags may still shift.
  • The anti-bot domain is adversarial: site-side countermeasures can break StealthyFetcher behavior at any time regardless of code quality.
  • Scraping third-party sites can violate terms of service or local law; the stealth focus (Cloudflare Turnstile bypass per docs/index.md) raises the stakes of each target you choose.
  • The fetchers extra widens your supply chain with curl_cffi, playwright, patchright, browserforge, apify-fingerprint-datapoints, msgspec, anyio, protego, and click — pin and audit those pins.
  • CONTRIBUTING.md shows bandit runs on every commit via pre-commit and on every PR via CI, so static security checks gate contributions.
  • The MCP server exposes fetch capability to agent clients; treat the scrapling-mcp endpoint like any networked service and restrict which URLs it may request.
  • Licensing is unambiguous: BSD 3-Clause (copyright 2024, Karim Shoair) permits internal and commercial use with notice retention.

Alternatives to compare

ApproachWhen to useTrade-off
Scrapy
Scheduled, high-volume crawls where stealth evasion is not required and you want a long-stable spider frameworkFree, open source
BeautifulSoup
Parsing HTML you have already fetched; no fetching, anti-bot, or crawl scheduling neededFree, open source
Browser automation and testing without packaged stealth fingerprinting — Scrapling itself builds on it via the fetchers extraFree, open source
Selenium
Legacy browser automation stacks and broad driver compatibilityFree, open source
Commercial scraping APIs (e.g., HyperSolutions)
Paying a vendor to handle Akamai, DataDome, Incapsula, or Kasada bypass instead of self-hostingCommercial, usage-based

What this trend reveals

Adaptive selectors as breakage insurance

The auto_save-then-adaptive sequence means a redesign does not force hand-rewriting selectors: the parser relocates matched elements. Data pipelines that break weekly on CSS churn can convert that into fewer patch releases.

Fetch one target page, select with css('.product', auto_save=True), rename the class in a saved local copy of the HTML, then re-run with adaptive=True and measure whether the same fields are recovered.

A self-hosted MCP fetch tool for agents

The scrapling-mcp console script plus the mcp>=2.0.0 dependency gives an agent client a fetch-and-extract call you host yourself, with no vendor API key in the middle.

Register scrapling-mcp with an MCP client using the io.github.D4Vinci/Scrapling name from the README, fetch one static page, and compare the returned content against the Python API output.

Consolidate a scattered fetch stack

Many setups juggle a request library, lxml, and a separate headless browser. Scrapling's three fetcher classes behind one response API with .css(auto_save=...) cover that range in a single dependency set pinned by pyproject.toml.

Port two existing scripts — one static HTML, one behind anti-bot — to Fetcher and StealthyFetcher, then count the dependencies you removed.

Best next action

Run a two-fetcher spike against one real target

Confirm the two claims that matter most — stealth fetching and adaptive relocation — on a single page before wiring Scrapling into anything scheduled.

  1. Set up a virtual environment on Python 3.10–3.13 and install Scrapling with the fetchers extra declared in pyproject.toml; add the ai extra if you want the MCP server.
  2. Fetch your target once with StealthyFetcher.fetch(url, headless=True, network_idle=True), mirroring the docs/index.md example.
  3. Run page.css('<your selector>', auto_save=True) on the fields you need and store the result.
  4. Edit a local copy of the HTML to move or rename those elements, re-select with adaptive=True, and record whether the same data comes back.
  5. If you use AI agents, start the scrapling-mcp console script and check it completes a fetch-and-extract request before any wider rollout.

RepoDaily verdict

Scrapling is one of the few scraping projects that bundles adaptive selectors, stealth fetching, a spider engine, and an agent-facing MCP server under a BSD 3-Clause license you deploy yourself. Treat the 0.4.14 Beta label and single-maintainer reality as reasons to pin versions and keep tests around your extractions — then let auto_save/adaptive earn their keep the next time your target site redesigns.

Sources