프록시-저렴함
통합
Crawl4AI: Complete guide to AI web scraping with proxies

Crawl4AI Proxy Integration

Crawl4AI is an open source Python crawler that turns any web page into clean, LLM ready Markdown, JSON, or structured data in a single async call. This guide covers proxy setup, rotation, and the errors you will hit once you scrape at real scale, plus how to pick the right Proxy-Cheap product for the job.
Crawl4AI: Complete guide to AI web scraping with proxies용 프록시 받기
Crawl4AI: Complete guide to AI web scraping with proxies
What is Crawl4AI?
Crawl4AI is an open source Python web crawler built for LLM workflows. It wraps Playwright to render JavaScript, then converts the page into clean Markdown, sanitized HTML, structured JSON, or screenshots, all returned in one CrawlResult object. It is maintained by UncleCode, has crossed 50,000+ GitHub stars, and ships a self hostable Docker image with a REST API for production use.

Key takeaways:

  • Crawl4AI is a free, Apache 2.0 Python library that converts any web page into LLM-ready Markdown, JSON, or structured data in a single async call.
  • For crawls at scale, proxies keep response success rates stable, return geo-accurate data, and help you respect site rate limits and regional requirements.
  • Configure proxies per request through CrawlerRunConfig.proxy_config (a dict or a ProxyConfig instance), the approach Crawl4AI officially recommends. RoundRobinProxyStrategy cycles through a proxy list you supply; it does not generate new IPs itself. The actual IP rotation is delivered by Proxy-Cheap, based on your credentials and session settings.
  • Match the Proxy-Cheap product to the workload: Static Residential (ISP) for persistent sessions, ISP for trusted long-lived sessions, Datacenter for high-volume public crawls.

What is Crawl4AI?

Crawl4AI is an open-source Python web crawler purpose-built for LLM workflows. It wraps Playwright to render JavaScript, then converts the rendered DOM to clean Markdown, filtered Markdown, sanitized HTML, structured JSON, or screenshots. All of it comes back in one CrawlResult object.

The project is maintained by UncleCode and has crossed 50,000+ GitHub stars. It runs on Python 3.10+, ships an async-first API (AsyncWebCrawler, arun, arun_many), and offers a self-hostable Docker image with a REST endpoint and dashboard. The current stable series is v0.9.x.

Two things set it apart from general-purpose scrapers. First, the Markdown output is tuned for token-efficient ingestion into LLMs and vector databases. Second, the 0.8.x line added adaptive crawling that stops once it has gathered enough information, and v0.8.5 introduced automatic anti-bot detection that can escalate through a proxy list you supply. Both carry into the current 0.9.x line.

What is Crawl4AI used for?

Crawl4AI is used wherever a developer needs the public web turned into structured input for an AI model. The most common applications:

  • LLM training and fine-tuning data. Crawl thousands of documentation pages, blog posts, or product catalogs and store the Markdown for downstream training.
  • Retrieval-augmented generation (RAG). Convert websites into clean chunks for vector stores like Chroma, Qdrant, or pgvector.
  • AI agents. Give an autonomous agent a live read of the web with deterministic Markdown formatting.
  • Market and competitive research. Track competitor pricing, product launches, and content updates at scale.
  • Documentation mirroring. Pull entire docs sites into a single Markdown corpus for in-house chatbots.
  • Lead generation and sentiment analysis. Extract structured fields from listings, reviews, or social pages with the LLM extraction strategy.

The unifying thread is AI web scraping. Developers want clean text more than they want raw HTML, and they want it in a format their model already understands.

Key features of Crawl4AI

The features that matter for production AI scraping:

  • LLM-friendly Markdown with a fit_markdown variant that prunes boilerplate to reduce token cost.
  • Async-first architecture with AsyncWebCrawler and a MemoryAdaptiveDispatcher for batched runs via arun_many.
  • JavaScript rendering through Playwright, with wait_for, wait_until, custom js_code, and Shadow DOM flattening for Web Components.
  • Multiple extraction strategies: CSS, XPath, regex, cosine similarity, and LLMExtractionStrategy for schema-driven structured output.
  • Deep crawling with BFS, DFS, and best-first traversal, plus filter chains and relevance scorers.
  • Native proxy support via CrawlerRunConfig.proxy_config, including list-based rotation strategies (RoundRobinProxyStrategy) and anti-bot retry that escalates through a proxy list you provide.
  • Docker image with a REST API, monitoring dashboard, and built-in playground UI.

Is Crawl4AI free to use?

Yes. Crawl4AI is fully free and open source under the Apache License 2.0. There is no API key, no rate limit, and no usage cap built into the library. Commercial use is permitted; attribution via the project badges is recommended but not required.

You install it with pip, run it on your own infrastructure, and pay nothing to the project. Operational costs come from three places only:

  1. The compute that runs Chromium (Playwright is memory-heavy at high concurrency).
  2. Optional LLM API calls if you use LLMExtractionStrategy with a hosted model.
  3. Proxies, once you crawl at any serious volume.

A hosted Crawl4AI Cloud is in closed beta, but the entire feature set discussed in this guide works on the free, self-hosted library.

How to install Crawl4AI

Installation is two commands plus a verification step. Run them in a clean virtual environment on Python 3.10 or newer.

# 1. Install the library
pip install -U crawl4ai

# 2. One-time setup: downloads Playwright Chromium and initializes the local cache DB
crawl4ai-setup

# 3. Verify the install end-to-end
crawl4ai-doctor

If crawl4ai-setup fails inside a container or restricted environment, install the browser manually:

python -m playwright install --with-deps chromium

For production workloads, the Crawl4AI Docker image is the cleaner option. It exposes a REST API on port 11235 and a playground UI:

docker pull unclecode/crawl4ai:latest
docker run -d -p 11235:11235 --name crawl4ai \
  --shm-size=1g unclecode/crawl4ai:latest
# Dashboard: http://localhost:11235/dashboard
# Playground: http://localhost:11235/playground

As of v0.9.0 the Docker API server is secure-by-default: authentication is on by default and the server binds to loopback unless you pass a token. If you expose the REST API beyond localhost, set a token / SECRET_KEY and review the v0.9.0 migration guide first. The core pip library is unaffected by this change.

Confirm the version you actually installed before copying code from any tutorial. The API surface changed meaningfully across the 0.5, 0.6, 0.8, and 0.9 lines:

pip show crawl4ai | grep Version

Quickstart: your first Crawl4AI scrape

The minimal working example. It launches headless Chromium, fetches a page, and prints the first 300 characters of clean Markdown.

import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode

async def main():
    # BrowserConfig governs the browser process (headless, viewport, user agent, proxy).
    browser_config = BrowserConfig(headless=True, verbose=True)

    # CrawlerRunConfig governs a single request (cache, timeouts, extraction).
    run_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)

    async with AsyncWebCrawler(config=browser_config) as crawler:
        result = await crawler.arun(
            url="https://example.com",
            config=run_config,
        )
        if result.success:
            # result.markdown is a MarkdownGenerationResult object with multiple variants.
            print(result.markdown.raw_markdown[:300])
        else:
            print("Crawl failed:", result.error_message)

if __name__ == "__main__":
    asyncio.run(main())

BrowserConfig governs the browser instance. CrawlerRunConfig governs each individual request, which is where Crawl4AI recommends putting proxy settings in v0.9.x.

The two-config split is important. BrowserConfig is per-browser; CrawlerRunConfig is per-request. In v0.9.x, the recommended pattern is to set proxies on the run config (CrawlerRunConfig.proxy_config) so each request can carry its own proxy.

What is the output of Crawl4AI?

A successful arun call returns a single CrawlResult object that exposes the page in many formats at once. You pick the field that matches your downstream use case.

FieldWhat it containsWhen to use it
result.markdown.raw_markdownFull Markdown of the rendered DOMMaximum recall, RAG ingestion
result.markdown.fit_markdownMarkdown after boilerplate pruningToken-sensitive LLM prompts
result.markdown.markdown_with_citationsMarkdown with numbered citation footnotesDocument-style outputs
result.cleaned_htmlSanitized HTML, scripts and styles strippedRe-parsing with BeautifulSoup
result.htmlRaw page HTMLArchival, forensic analysis
result.extracted_contentJSON string from your extraction_strategyDirect DB or app ingestion
result.linksDict of internal and external linksLink graphs, deep crawling
result.mediaImage, audio, and video metadataMultimodal pipelines
result.screenshotBase64 PNG of the full pageVisual archive
result.pdfPDF bytes of the rendered pageLong-page archival
result.metadataTitle, description, language, OG tagsCategorization, SEO research

The default Crawl4AI output is Markdown, and that is what makes it different from Scrapy or raw requests. The fit_markdown variant in particular is worth pointing your LLM at first. It strips navigation, footers, cookie banners, and ad slots, which routinely cut token usage by 40% or more on real-world pages.

Why you need a proxy with Crawl4AI

A direct connection from your server's IP is fine for a handful of pages. It stops being fine the moment you scale to a real workload, which is why every serious AI scraping pipeline runs through a residential proxy network.

For a broader comparison of what to look for, the best proxies for web scraping guide covers provider selection criteria across the industry. Three concrete reasons to put a proxy in front of every Crawl4AI run:

  1. Reliable connectivity and stable session quality. Many sites throttle repeated requests from a single IP. Proxies distribute requests across a pool of IPs, which keeps response success rates high across long crawls.
  2. Accurate geo-specific data. Prices, product listings, and search results vary by country and even city. Routing through an IP in the target market is the only way to capture what a user there actually sees. This matters for localization testing, ad verification, and market research.
  3. Compliance with site terms and regional requirements. Public-web data collection done at scale should respect rate limits, regional regulations, and the contractual posture each site sets. A proxy network with country and city targeting lets you stay within those boundaries while still gathering the data you need.

Proxy-Cheap offers four product lines that map cleanly onto Crawl4AI's typical workloads:

Each product line is optimized for a different shape of workload. The links below match the cards above.

  • Static Residential (ISP) proxies for account-bound scraping, persistent sessions, and authenticated dashboards. Same IP held for the full session, ISP-trusted, ideal for LLM agents that maintain state.
  • ISP proxies for long-lived sessions with flexible authentication. Static and rotating ISP options, HTTP/SOCKS5, multiple authentication methods.
  • Datacenter proxies for high-volume crawls of documentation, public catalogs, and static content. High throughput, predictable per-IP pricing.
  • Rotating Residential proxies for geo-sensitive content, dynamic pricing pages, and large-scale rotation. 155M+ IPs across 180+ countries, automatic IP refresh per request or sticky session.

For most AI scraping projects the answer is one of these four, picked by the shape of the target site rather than the size of the budget.

How to set up a proxy with Crawl4AI

A single Crawl4AI call routes through the Proxy-Cheap gateway, fetches the target site through a residential IP, and returns a fully processed CrawlResult.

Crawl4AI gives you two equivalent ways to configure a proxy: a plain dict or the ProxyConfig class. Both attach to CrawlerRunConfig.proxy_config in v0.9.x, which is the method Crawl4AI recommends. Setting a proxy on BrowserConfig still works, but the per-request run-config approach is preferred because it enables rotation and per-site isolation.

Before the code, a quick note on Proxy-Cheap's endpoint format because it differs by product:

  • Rotating residential uses a shared regional gateway. The endpoint is thehub.proxy-cheap.com on port 8080. Each request through the gateway gets a new IP automatically, and you select the country in the dashboard when you generate the credentials. Each request through the gateway gets a new IP automatically.
  • Static residential, ISP, and datacenter assign a unique IP and port to each proxy you order. You copy the exact host and port from the dashboard's proxy list.

The dict format is the simplest entry point. The example below uses the US rotating gateway:

import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig

async def main():
    # The dict format CrawlerRunConfig.proxy_config expects.
    # Keys are exactly: server, username, password.
    proxy_config = {
        "server":   "http://thehub.proxy-cheap.com:8080",   # or proxy-eu...
        "username": "<your-proxycheap-username>",
        "password": "<your-proxycheap-password>",
    }

    run_config = CrawlerRunConfig(proxy_config=proxy_config)

    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        result = await crawler.arun(
            url="https://httpbin.org/ip",   # echoes the outbound IP
            config=run_config,
        )
        print("Outbound HTML:", result.html[:200])

if __name__ == "__main__":
    asyncio.run(main())

The ProxyConfig class adds parsing helpers and is the cleaner option once you graduate beyond a single proxy:

import asyncio
from crawl4ai import (
    AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, ProxyConfig
)

async def main():
    # Equivalent to the dict above, but typed and importable.
    proxy = ProxyConfig(
        server="http://thehub.proxy-cheap.com:8080",
        username="<your-proxycheap-username>",
        password="<your-proxycheap-password>",
    )

    # ProxyConfig.from_string parses several common formats:
    #   "http://user:pass@host:port"
    #   "host:port:user:pass"
    #   "socks5://host:port"
    # proxy = ProxyConfig.from_string("thehub.proxy-cheap.com:8080:user:pass")

    run_config = CrawlerRunConfig(proxy_config=proxy)

    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        result = await crawler.arun(
            url="https://httpbin.org/ip",
            config=run_config,
        )
        print(result.success, result.status_code)

if __name__ == "__main__":
    asyncio.run(main())

Three practical notes from production use:

  • Put credentials in `username` and `password`, not in the `server` URL. Embedding them inline (http://user:pass@host:port on the server field) can trigger ERR_INVALID_AUTH_CREDENTIALS because Playwright handles proxy auth out-of-band.
  • Use `http://` for the `server` scheme. Proxy-Cheap supports HTTP/SOCKS5 across the product line. HTTP is the most compatible choice for Crawl4AI because Playwright handles HTTP proxy authentication natively, while its SOCKS5 username and password support has known limitations.
  • IP whitelist authentication is supported on static residential, ISP, datacenter, static mobile and unlimited-bandwidth. If your crawler runs from a stable IP, whitelist it in the dashboard and omit username and password entirely. Useful for headless servers and CI pipelines.

Proxy-Cheap delivers credentials in standard host:port plus username:password form, generated in the dashboard's "Setup Credentials" panel. That maps one-to-one onto the dict above. For a deeper walkthrough of credential generation, see the step-by-step guide to using residential proxies.

Rotating proxies with Crawl4AI

You get rotation in two different ways depending on the product line:

  • Rotating residential: rotation happens at the Proxy-Cheap gateway, not in Crawl4AI. Each request through the rotating-residential gateway is assigned a new exit IP by Proxy-Cheap, based on the credentials and session type you set in the dashboard (per-request by default, or a sticky session). You only need one entry in your config.
  • Static residential, ISP, or datacenter: each proxy is a fixed IP. To rotate, you maintain a list of proxies and cycle through them in your code.

Crawl4AI ships a built-in rotation strategy for the second case. Pair RoundRobinProxyStrategy with arun_many and the crawler cycles through the proxies in your list in order. Each entry is a distinct Proxy-Cheap IP, so the variety of exit IPs comes from the list you build, not from Crawl4AI.
Important: Crawl4AI does not perform IP rotation on its own. RoundRobinProxyStrategy only cycles through the proxies you provide. The change of exit IP is delivered by Proxy-Cheap, either by the rotating-residential gateway (a new IP per request, or a sticky session you configure) or by the distinct static IPs you list. IP rotation therefore depends on your Proxy-Cheap credentials and session settings, not on Crawl4AI.
The example below uses both rotating residential gateways for geographic diversity, but the same pattern works for any list of static proxies:

import asyncio, os, re
from crawl4ai import (
    AsyncWebCrawler, BrowserConfig, CrawlerRunConfig,
    CacheMode, ProxyConfig,
)
from crawl4ai.proxy_strategy import RoundRobinProxyStrategy

async def main():
    # PROXIES env var format: "host:port:user:pass,host:port:user:pass,..."
    # The example below uses the US and EU rotating gateways for geographic
    # rotation. For static residential / ISP / datacenter, list the unique
    # IPs and ports from your dashboard instead.
    os.environ.setdefault(
        "PROXIES",
        "thehub.proxy-cheap.com:8080:<user>:<session-1-pass>,"
        "thehub.proxy-cheap.com:8080:<user>:<session-2-pass>",
    )

    proxies = ProxyConfig.from_env()   # parses PROXIES into list[ProxyConfig]
    if not proxies:
        raise SystemExit("Set the PROXIES env var first.")

    strategy = RoundRobinProxyStrategy(proxies)

    run_config = CrawlerRunConfig(
        cache_mode=CacheMode.BYPASS,
        proxy_rotation_strategy=strategy,   # rotates per request
    )

    urls = ["https://httpbin.org/ip"] * (len(proxies) * 3)

    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        results = await crawler.arun_many(urls=urls, config=run_config)

        for i, r in enumerate(results):
            if r.success:
                ip = re.search(r"(?:\d{1,3}\.){3}\d{1,3}", r.html)
                print(f"[{i+1}] outbound IP -> {ip.group(0) if ip else '?'}")
            else:
                print(f"[{i+1}] FAIL -> {r.error_message}")

if __name__ == "__main__":
    asyncio.run(main())

The strategy hands out the next proxy from your list each time the crawler requests one. With arun_many, every URL in the batch gets the next entry in sequence.

When to use rotating residential versus static residential depends on the workload:

  • Rotating residential is the default for high-volume AI scraping. The gateway hands you a new IP per request (or a sticky session for around 30 minutes), the pool is large, and response quality stays stable across long jobs. Use this for LLM training data, RAG indexing, market research, and any crawl that touches more than a few thousand pages.
  • Static residential (ISP) is the right choice when you need the same IP to persist for hours or days. Think account-bound scraping, session-aware dashboards, or workflows that depend on a stable IP mid-session.

For a deeper comparison, read static vs rotating proxies: which type works better and how to rotate proxies in Python with Requests and AIOHTTP.

A protocol note: Proxy-Cheap supports HTTP/SOCKS5 across rotating and static products. For Crawl4AI integrations, HTTP with username and password is the most compatible option because Playwright handles HTTP proxy authentication natively. Static residential, ISP, and datacenter products also support IP whitelist authentication if you prefer not to embed credentials.

Advanced: LLM extraction with proxied requests

The real payoff of Crawl4AI is LLMExtractionStrategy. Define a Pydantic schema, point the strategy at any model supported by LiteLLM, and the crawler returns validated structured JSON. Combine it with a proxy and you have a fully production-ready AI scraping pipeline.

import asyncio, os, json
from pydantic import BaseModel, Field
from crawl4ai import (
    AsyncWebCrawler, BrowserConfig, CrawlerRunConfig,
    CacheMode, LLMConfig, ProxyConfig,
)
from crawl4ai.extraction_strategy import LLMExtractionStrategy

class Product(BaseModel):
    title: str = Field(..., description="Product name")
    price: str = Field(..., description="Price including currency symbol")
    description: str = Field("", description="Short product description")

async def main():
    # 1. Schema-driven LLM extraction. Works with OpenAI, Anthropic, Groq, Ollama, etc.
    llm_strategy = LLMExtractionStrategy(
        llm_config=LLMConfig(
            provider="openai/gpt-4o-mini",
            api_token=os.environ["OPENAI_API_KEY"],
        ),
        schema=Product.model_json_schema(),
        extraction_type="schema",
        instruction="Extract every product on the page as a Product object.",
    )

    # 2. Residential proxy for accurate, geo-specific pricing.
    proxy = ProxyConfig(
        server="http://thehub.proxy-cheap.com:8080",
        username=os.environ["PROXY_USER"],
        password=os.environ["PROXY_PASS"],
    )

    run_config = CrawlerRunConfig(
        cache_mode=CacheMode.BYPASS,
        extraction_strategy=llm_strategy,
        proxy_config=proxy,
    )

    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        result = await crawler.arun(
            url="https://web-scraping.dev/products",
            config=run_config,
        )
        if result.success:
            products = json.loads(result.extracted_content)
            print(json.dumps(products[:3], indent=2))
        else:
            print("Extraction failed:", result.error_message)

if __name__ == "__main__":
    asyncio.run(main())

Two things to verify when you put this into production:

  1. LLM cost. Every page sent to a hosted model costs tokens. Use fit_markdown or a CSS pre-filter to trim the input before it hits the LLM. For local-only pipelines, switch the provider to ollama/llama3.3 with api_token=None.
  2. Proxy choice. For pricing-sensitive extractions like e-commerce, travel, or ads, a residential IP in the target country returns the most accurate data because it matches what a real user there sees. For documentation and unprotected public content, datacenter IPs deliver the same content at higher throughput.

Common errors and how to fix them

Most Crawl4AI failures fall into one of seven buckets. Each has a simple fix.

1. `BrowserType.launch: Executable doesn't exist`. Playwright Chromium is missing or installed to a path the runtime cannot see. Run crawl4ai-setup. If that fails, run python -m playwright install --with-deps chromium. In containers, install libnss3 and libatk1.0-0.

2. `Page.wait_for_selector: Timeout 30000ms exceeded`. The default page timeout is too short for JavaScript-heavy pages. Extend it on the run config:

run_config = CrawlerRunConfig(
    page_timeout=120_000,        # 2 minutes
    wait_until="networkidle",    # wait for the network to settle
    wait_for="css:main",         # or a specific selector
)

3. `net::ERR_INVALID_AUTH_CREDENTIALS` with an authenticated proxy. You put credentials inline in the server URL. Move them into explicit username and password fields on the dict or ProxyConfig. Playwright handles proxy auth out-of-band and ignores credentials embedded in the URL on some platforms.

4. Rotating proxy snippet from an old tutorial does not run. Older docs passed BrowserConfig to arun(config=...), which arun rejects. The current pattern is CrawlerRunConfig(proxy_rotation_strategy=RoundRobinProxyStrategy(proxies)), as shown in the rotation section above.

5. `Object of type ProxyConfig is not JSON serializable` against the Docker REST API. When calling the containerized REST endpoint, send proxy_config as a plain dict in the JSON body, not as a ProxyConfig instance. The serialization gap was fixed in v0.7.8 and remains fixed in the current 0.9.x line. Sending proxy_config as a plain dict in the JSON body is still the safe pattern for the Docker REST API.

6. Missing links or empty Markdown on JavaScript-rendered pages. The page is not fully hydrated when extraction runs. Add wait_until="networkidle", delay_before_return_html=3, and flatten_shadow_dom=True for sites that use Web Components.

7. High memory usage during concurrent crawls. Each Chromium instance is heavy. Use arun_many with the built-in MemoryAdaptiveDispatcher rather than wrapping arun in asyncio.gather, share a single AsyncWebCrawler per worker, and pass extra_args=["--no-sandbox", "--disable-dev-shm-usage"] to BrowserConfig.

Is Crawl4AI good? Pros and cons

For AI-focused workloads, Crawl4AI is the strongest open-source choice today. It is free under Apache 2.0, the API is sensibly designed, the Markdown output saves real money on LLM tokens, and the maintainer ships features at a pace most paid scrapers cannot match. A few honest tradeoffs to weigh before you commit.

Pros

  • Apache 2.0, no usage caps, commercial use permitted.
  • Markdown-first output drops token costs in LLM pipelines.
  • Async-first architecture with adaptive dispatch for high concurrency.
  • First-class proxy support with rotation strategies built in.
  • Docker image with REST API for production deployments.
  • Active development; the most-starred open-source AI crawler on GitHub.

Cons

  • You manage the infrastructure: browsers, scaling, proxies, and observability.
  • Playwright's memory footprint adds up quickly at high concurrency.
  • For the most challenging public sites, Crawl4AI on its own may not be enough; pairing it with a strong residential proxy network keeps success rates stable and returns geo-specific data.
  • A few documentation samples lag the current API; always check your installed version.

The middle ground for most teams: run Crawl4AI yourself, pair it with a value-tier proxy provider on pay-as-you-go billing, and you cover 90% of AI scraping use cases without enterprise-tier overhead.

Crawl4AI vs Scrapy. The two tools solve different problems. Scrapy is the right choice for high-volume HTTP-only crawls of static pages where you need maximum throughput and don't need a browser. Crawl4AI is the right choice when the target sites require JavaScript rendering, when you want Markdown output ready for an LLM, or when you want schema-driven structured extraction without writing parsers by hand. Many teams use both: Scrapy for discovery and link harvesting across millions of URLs, then Crawl4AI for AI-ready content extraction on the pages that matter.

자주 묻는 질문

Python 3.10 or newer. After pip install -U crawl4ai, run crawl4ai-setup to download the Playwright Chromium build and initialize the local cache database.

Yes. Crawl4AI is released under the Apache License 2.0, which permits commercial use. Attribution via project badges is recommended but not required, and there is no API key, rate limit, or usage cap built into the library itself.

raw_markdown is the full Markdown conversion of the rendered DOM, used when you want maximum content recall. fit_markdown is the same content with navigation, footers, cookie banners, and ad slots stripped out, which routinely cuts LLM token usage by 40 to 50 percent on real-world pages.

Crawl4AI wraps Playwright to fully render JavaScript before extraction. For heavy SPAs, set page_timeout=120_000, wait_until="networkidle", and wait_for="css:main" on your CrawlerRunConfig so the page is fully hydrated before the Markdown is generated.

Yes. The Proxy-Cheap rotating residential gateway supports both per-request rotation (default) and sticky sessions of around 30 minutes when you need the same exit IP across multiple requests in a single workflow. This behavior is controlled by your Proxy-Cheap credentials and session settings, not by Crawl4AI.

Yes. Configure a single proxy through CrawlerRunConfig.proxy_config as either a dict or a ProxyConfig instance, the method Crawl4AI recommends. For multiple proxies, use RoundRobinProxyStrategy with proxy_rotation_strategy, which cycles through the proxy list you supply. Crawl4AI itself does not create new IPs; the exit IP comes from Proxy-Cheap (the rotating gateway or your list of static IPs).

A dict with three keys: server (for example, http://host:port), username, and password. The ProxyConfig class accepts the same parameters and includes from_string and from_env helpers for parsing list formats.

Yes. Proxy-Cheap supports HTTP/SOCKS5 across rotating and static products. For Crawl4AI specifically, HTTP is the most compatible choice because Playwright handles HTTP proxy authentication natively. Static products also support IP whitelist authentication.

Yes. LLMExtractionStrategy uses LiteLLM under the hood, so any provider works, including Ollama. Set LLMConfig(provider="ollama/llama3.3", api_token=None) to run locally.

Use Scrapy for high-volume HTTP-only crawls of static pages. Use Crawl4AI when you need JavaScript rendering, Markdown output, or LLM-driven extraction. Many teams use Scrapy for discovery and Crawl4AI for AI-ready content extraction.