CrawleeβA web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
This repository does not currently contain dedicated AI agent rule files (.cursorrules, CLAUDE.md, AGENTS.md).
Read Full Technical Documentation