Introduction
Open-source web scraping engine for LLMs - Node library, CLI, and deployment scripts.
Reader is an open-source web scraping engine built for LLMs, distributed under Apache 2.0. Two primitives, clean markdown, ready for your agents.
reader on GitHub
Full source, issues, Dockerfile, and examples.
What Reader gives you
ReaderClient- high-level API with lazy initialization, browser pool management, and proxy rotationscrape()andcrawl()- the two core primitives for turning URLs into clean content- Playwright browser engine - full headless Chrome with JavaScript execution and anti-bot bypass via stealth plugin
- Proxy tiers - standard (datacenter) and premium (residential) proxy pools, selectable per request
- Browser pool - recycled Chrome instances with health checks and graceful retirement
- CLI - one-off scrapes, crawls, and a daemon mode with shared pool
- Pluggable config - domain profiles, block detection, and URL rewriters are all caller-provided
- Deployment scripts - production-ready Dockerfile and Docker Compose setup