Sriram-PR/doc-scraper
by Various
Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
MCP
Sriram-PR/doc-scraper
Added 8 Sept 2026
Overview
Go web crawler that scrapes documentation sites and converts the content into clean Markdown for LLM ingestion, such as RAG pipelines or training data. It targets documentation pages specifically and outputs structured Markdown files.
Best for
Best for
Developers who need a quick, no-frills way to turn documentation sites into Markdown for LLM-based applications.
Use cases
- Scrape documentation sites into Markdown for RAG knowledge bases
- Convert multiple doc pages into a consistent format for LLM fine-tuning
- Build local corpora from public docs for offline retrieval
How to use
Tools exposed
default_user_agentdefault_delay_per_hostnum_workersnum_image_workersmax_requestsmax_requests_per_hostoutput_base_dirstate_dirmax_retriesinitial_retry_delaymax_retry_delayglobal_crawl_timeoutper_page_timeoutskip_imagesmax_image_size_bytesmax_page_size_bytesenable_jsonl_outputjsonl_output_filenameenable_incrementalcrawl_history_retention
Tested with
Claude Desktop, Claude Code, Cursor, Continue
Notes
Go web crawler that scrapes documentation sites and converts the content into clean Markdown for LLM ingestion, such as RAG pipelines or training data. It targets documentation pages specifically and outputs structured Markdown files.
98 stars on GitHub. Last updated 2026-09-06. Licensed Apache-2.0.
Use cases
- Scrape documentation sites into Markdown for RAG knowledge bases
- Convert multiple doc pages into a consistent format for LLM fine-tuning
- Build local corpora from public docs for offline retrieval
Pros
- Simple, focused tool written in Go with fast execution
- Produces clean Markdown suitable for LLM context windows
- Lightweight and easy to integrate into existing pipelines
Cons
- Limited to documentation sites; not a general-purpose scraper
- Small community and minimal documentation beyond the repo
- No built-in scheduling or incremental crawling features
Indexed from awesome-mcp-servers-punkpeye and enriched against its public facts.
Pros
- Simple, focused tool written in Go with fast execution
- Produces clean Markdown suitable for LLM context windows
- Lightweight and easy to integrate into existing pipelines
Cons
- Limited to documentation sites; not a general-purpose scraper
- Small community and minimal documentation beyond the repo
- No built-in scheduling or incremental crawling features
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
