Sriram-PR/doc-scraper
by Various
Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
MCP
Sriram-PR/doc-scraper
Added 8 Sept 2026
Overview
Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
How to use
Tools exposed
default_user_agentdefault_delay_per_hostnum_workersnum_image_workersmax_requestsmax_requests_per_hostoutput_base_dirstate_dirmax_retriesinitial_retry_delaymax_retry_delayglobal_crawl_timeoutper_page_timeoutskip_imagesmax_image_size_bytesmax_page_size_bytesenable_jsonl_outputjsonl_output_filenameenable_incrementalcrawl_history_retention
Tested with
Claude Desktop, Claude Code, Cursor, Continue
Notes
Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
98 stars on GitHub. Last updated 2026-09-06. Licensed Apache-2.0.
Indexed from awesome-mcp-servers-punkpeye. Verified against the live GitHub repo.
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
Anything LLM
Community
The all-in-one AI productivity accelerator. On device and privacy first with no annoying setup or configuration.
Private GPT
Community
Interact with your documents using the power of GPT, 100% privately, no data leaks
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
