🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
-
Updated
Mar 23, 2026 - JavaScript
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
Open-source, production-grade web scraping engine built for LLMs. Scrape and crawl the entire web, clean markdown, ready for your agents.
Model Context Protocol (MCP) Server for Graphlit Platform
A fork of Dragnet that also extract author, headline, date, keywords from context, as well as built in metadata extraction all in one package
Full-content web fetcher for AI agents — Chrome TLS fingerprinting, browser impersonation, and multi-strategy article extraction
A powerful MCP server extension providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content extraction into your MCP environment, enabling AI assistants to search the web and extract webpage content programmatically.
Readability2 converts HTML to plain text.
Next.js template for seamless PDF parsing using pdf2json and FilePond. Ideal for developers seeking a ready-to-use solution for PDF content extraction in Next.js projects.
A collection of OpenClaw Agent Skills — search, analysis, content extraction, and more.
Pure ruby implementation of the Boilerpipe content extraction algorithm tuned for online articles
DOM Based Content Extraction via Text Density
Web content extraction using machine learning
🔍 Model Context Protocol (MCP) tool for parsing websites using the Jina.ai Reader
Local browser toolkit for AI agents: deep research and browser use automation with local Chrome (CDP) + Playwright. Flexible, extensible scripts for web navigation, extraction and workflow automatization - built for reproducible research and agent-driven browsing.
Configurable web access extension for pi that routes search, contents, answers, and research across Claude, Codex, Exa, Gemini, Parallel, and Valyu providers.
Tool to extracts the text from a web article urls and get frequency words, entities recognition, automatic summary and more
Pure Rust document-to-Markdown converter for LLM workflows (DOCX, PPTX, XLSX, HTML, CSV, JSON, XML, images).
Make PDF Files Accessible, Extract Data from PDF, Convert PDF to HTML, Fill-in PDF Form, Stamp PDF and more...
Benson turns a list of URLs into mp3s of the contents of each web page - take control over your reading backlog!
A userscript that adds a button to YouTube video pages for copying the transcript with or without timestamps.
Add a description, image, and links to the content-extraction topic page so that developers can more easily learn about it.
To associate your repository with the content-extraction topic, visit your repo's landing page and select "manage topics."