h

huoshui-fetch

@huoshuiai42/huoshui-fetch
Hosted
0 Stars 6 次浏览 huoshuiai42 更新于 2026-08-23

An MCP server that provides tools for fetching, converting, and extracting data from web pages.

MCP 服务配置

复制以下 JSON 到 OPClaw 或其他 MCP 客户端的配置文件中即可使用

{
  "mcpServers": {
    "huoshui-fetch": {
      "args": [
        "huoshui-fetch@1.0.0"
      ],
      "command": "uvx"
    }
  }
}

可用工具 (5 个)

该服务在 MCP 协议中暴露的工具,AI 可按需调用

tavily_search 14 个参数 需填 1 项

Search the web for current information on any topic. Use for news, facts, or data beyond your knowledge cutoff. Returns snippets and source URLs.

必填参数:query

tavily_extract 6 个参数 需填 1 项

Extract content from URLs. Returns raw page content in markdown or text format.

必填参数:urls

tavily_crawl 11 个参数 需填 1 项

Crawl a website starting from a URL. Extracts content from pages with configurable depth and breadth.

必填参数:url

tavily_map 8 个参数 需填 1 项

Map a website's structure. Returns a list of URLs found starting from the base URL.

必填参数:url

tavily_research 2 个参数 需填 1 项

Perform comprehensive research on a given topic or question. Use this tool when you need to gather information from multiple sources to answer a question or complete a task. Returns a detailed response based on the research findings.

必填参数:input

服务介绍

huoshui-fetch

A dedicated web content fetching and conversion MCP (Model Context Protocol) server that provides tools for fetching, converting, and extracting data from web pages.

# Features

# # Fetching Tools

  • fetch_url: Fetch content from URLs with customizable timeout, redirect handling, and user-agent
  • fetch_with_headers: Fetch URLs with custom headers for authenticated requests

# # Conversion Tools

  • html_to_markdown_tool: Convert HTML to clean Markdown format
  • html_to_text_tool: Extract plain text from HTML
  • clean_html_tool: Remove scripts/styles and sanitize HTML
  • json_to_markdown_tool: Convert JSON data to readable Markdown

# # Extraction Tools

  • extract_article_tool: Extract main article content using readability
  • extract_links_tool: Extract all links with filtering options
  • extract_metadata_tool: Extract page metadata (title, description, OG tags)
  • extract_images_tool: Extract images with size filtering
  • extract_structured_data_tool: Extract JSON-LD and microdata

# Installation

From MCP Registry (Recommended)

This server is available in the Model Context Protocol Registry. Install it using your MCP client.

mcp-name: io.github.huoshuiai42/huoshui-fetch

#  Using uv (recommended)
uv sync

#  Or install from GitHub
pip install git+https://github.com/yourusername/huoshui-fetch.git

# Usage

# # Run with uvx (recommended for one-time use)

#  From the repository
uvx - -from . huoshui-fetch

#  From GitHub (once published)
uvx - -from git+https://github.com/yourusername/huoshui-fetch.git huoshui-fetch

# # Run directly

#  Using uv
uv run python -m huoshui_fetch

#  Or if installed
python -m huoshui_fetch

The server communicates via standard input/output, making it perfect for integration with Claude Desktop and other MCP-compatible clients.

# Configuration for Claude Desktop

Add to your Claude Desktop configuration:

{
  "mcpServers": {
    "huoshui-fetch": {
      "command": "uvx",
      "args": ["- -no-cache", "- -from", ".", "huoshui-fetch"],
      "cwd": "/path/to/huoshui-fetch"
    }
  }
}

Or if installed from GitHub:

{
  "mcpServers": {
    "huoshui-fetch": {
      "command": "uvx",
      "args": [
        "- -from",
        "git+https://github.com/yourusername/huoshui-fetch.git",
        "huoshui-fetch"
      ]
    }
  }
}

# Example Usage

Once configured, you can use the tools in Claude Desktop:

// Fetch a webpage
fetch_url("https://example.com")

// Convert HTML to Markdown
html_to_markdown_tool("<h1>Hello</h1><p>World</p>")

// Extract article content
extract_article_tool(html_content, "https://example.com/article")

# Requirements

  • Python 3.11+
  • Dependencies listed in pyproject.toml

# Development & Publishing

This project includes comprehensive automation for building and publishing to PyPI.

# # Automated Publishing Workflow

#  Complete automated workflow (TestPyPI + PyPI)
uv run python scripts/publish.py - -include-pypi

#  TestPyPI only (recommended for testing)
uv run python scripts/publish.py

#  Bump version and publish
uv run python scripts/publish.py - -version-bump patch - -include-pypi

# # Individual Commands

#  Version management
uv run python scripts/version_manager.py - -check
uv run python scripts/version_manager.py - -bump patch

#  Setup PyPI credentials (first time)
uv run python scripts/credentials_setup.py

#  Build package
uv run python scripts/build.py

#  Run comprehensive tests
uv run python scripts/test.py

#  Upload to PyPI
uv run python scripts/upload.py

# # Features

  • Version Management: Automatic synchronization across all files
  • Quality Checks: Ruff linting and MyPy type checking
  • Build Automation: Clean builds with validation
  • Testing Suite: Comprehensive package and functionality tests
  • Publishing Workflow: TestPyPI → PyPI using uv publish (supports .pypirc files)
  • Error Recovery: Built-in error handling and recovery options

See PUBLISHING.md for detailed documentation.

# DXT Extension

This project supports DXT (Desktop Extensions) format for easy distribution and installation.

To build the DXT extension:

python build_dxt.py

This will create a huoshui-fetch-{version}.dxt file that can be installed in compatible AI desktop applications.

# License

MIT

相关 MCP 服务