数据集浏览器
启用与 Hugging Face 数据集查看器 API 的交互,允许用户浏览、搜索、过滤和分析托管在 Hugging Face Hub 上的数据集。
MCP 服务配置
复制以下 JSON 到 OPClaw 或其他 MCP 客户端的配置文件中即可使用
{
"mcpServers": {
"dataset-viewer": {
"args": [
"run",
"dataset-viewer"
],
"command": "uv"
}
}
}
该服务需要配置环境变量:HUGGINGFACE_TOKEN
可用工具 (8 个)
该服务在 MCP 协议中暴露的工具,AI 可按需调用
get_info 2 个参数 需填 1 项
Get detailed information about a Hugging Face dataset including description, features, splits, and statistics. Run validate first to check if the dataset exists and is accessible.
必填参数:dataset
get_rows 5 个参数 需填 3 项
Get paginated rows from a Hugging Face dataset
必填参数:dataset、config、split
get_first_rows 4 个参数 需填 3 项
Get first rows from a Hugging Face dataset split
必填参数:dataset、config、split
search_dataset 5 个参数 需填 4 项
Search for text within a Hugging Face dataset
必填参数:dataset、config、split、query
filter 7 个参数 需填 4 项
Filter rows in a Hugging Face dataset using SQL-like conditions
必填参数:dataset、config、split、where
get_statistics 4 个参数 需填 3 项
Get statistics about a Hugging Face dataset
必填参数:dataset、config、split
get_parquet 2 个参数 需填 1 项
Export Hugging Face dataset split as Parquet file
必填参数:dataset
validate 2 个参数 需填 1 项
Check if a Hugging Face dataset exists and is accessible
必填参数:dataset
服务介绍
Dataset Viewer MCP Server
一个用于与 Hugging Face Dataset Viewer API 交互的MCP服务器,提供了浏览和分析托管在Hugging Face Hub上的数据集的功能。
功能
资源
- 使用
dataset://URI 方案访问 Hugging Face 数据集 - 支持数据集配置和分割
- 提供分页访问数据集内容
- 处理私有数据集的身份验证
- 支持搜索和过滤数据集内容
- 提供数据集统计和分析
工具
服务器提供以下工具:
-
validate
- 检查数据集是否存在且可访问
- 参数:
dataset: 数据集标识符(例如 'stanfordnlp/imdb')auth_token(可选): 用于私有数据集
-
get_info
- 获取关于数据集的详细信息
- 参数:
dataset: 数据集标识符auth_token(可选): 用于私有数据集
-
get_rows
- 获取数据集的分页内容
- 参数:
dataset: 数据集标识符config: 配置名称split: 分割名称page(可选): 页码(从0开始)auth_token(可选): 用于私有数据集
-
get_first_rows
- 从数据集分割中获取前几行
- 参数:
dataset: 数据集标识符config: 配置名称split: 分割名称auth_token(可选): 用于私有数据集
-
get_statistics
- 获取关于数据集分割的统计信息
- 参数:
dataset: 数据集标识符config: 配置名称split: 分割名称auth_token(可选): 用于私有数据集
-
search_dataset
- 在数据集中搜索文本
- 参数:
dataset: 数据集标识符config: 配置名称split: 分割名称query: 要搜索的文本auth_token(可选): 用于私有数据集
-
filter
- 使用类似SQL的条件过滤行
- 参数:
dataset: 数据集标识符config: 配置名称split: 分割名称where: SQL WHERE 子句(例如 "score > 0.5")orderby(可选): SQL ORDER BY 子句page(可选): 页码(从0开始)auth_token(可选): 用于私有数据集
-
get_parquet
- 以Parquet格式下载整个数据集
- 参数:
dataset: 数据集标识符auth_token(可选): 用于私有数据集
安装
前提条件
- Python 3.12 或更高版本
- uv - 快速的Python包安装器和解析器
设置
- 克隆仓库:
git clone https://github.com/privetin/dataset-viewer.git
cd dataset-viewer
- 创建虚拟环境并安装:
# Create virtual environment
uv venv
# Activate virtual environment
# On Unix:
source .venv/bin/activate
# On Windows:
.venv\Scripts\activate
# Install in development mode
uv add -e .
配置
环境变量
HUGGINGFACE_TOKEN: 用于访问私有数据集的Hugging Face API令牌
Claude Desktop集成
将以下内容添加到您的Claude Desktop配置文件中:
在 Windows 上: %APPDATA%\Claude\claude_desktop_config.json
在 MacOS 上: ~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"dataset-viewer": {
"command": "uv",
"args": [
"run",
"dataset-viewer"
]
}
}
}
使用示例
- 验证数据集:
{
"dataset": "stanfordnlp/imdb"
}
- 获取数据集信息:
{
"dataset": "stanfordnlp/imdb"
}
- 搜索数据集内容:
{
"dataset": "stanfordnlp/imdb",
"config": "plain_text",
"split": "train",
"query": "great movie"
}
- 筛选并排序行:
{
"dataset": "stanfordnlp/imdb",
"config": "plain_text",
"split": "train",
"where": "label = 'positive'",
"orderby": "text DESC",
"page": 0
}
- 获取数据集统计信息:
{
"dataset": "stanfordnlp/imdb",
"config": "plain_text",
"split": "train"
}
许可证
MIT 许可证 - 详情请参阅 LICENSE