haiku-rag
Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling
MCP 服务配置
复制以下 JSON 到 OPClaw 或其他 MCP 客户端的配置文件中即可使用
{
"mcpServers": {
"haiku-rag": {
"args": [
"haiku-rag@0.28.0"
],
"command": "uvx"
}
}
}
可用工具 (5 个)
该服务在 MCP 协议中暴露的工具,AI 可按需调用
tavily_search 14 个参数 需填 1 项
Search the web for current information on any topic. Use for news, facts, or data beyond your knowledge cutoff. Returns snippets and source URLs.
必填参数:query
tavily_extract 6 个参数 需填 1 项
Extract content from URLs. Returns raw page content in markdown or text format.
必填参数:urls
tavily_crawl 11 个参数 需填 1 项
Crawl a website starting from a URL. Extracts content from pages with configurable depth and breadth.
必填参数:url
tavily_map 8 个参数 需填 1 项
Map a website's structure. Returns a list of URLs found starting from the base URL.
必填参数:url
tavily_research 2 个参数 需填 1 项
Perform comprehensive research on a given topic or question. Use this tool when you need to gather information from multiple sources to answer a question or complete a task. Returns a detailed response based on the research findings.
必填参数:input
服务介绍
Haiku RAG
Agentic RAG built on LanceDB, Pydantic AI, and Docling.
Features
- Hybrid search 鈥� Vector + full-text with Reciprocal Rank Fusion
- Question answering 鈥� QA agents with citations (page numbers, section headings)
- Reranking 鈥� MxBAI, Cohere, Zero Entropy, or vLLM
- Research agents 鈥� Multi-agent workflows via pydantic-graph: plan, search, evaluate, synthesize
- Conversational RAG 鈥� Chat TUI and web application for multi-turn conversations with session memory
- Document structure 鈥� Stores full DoclingDocument, enabling structure-aware context expansion
- Multiple providers 鈥� Embeddings: Ollama, OpenAI, VoyageAI, LM Studio, vLLM. QA/Research: any model supported by Pydantic AI
- Local-first 鈥� Embedded LanceDB, no servers required. Also supports S3, GCS, Azure, and LanceDB Cloud
- CLI & Python API 鈥� Full functionality from command line or code
- MCP server 鈥� Expose as tools for AI assistants (Claude Desktop, etc.)
- Visual grounding 鈥� View chunks highlighted on original page images
- File monitoring 鈥� Watch directories and auto-index on changes
- Time travel 鈥� Query the database at any historical point with
--before - Inspector 鈥� TUI for browsing documents, chunks, and search results
Installation
Python 3.12 or newer required
Full Package (Recommended)
pip install haiku.rag
Includes all features: document processing, all embedding providers, and rerankers.
Using uv? uv pip install haiku.rag
Slim Package (Minimal Dependencies)
pip install haiku.rag-slim
Install only the extras you need. See the Installation documentation for available options.
Quick Start
Note: Requires an embedding provider (Ollama, OpenAI, etc.). See the Tutorial for setup instructions.
# Index a PDF
haiku-rag add-src paper.pdf
# Search
haiku-rag search "attention mechanism"
# Ask questions with citations
haiku-rag ask "What datasets were used for evaluation?" --cite
# Deep QA 鈥� decomposes complex questions into sub-queries
haiku-rag ask "How does the proposed method compare to the baseline on MMLU?" --deep
# Research mode 鈥� iterative planning and search
haiku-rag research "What are the limitations of the approach?"
# Interactive chat 鈥� multi-turn conversations with memory
haiku-rag chat
# Watch a directory for changes
haiku-rag serve --monitor
See Configuration for customization options.
Python API
from haiku.rag.client import HaikuRAG
async with HaikuRAG("research.lancedb", create=True) as rag:
# Index documents
await rag.create_document_from_source("paper.pdf")
await rag.create_document_from_source("https://arxiv.org/pdf/1706.03762")
# Search 鈥� returns chunks with provenance
results = await rag.search("self-attention")
for result in results:
print(f"{result.score:.2f} | p.{result.page_numbers} | {result.content[:100]}")
# QA with citations
answer, citations = await rag.ask("What is the complexity of self-attention?")
print(answer)
for cite in citations:
print(f" [{cite.chunk_id}] p.{cite.page_numbers}: {cite.content[:80]}")
For research agents and chat, see the Agents docs.
MCP Server
Use with AI assistants like Claude Desktop:
haiku-rag serve --mcp --stdio
Add to your Claude Desktop configuration:
{
"mcpServers": {
"haiku-rag": {
"command": "haiku-rag",
"args": ["serve", "--mcp", "--stdio"]
}
}
}
Provides tools for document management, search, QA, and research directly in your AI assistant.
Examples
See the examples directory for working examples:
- Docker Setup - Complete Docker deployment with file monitoring and MCP server
- Web Application - Full-stack conversational RAG with CopilotKit frontend
Documentation
Full documentation at: https://ggozad.github.io/haiku.rag/
- Installation - Provider setup
- Architecture - System overview
- Configuration - YAML configuration
- CLI - Command reference
- Python API - Complete API docs
- Agents - QA, chat, and research agents
- Applications - Chat TUI, web app, and inspector
- Server - File monitoring and MCP
- MCP - Model Context Protocol integration
- Benchmarks - Performance benchmarks
- Changelog - Version history
License
This project is licensed under the MIT License.
mcp-name: io.github.ggozad/haiku-rag