MCP Hub
Back to servers

@staticn0va/wigolo

Validated

Local-first web search MCP server for AI coding agents

npm568/wk
Stars
1
Tools
8
Updated
Apr 14, 2026
Validated
Apr 26, 2026
Validation Details

Duration: 18.8s

Server: wigolo v0.6.6

Quick Install

npx -y @staticn0va/wigolo

wigolo

Local-first web search MCP server for AI coding agents.

Search, fetch, crawl, cache, and extract — zero API keys, zero cloud, zero cost.

License: BSL 1.1 Node.js TypeScript

Quick Start · Features · Why wigolo? · Roadmap

$ claude mcp add wigolo -- npx @staticn0va/wigolo
Added MCP server wigolo

$ # That's it. Your agent now has web search.

What is this?

wigolo gives AI coding agents (Claude Code, Cursor, Gemini CLI, Codex, Windsurf) web search, page fetching, site crawling, content extraction, and a local knowledge cache. It runs entirely on your machine. No API keys, no cloud, no cost — works out of the box with npx.

Quick Start

Claude Code:

claude mcp add wigolo -- npx @staticn0va/wigolo

Cursor / VS Code / any MCP client:

{
  "mcpServers": {
    "wigolo": {
      "command": "npx",
      "args": ["@staticn0va/wigolo"]
    }
  }
}

Optional warmup (improves quality on first use):

npx @staticn0va/wigolo warmup          # Downloads Playwright + SearXNG
npx @staticn0va/wigolo warmup --all  # + ML reranking + Trafilatura extraction
npx @staticn0va/wigolo warmup --force     # Wipe SearXNG state/install/locks and re-bootstrap

Diagnostics

Run npx @staticn0va/wigolo doctor to see the health of every component (Python, Docker, Playwright, Trafilatura, FlashRank, SearXNG install + process). Exits 0 when healthy, 1 when any required component is degraded. Usable in scripts: npx @staticn0va/wigolo doctor && my-agent.

Daemon Mode

Run wigolo as a persistent HTTP server for lower latency and shared infrastructure:

Start the daemon

npx @staticn0va/wigolo serve
npx @staticn0va/wigolo serve --port 4444 --host 0.0.0.0

The daemon exposes:

  • POST /mcp -- StreamableHTTP MCP transport (preferred)
  • GET /sse -- SSE MCP transport (legacy compatibility)
  • GET /health -- Health check endpoint

Check health

npx @staticn0va/wigolo health
# or
curl http://127.0.0.1:3333/health

Returns:

{
  "status": "healthy",
  "searxng": "active",
  "browsers": "ready",
  "cache": "active",
  "uptime_seconds": 3600
}

Auto-connect

When starting in stdio mode, wigolo checks if a daemon is already running on WIGOLO_DAEMON_PORT. If detected, a notice is printed to stderr. Full stdio-to-daemon proxy is planned for v2.1.

Prerequisites

  • Node.js 20+Download or brew install node (macOS) / winget install OpenJS.NodeJS (Windows) / sudo apt install nodejs (Ubuntu/Debian)
  • Python 3.8+ (recommended)Download or brew install python3 (macOS) / winget install Python.Python.3 (Windows) / sudo apt install python3 (Ubuntu/Debian)
  • Docker (optional) — Alternative to Python for running SearXNG.

Everything else (Playwright, SearXNG) is downloaded automatically on first use or via npx @staticn0va/wigolo warmup.

What works without Python?

Everything except embedded SearXNG. Without Python, search falls back to direct scraping of Bing, DuckDuckGo, and Startpage — functional but less reliable. All other tools (fetch, crawl, cache, extract) work fully with just Node.js.

Features

search

Search the web and get full markdown content in one call — not snippets.

search("React Server Components best practices", { max_results: 5 })
→ titles, URLs, relevance scores, and full extracted markdown per result
  • Domain filtering: include_domains: ["react.dev"], exclude_domains: ["medium.com"]
  • Date filtering: from_date: "2024-01-01", to_date: "2025-01-01"
  • Category search: general, news, code, docs, papers
  • ML reranking with FlashRank when installed
  • Falls back to direct engine scraping when SearXNG is unavailable

fetch

Fetch any URL and get clean markdown. The page-fetching engine behind search.

fetch("https://docs.react.dev/reference/react/useState")
→ clean markdown, links, images, metadata, cached for future use
  • Smart routing: HTTP first, Playwright fallback for JS-rendered pages (auto-detected)
  • Section targeting: section: "Parameters" extracts content under that heading
  • Authenticated browsing: use_auth: true with stored session or Chrome profile
  • PDF support: text extraction via pdf-parse

crawl

Crawl a site from a seed URL — documentation sites, wikis, anything.

crawl("https://docs.example.com", { strategy: "sitemap", max_pages: 50 })
→ array of pages with titles, markdown, depth
  • Strategies: bfs, dfs, sitemap, map (URL discovery only — no content, faster)
  • URL filtering with include/exclude patterns (regex)
  • robots.txt compliance
  • Cross-page content deduplication (strips repeated nav/header/footer)
  • Total character budget to prevent context overflow

cache

Query previously fetched content without hitting the network.

cache({ query: "React hooks", url_pattern: "*react.dev*" })
→ matching cached pages with full markdown
  • SQLite FTS5 full-text search over all cached content
  • Combined filters: text query + URL pattern + date range
  • Cache stats and selective clearing

extract

Structured data extraction from any URL or HTML.

extract("https://example.com/product", { mode: "schema", schema: { price: "string", name: "string" } })
→ { price: "$29.99", name: "Widget Pro" }

Modes:

  • selector — CSS selector → text content
  • tables — HTML tables → structured row objects
  • metadata — title, description, author, date, JSON-LD
  • schema — JSON Schema → heuristic field matching from page content

Why wigolo?

wigoloTavilyFirecrawlExa
CostFree$30–500/mo$16–500/mo$7/1K queries
API key requiredNoneYesYesYes
Authenticated browsingYesNoNoNo
Localhost accessYesNoNoNo
Local cache + FTSYesNoNoNo
Search + extract unifiedYesYesPartialPartial
ML rerankingLocalProprietaryNoNeural index
Rate limitsNoneTieredTieredTiered

Configuration

wigolo works with zero configuration. For advanced use:

# Use an existing SearXNG instance instead of the embedded one
SEARXNG_URL=http://localhost:8888

# Authenticated browsing — export session state via Playwright
WIGOLO_AUTH_STATE_PATH=~/.wigolo/auth.json

# Or use your Chrome profile directly (close Chrome first)
WIGOLO_CHROME_PROFILE_PATH=~/.config/google-chrome/Default

# ML reranking (install with: npx @staticn0va/wigolo warmup --reranker)
WIGOLO_RERANKER=flashrank

# Tune extraction — auto/always/never
WIGOLO_TRAFILATURA=auto

# Logging
LOG_LEVEL=info          # debug, info, warn, error
LOG_FORMAT=json         # json, text

Full list of env vars:

VariableDefaultDescription
SEARXNG_URL(auto)External SearXNG URL
SEARXNG_MODEnativenative or docker
SEARXNG_PORT8888Port for embedded SearXNG
WIGOLO_DATA_DIR~/.wigoloData + cache directory
WIGOLO_AUTH_STATE_PATHPlaywright storage state JSON
WIGOLO_CHROME_PROFILE_PATHChrome user data directory
WIGOLO_RERANKERnoneflashrank or none
WIGOLO_TRAFILATURAautoauto, always, or never
MAX_BROWSERS3Concurrent Playwright contexts
FETCH_TIMEOUT_MS10000HTTP fetch timeout
CRAWL_CONCURRENCY2Concurrent crawl requests
RESPECT_ROBOTS_TXTtrueHonor robots.txt
WIGOLO_BOOTSTRAP_MAX_ATTEMPTS3Cap on SearXNG bootstrap auto-retries
WIGOLO_BOOTSTRAP_BACKOFF_SECONDS30,3600,86400Backoff seconds for retry attempts 1, 2, 3
WIGOLO_HEALTH_PROBE_INTERVAL_MS30000Interval between SearXNG /healthz probes
WIGOLO_DAEMON_PORT3333HTTP server port for daemon mode
WIGOLO_DAEMON_HOST127.0.0.1HTTP server bind address for daemon mode

How it works

search query
    → SearXNG (70+ engines) or direct scraping (Bing/DDG/Startpage)
    → deduplicate by URL
    → domain/date/category filters
    → ML reranking (FlashRank, optional)
    → link validation
    → fetch + extract top N results in parallel
    → return markdown

Each step degrades gracefully:
  SearXNG down?        → direct scraping fallback
  Page needs JS?       → auto-detected, Playwright used transparently
  Extractor fails?     → ensemble: site-specific → Defuddle → Trafilatura → Readability → Turndown
  Already fetched?     → served from SQLite cache with FTS5

SearXNG bootstrap failures are self-healing: wigolo retries after 30 seconds, 1 hour, and 24 hours on successive server restarts. Once attempts are exhausted, direct-scraping stays permanent until the user runs warmup --force. Tool responses include a one-time fallback warning so agents can surface the recovery command. See doctor for the full state.

Extraction ensemble — every page runs through multiple extractors in order, falling back if content is below threshold:

  1. Site-specific extractors (GitHub, Stack Overflow, MDN, docs frameworks)
  2. Defuddle — markdown-aware, site-adaptive
  3. Trafilatura — high-precision article extraction (Python, optional)
  4. Readability.js — battle-tested Mozilla algorithm
  5. Raw Turndown — last resort HTML-to-markdown

Roadmap

v2.1 — Next

  • Daemon mode — persistent HTTP server, zero startup latency
  • Browser interaction — click, type, scroll before extraction
  • Content change detection — diff monitoring for cached pages
  • CDP session discovery — attach to running Chrome for seamless auth
  • Plugin system — community extractors and search engines

v2.2

  • Multi-browser pool — Chromium + Firefox for fingerprint diversity
  • Interactive REPL (wigolo shell)
  • Agent skill distribution — MCP registry listings, SKILL.md

v3 — The Knowledge Engine

  • Answer synthesis — search + LLM = direct answers with citations (bring your own key)
  • Semantic search — local vector embeddings over cached content (findSimilar)
  • Agent endpoint — describe what you need, no URLs required
  • Streaming answers — real-time generation as results come in
  • Knowledge graph — entity and relationship extraction from crawled content
  • Auto re-crawl scheduler — keep documentation fresh automatically
  • Lightpanda browser — optional ultra-lightweight headless browser (11x less RAM than Chrome)
  • Cloud sync — share cache across machines via rclone (S3, Drive, Dropbox)
  • Team knowledge base — shared indexed content across team members

Discovery

wigolo is listed on MCP server registries for agent discovery:

  • SKILL.md -- machine-readable tool description at repo root
  • npm -- npm info @staticn0va/wigolo or search for mcp-server keyword

To add wigolo to your agent's toolset:

claude mcp add wigolo -- npx @staticn0va/wigolo

See SKILL.md for the full tool schema in agent-discovery format.

Troubleshooting

SearXNG won't start Make sure python3 is on your PATH and version 3.8+. Check with python3 --version. Alternatively, set SEARXNG_MODE=docker if Docker is available.

Playwright browser not found Run npx @staticn0va/wigolo warmup to download Chromium. This is done automatically on first use but can fail behind corporate proxies.

Search returns no results If SearXNG and all fallback engines fail, check your network connection. Behind a proxy? Set PROXY_URL=http://your-proxy:port.

Permission errors on ~/.wigolo/ wigolo stores its cache and SearXNG installation in ~/.wigolo/. Ensure your user has write access. Override with WIGOLO_DATA_DIR=/your/path.

Contributing

PRs welcome. Open an issue first to discuss what you'd like to change.

git clone https://github.com/KnockOutEZ/wigolo
cd wigolo
npm install
npm test

Releasing

Releases are triggered by pushing a version tag. CI handles the rest.

# on main, all changes committed and pushed
make release-patch   # or: release-minor / release-major

Run make help for all targets, or make release-dry-run to preview the npm tarball.

The release workflow will:

  1. Build a clean dist/
  2. Verify the tag matches package.json version
  3. Publish to npm with provenance
  4. Create a GitHub Release with auto-generated notes

Requires the NPM_TOKEN repository secret (npm automation token with publish scope).

License

BSL 1.1 — free for individuals, small teams (under $1M revenue), education, and open source. Converts to MIT on 2029-04-12.

Reviews

No reviews yet

Sign in to write a review