data-ai MCP Server
Give your AI assistant real web search, full-page reading, and multi-source research with citations that are never fabricated.
Discovered via github-topic:model-context-protocol and last synced 1d ago.
1. Install the package
uvx web-researcher-mcp
2. Add to claude_desktop_config.json
{
"mcpServers": {
"web-researcher-mcp": {
"command": "uvx",
"args": [
"web-researcher-mcp"
]
}
}
}Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)
pages, PDFs, DOCX, PPTX, YouTube transcripts |
Perplexity
No
No (separate app)
**Free forever** (open source)
What it does
Search the web — optionally restricted to only the sources you trust via lenses
Search and then read the best results — with quality scoring to surface the most reliable sources
Find images by size, type, color, or format
Search recent news with date controls and source filtering
Find real papers with real DOIs — authors, citation counts, open-access links
Fetch a paper's full text in one call from its DOI, Semantic Scholar ID, or URL — no need to chain `academic_search` then `scrape_page`
Walk a paper's citation neighborhood — works it cites and works that cite it, with intent/influence signals
Search patent offices (US, Europe, international) with classification codes
Search SEC EDGAR for US public-company filings (10-K, 10-Q, 8-K, …) — or pull structured XBRL company facts
Search US court opinions and dockets via CourtListener — real cases with real citations
OSINT company reconnaissance — Certificate Transparency log SANs, Wayback Machine historical URL inventory, derived subdomains, and a web-search company summary. Each phase fails soft and is independently selectable
Export a research session as a shareable report (markdown or JSON), with full per-step provenance
Research a brand and produce use-case-specific creative direction (landing page, email, video brief) — calls `brand_research` and interprets the structured JSON for you
`SEARCHAPI_API_KEY`
Rare-disease and biomedical knowledge-graph sources — ontology portals, gene-disease databases, curated rare-disease registries
Technology industry
Look up economic data — World Bank global development indicators, OECD economic indicators, Eurostat European statistics (all keyless), and FRED US macro series (GDP, CPI, unemployment, rates; requires FRED_API_KEY)
Check a citation before you rely on it — does it exist, match a real record, and is it retracted or a dead link? Evidence, not a verdict
Turn collected sources into a formatted bibliography — APA, MLA, BibTeX, RIS, or CSL-JSON (Zotero/EndNote/Mendeley-ready)
Deep OSINT reconnaissance on a company — maps infrastructure, filings, personnel, and public footprint
`searxng`
Notes
Clinical trials, drug safety, evidence-based medicine
Law, cases, statutes
Search ClinicalTrials.gov — clinical-trial registrations with status, phase, sponsor, and whether results are posted (discovery, not medical advice)
Audit a whole reference list in one pass — paste a CSL-JSON/RIS/BibTeX file (or a session) and get per-entry + corpus-level flags for retracted, dead-link, and unverifiable citations
Ask the same question to a panel of independently configured LLMs and compare answers — consensus, contradictions, and model-unique points, computed deterministically, never smoothed over by an arbiter model
Research a subject's syllabus coverage, institutional climate, and academic-freedom context — calls `web_search` with the `curriculum` lens
`tavily`
Zero-config (public REST Search API); searches issues/PRs, not the full web
Academic curriculum data, institutional free speech climate, and global education statistics
Health, medicine
Query the Monarch Initiative biomedical knowledge graph — rank diseases and genes by phenotype similarity, look up disease/gene/phenotype entities, traverse gene-disease-phenotype associations
Audit an AI-generated recommendation list (listicle, product ranking) for self-promotion, author conflicts of interest, domain reputation, and dead links — catches GEO-gamed picks. Evidence, not a verdict
What it guides your AI to do
Whole-Web
`exa`
Public records, corporate filings, FOIA
Research, papers
Search the ecosyste.ms Awesome API for community-curated "awesome-*" lists on a GitHub topic — structured, filterable coverage (stars, curated-entry count, topics) beyond free-text search
Capture a fresh Internet Archive (Wayback Machine) snapshot of a URL via Save Page Now so a cited source stays verifiable if the page later changes or disappears — returns snapshot URL + timestamp (write tool)
Run a structured, multi-step deep dive on a topic
`duckduckgo`
none
Focus
Code docs, tutorials, Q&A
Policy, regulations
Search for physical places (restaurants, shops, services, points of interest) by local intent query — structured POI details and descriptions. Requires `BRAVE_API_KEY`
Multi-step deep research — your AI remembers what it already found and builds on it
Verify a claim against multiple independent sources
`GOOGLE_CUSTOM_SEARCH_API_KEY` + `GOOGLE_CUSTOM_SEARCH_ID`
`reddit`
Official documentation and API references only
Developer-first results re-ranked by Brave's Programming Goggle — surfaces docs, repos, and authoritative technical content (requires Brave)
Open-source intelligence — public records, corporate registries, social footprint, infrastructure
Research a company's complete brand identity — colors (hex), logos, typography, tone of voice, and social handles — from any domain or company name. Returns structured JSON for AI content generation. No API key required; BrandFetch key optional for richer data
Recover a research session after context loss — picks up right where you left off
Systematically review academic literature on a topic
`serper`
`github`
Preprint servers, OA aggregators, and repositories beyond core journal indexes
Current events, journalism
Description
Size up a company and its market (news, patents, web)
`brave`
`bluesky`
Preprint servers, repositories, open-access journals
Infrastructure and operations — Kubernetes, Docker, Terraform, cloud, CI/CD
Community-curated "awesome-*" lists on GitHub — PR-reviewed tool and resource collections across every domain
CVEs, advisories, vulnerability research
Markets, filings
Read any URL in full — web pages, PDFs, Word docs, slideshows, YouTube transcripts, Hacker News threads (read natively via the HN API); supports `mode: raw` for verbatim, unsanitized source (e.g. inspecting JSON or HTML)
`xquik`
`YOUDOTCOM_API_KEY`