NetLens

NetLens

An MCP server that fetches any URL and returns clean Markdown, plus web search with real links.

Category
访问服务器

README

NetLens

npm PyPI CI License: Apache 2.0 Python 3.10+

An MCP server for unobstructed web reading. It fetches any URL directly with browser-like headers — past robots.txt and naive bot blocks — and returns the full page as clean, ad-stripped Markdown, not a summary. Plus web search that returns real links. Zero dependencies: pure Python standard library.

Built for AI Agents

AI agents constantly hit pages their built-in tools can't read. NetLens fixes the three usual reasons a fetch comes back empty or useless:

Native web tools NetLens
Honor robots.txt, so crawler-disallowed pages return nothing Reads like the browser you'd open yourself — doesn't consult robots.txt
Blocked by header/User-Agent bot filters (403/202 to non-browser clients) Sends real browser headers via the system curl; commonly turns 403 → 200
Return a summary of the page Returns the full page content as Markdown
Leave ads, cookie banners, nav, and related-links chrome in the output Strips boilerplate locally so only the content reaches your context

It does not try to defeat JavaScript/Cloudflare challenge pages or CAPTCHAs — that's out of scope by design. When a page is a hard block, the HTTP status is surfaced honestly rather than faked.

Installation

npm (via npx):

{
  "mcpServers": {
    "netlens": {
      "command": "npx",
      "args": ["-y", "netlens-mcp"]
    }
  }
}

PyPI (via uvx):

{
  "mcpServers": {
    "netlens": {
      "command": "uvx",
      "args": ["netlens-mcp"]
    }
  }
}

Add either to your MCP client config (e.g. .mcp.json for Claude Code), then restart the session so the tools load.

Tools

web_search

Search the web and return real result links (title, URL, snippet), parsed locally — links, not summaries. Follow up with web_fetch to read a result.

Argument Type Description
query string (required) The search query
limit integer Optional cap; default returns the full first page (~10)
engine string auto (default), duckduckgo, bing, mojeek, searxng
page integer Result page, 1-based. SearXNG only
time_range string day, week, month, year. SearXNG only
categories string e.g. it, science, news. SearXNG only

The result reports which engine answered and what happened to any that were skipped, so falling through to a different backend is visible rather than silent:

{
  "query": "…",
  "engine": "mojeek",
  "results": [],
  "engines_skipped": [ { "engine": "duckduckgo", "outcome": "HTTP 202" } ]
}

Outcomes are observations, not conclusions — HTTP 202 is what the server sent; whether that is throttling, changed markup or genuinely no matches cannot be determined from the response. With SearXNG configured, direct answers, infoboxes and suggestions appear alongside the results.

A search fetches a single result page (~10 results), returned in full by default so nothing at position 9/10 is dropped. There's no deep pagination — if the answer isn't in the first page, refine the query.

web_fetch

Fetch any page and return its full content as clean Markdown.

Argument Type Description
url string (required) URL to fetch (scheme optional; https assumed)
mode string article (main content only, default), full (whole body), raw (unconverted HTML), outline (heading structure)
section string Return only this heading's content, plus anything nested under it
links string inline (default) keeps link targets; none keeps link text but drops URLs
max_chars integer Optional cap on returned characters (truncates with a note)

Reading part of a long page

A large article can be tens of thousands of characters when you want one part of it. mode="outline" returns its shape, and section returns just that piece — on a large encyclopedia article that is 153,000 characters full, 1,000 as an outline, and 7,400 for the section actually wanted.

web_fetch(url=…, mode="outline")      → headings with each section's size
web_fetch(url=…, section="Gameplay")  → that section and its subsections

On link-heavy pages the URLs themselves are a large share of the output — around 40% of a big encyclopedia article — so links="none" roughly halves it when you only need the prose. Links pointing back into the same page are always rendered as plain text.

If a page turns out to be a client-rendered shell, web_fetch says so rather than returning an empty result as a success:

(note: Only 10 characters of readable text were found in 6,856 characters of HTML. This page appears to be rendered client-side by JavaScript…)

Workflow: web_search to find pages, then web_fetch to read them.

Search engines

Search is a pluggable, selectable registry. In auto mode NetLens tries engines in order and returns the first with results, so a rate-limit/challenge page on one falls through to the next.

Engine Notes
duckduckgo Default; html.duckduckgo.com endpoint
bing Automatic fallback
mojeek Independent index; automatic fallback
searxng Self-hosted/public SearXNG JSON API — set NETLENS_SEARXNG_URL

Pick per call with the engine argument, or set a default with NETLENS_SEARCH_ENGINE.

Configuration

Environment Variable Default Description
NETLENS_SEARCH_ENGINE auto Default search backend
NETLENS_SEARXNG_URL SearXNG base URL, e.g. http://192.168.1.10:8888
NETLENS_SEARXNG_TIMEOUT 8 Seconds before an unreachable SearXNG is skipped
NETLENS_USER_AGENT Chrome UA Override the request User-Agent
NETLENS_MAX_BYTES 10485760 Cap on a single response; larger ones are truncated
NETLENS_CACHE_TTL 300 Seconds to reuse a fetched page; 0 disables caching
NETLENS_HOST_DELAY 0.5 Minimum seconds between requests to the same host
NETLENS_REQUEST_TIMEOUT 120 Ceiling on a single tool call

Self-hosted SearXNG

Point NETLENS_SEARXNG_URL at an instance with the JSON API enabled (search.formats must include json in its settings.yml). It is then tried first in auto mode, which removes the HTML scraping — and the rate limiting that comes with it — from the common path.

Because it is tried first, it must fail fast when the box is off: the connect timeout is bounded separately so an unreachable instance is skipped in a few seconds rather than stalling every search, and the skip is reported in the result.

How it works

  • Direct fetch. Requests go straight to the target site via the system curl (better TLS/HTTP-2/compression, so it looks like a real browser), falling back to urllib. No third-party proxy or reader is involved.
  • Local conversion. HTML → Markdown happens in-process with a hand-rolled html.parser converter — headings, lists, links (relative URLs resolved), code blocks, and GFM tables with colspan/rowspan.
  • Content selection, not deletion. NetLens picks the page's main content region — the HTML5 landmark (<main> / <article> / [role=main]) when one exists, otherwise the subtree holding the most prose relative to its link density — and converts only that. Because it selects a winner rather than deleting anything that matches a name pattern, extraction cannot silently return an empty page. Nothing inspects CSS class or id names to decide what is content.
  • Pruning by measurement. Within that region, blocks that are overwhelmingly link anchors (navboxes, tag clouds, "more from this site" grids) are dropped based on their link density. Non-rendering elements (<script>, <style>, …), explicitly hidden elements, and third-party ad-network slots (identified by vendor names like adsbygoogle, which cannot collide with real prose) are removed outright.
  • Response charset is honored (from Content-Type or <meta>), so non-UTF-8 pages don't come back garbled.

Usage from the CLI

The server is also a plain script — handy for testing before a client loads it:

python -m netlens_mcp.server search "http caching best practices"
python -m netlens_mcp.server fetch  https://example.com/article
python -m netlens_mcp.server full   https://example.com   # whole body
python -m netlens_mcp.server raw    https://example.com   # unconverted HTML

python -m netlens_mcp runs the stdio MCP server; python -m netlens_mcp.server <cmd> runs the CLI.

Development

pip install -e ".[dev]"
python -m pytest        # run the test suite
ruff check .            # lint

Requirements

  • Python 3.10+ (and the system curl, which ships with modern Windows/macOS/Linux; falls back to urllib if absent)

License

Apache License 2.0 — see LICENSE and NOTICE.

<!-- mcp-name: io.github.pzalutski-pixel/netlens -->

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选