arquivo-pt-mcp

arquivo-pt-mcp

An MCP server that enables Claude and other LLMs to search and access archived Portuguese web content from Arquivo.pt, including full-text search, image search, version listing, and snapshot retrieval.

Category
访问服务器

README


title: Arquivo Pt MCP emoji: 📚 colorFrom: blue colorTo: red sdk: docker app_port: 7860 pinned: false license: mit

arquivo-pt-mcp

Python Versions License: MIT CI

Um servidor Model Context Protocol (MCP) para o Arquivo.pt — o arquivo web português. Permite que o Claude (ou qualquer outro LLM compatível com MCP) pesquise e leia conteúdo web português arquivado.


🇵🇹 Em Português

O que faz

O arquivo-pt-mcp expõe seis ferramentas ao modelo de linguagem, permitindo-lhe consultar o Arquivo.pt como se fosse uma base de dados nativa:

Ferramenta Descrição
search Pesquisa em texto integral no arquivo, com filtros opcionais de intervalo de datas e site
image_search Pesquisa em mais de 1,8 mil milhões de imagens arquivadas
list_versions Lista todas as capturas de um determinado URL através do servidor CDX
get_snapshot Obtém uma página arquivada específica a partir de um URL + timestamp
extract_text Obtém uma página arquivada e devolve o texto legível (HTML removido)
get_screenshot Obtém o URL de uma captura PNG renderizada de uma página arquivada (opcionalmente com os bytes inline)

Instalação

pip install arquivo-pt-mcp

Ou, se preferir usar o uv:

uv add arquivo-pt-mcp

Para desenvolvimento (instalação a partir do código fonte):

git clone https://github.com/thaenor/arquivo-pt-mcp.git
cd arquivo-pt-mcp
pip install -e ".[dev]"

Configuração

🚀 Caminho mais fácil — sem instalação

Existe uma instância pública alojada em Hugging Face Spaces. Aponte qualquer cliente MCP para:

https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

Sem Python, sem pip install, sem terminal. Basta o URL. (Ver instalação local abaixo se preferir correr no seu próprio computador.)

Claude.ai (web ou desktop) — Pro/Max/Team

A configuração mais simples para utilizadores não-técnicos.

  1. Vá a Settings → Connectors → Add custom connector.
  2. Name: arquivo-pt
  3. URL: https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp
  4. Guarde. As seis ferramentas ficam disponíveis em qualquer nova conversa.

Claude Desktop

Edite o ficheiro claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%\Claude\claude_desktop_config.json
  • Linux: ~/.config/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "arquivo-pt": {
      "url": "https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp"
    }
  }
}

Reinicie o Claude Desktop.

Claude Code (CLI)

Um único comando:

claude mcp add --transport http arquivo-pt https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

Cursor

Settings → MCP → Add server:

  • Name: arquivo-pt
  • URL: https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

ChatGPT — Plus/Team/Enterprise

Settings → Connectors → Add → MCP server URL:

https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

Outros clientes MCP (Zed, Cline, Windsurf, Continue, …)

Qualquer cliente que suporte o transporte Streamable HTTP do MCP pode usar o mesmo URL acima. Consulte a documentação do seu cliente para saber onde o colar.


Instalação local

A instância pública é adequada para uso casual, mas é uma máquina partilhada gratuita sem garantia de disponibilidade. Corra o servidor localmente se precisar de privacidade, cache própria, ou disponibilidade garantida.

Como servidor stdio local (Claude Desktop, Cursor, etc.):

{
  "mcpServers": {
    "arquivo-pt": {
      "command": "uvx",
      "args": ["arquivo-pt-mcp"]
    }
  }
}

O uvx (do uv) instala o pacote automaticamente na primeira execução — só precisa de ter o uv instalado.

Como servidor HTTP de longa duração:

pip install arquivo-pt-mcp        # ou: uv add arquivo-pt-mcp
arquivo-pt-mcp --transport http --host 127.0.0.1 --port 8000

Depois aponte o seu cliente para http://127.0.0.1:8000/mcp. Para expor publicamente, coloque-o atrás de um proxy inverso com TLS (Caddy, nginx, Traefik) e passe --allowed-host <hostname>.

Nota: em modo HTTP as caches em memória são partilhadas entre todos os clientes ligados.

Exemplos de utilização

Uma vez configurado, pode pedir ao Claude coisas como:

  • “Pesquisa no Arquivo.pt por ‘eleições 2005’ e mostra-me os primeiros três resultados.”
  • “Mostra-me como a página inicial do Público era a 1 de janeiro de 2010.”
  • “Quantas vezes foi o Expresso arquivado em 2008?”
  • “Extrai o texto do snapshot mais antigo do sapo.pt.”
  • ”Procura imagens do Terreiro do Paço arquivadas antes de 2010.”
  • ”Mostra-me uma screenshot da homepage do Público em 1 de janeiro de 2010.”

Endpoints da API utilizados

  • https://arquivo.pt/textsearch — pesquisa em texto
  • https://arquivo.pt/imagesearch — pesquisa de imagens
  • https://arquivo.pt/wayback/cdx — índice de capturas (CDX)
  • https://arquivo.pt/wayback/{timestamp}/{url} — obtenção de snapshots
  • https://arquivo.pt/wayback/noFrame/{timestamp}/{url} — snapshot limpo para extração de texto
  • https://arquivo.pt/screenshot?url=... — captura PNG renderizada
  • https://arquivo.pt/noFrame/replay/{timestamp}/{url} — replay sem frame para screenshot

Documentação oficial: https://github.com/arquivo/pwa-technologies/wiki/Arquivo.pt-API

Desenvolvimento

git clone https://github.com/thaenor/arquivo-pt-mcp.git
cd arquivo-pt-mcp
uv sync --extra dev
pytest -q

O projeto usa:

  • pytest + pytest-asyncio para testes
  • pytest-cov para cobertura
  • ruff para lint e formatação
ruff check src tests
ruff format src tests
pytest --cov=arquivo_pt_mcp

Testes de integração

Opcionalmente, pode executar os testes de integração contra a API real do Arquivo.pt:

RUN_INTEGRATION=1 pytest -m integration -v

Estes testes estão marcados com @pytest.mark.integration e são ignorados por padrão. Apenas executam quando a variável de ambiente RUN_INTEGRATION=1 está definida. No GitHub Actions podem ser disparados manualmente via workflow_dispatch.

Nota sobre o GitHub Actions: Os runners padrão do GitHub estão alojados nos EUA. A conectividade TCP transatlântica para arquivo.pt (alojado em Portugal) é por vezes pouco fiável — os testes podem falhar com httpx.ConnectError ou httpx.ConnectTimeout independentemente da qualidade do código. Por isso, os testes de integração foram removidos do agendamento noturno no CI. Podem ser disparados manualmente via workflow_dispatch quando necessário. Para execuções locais (a partir de qualquer localização na Europa) os testes passam de forma consistente.

Prémio Arquivo.pt

Este projeto foi desenvolvido para participar no Prémio Arquivo.pt, que incentiva a criação de ferramentas e aplicações que aproveitam o arquivo web português para fins educativos, científicos, culturais e técnicos.

Roadmap

  • [ ] Guardar Página Agora (Save Page Now) — requer credenciais de API
  • [x] Cache com TTL para reduzir chamadas ao Arquivo.pt
  • [x] Gestão de rate limits com retentativas exponenciais
  • [x] Validação de inputs com modelos Pydantic
  • [x] Suporte para pesquisa avançada por domínio e coleção
  • [x] Integração com outros clientes MCP (Zed, Cline, Windsurf)

Licença

MIT


🇬🇧 In English

What it does

arquivo-pt-mcp exposes six tools to the language model, letting it query Arquivo.pt as if it were a native data source:

Tool Description
search Full-text search across the archive, with optional date range and site filters
image_search Search across 1.8B+ archived images
list_versions Every capture of a given URL, via the CDX server
get_snapshot Resolve a URL + timestamp to a specific archived page
extract_text Fetch an archived page and return its readable text (HTML stripped)
get_screenshot Get the PNG render URL of an archived page (optionally embed the bytes inline)

Installation

pip install arquivo-pt-mcp

Or with uv:

uv add arquivo-pt-mcp

For development (install from source):

git clone https://github.com/thaenor/arquivo-pt-mcp.git
cd arquivo-pt-mcp
pip install -e ".[dev]"

Configuration

🚀 Easiest path — no install required

A public hosted instance runs on Hugging Face Spaces. Point any MCP client at:

https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

No Python, no pip install, no terminal. Just the URL. (See self-hosted setup below if you'd rather run it locally.)

Claude.ai (web or desktop) — Pro/Max/Team

The simplest setup for non-technical users.

  1. Go to Settings → Connectors → Add custom connector.
  2. Name: arquivo-pt
  3. URL: https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp
  4. Save. The six tools are now available in any new chat.

Claude Desktop

Edit your claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%\Claude\claude_desktop_config.json
  • Linux: ~/.config/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "arquivo-pt": {
      "url": "https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp"
    }
  }
}

Restart Claude Desktop.

Claude Code (CLI)

One command:

claude mcp add --transport http arquivo-pt https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

Cursor

Settings → MCP → Add server:

  • Name: arquivo-pt
  • URL: https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

ChatGPT — Plus/Team/Enterprise

Settings → Connectors → Add → MCP server URL:

https://decaf-squirrel-arquivo-pt-mcp.hf.space/mcp

Other MCP clients (Zed, Cline, Windsurf, Continue, …)

Any client that speaks the MCP Streamable HTTP transport can use the same URL above. Refer to your client's docs for where to paste it.


Self-hosted setup

The hosted instance is fine for casual use, but it's a free shared CPU box with no SLA. Run it yourself if you need privacy, your own caching, or guaranteed availability.

As a local stdio server (Claude Desktop, Cursor, etc.):

{
  "mcpServers": {
    "arquivo-pt": {
      "command": "uvx",
      "args": ["arquivo-pt-mcp"]
    }
  }
}

uvx (from uv) auto-installs the package on first run — nothing to install manually beyond uv itself.

As a long-lived HTTP server:

pip install arquivo-pt-mcp        # or: uv add arquivo-pt-mcp
arquivo-pt-mcp --transport http --host 127.0.0.1 --port 8000

Then point your client at http://127.0.0.1:8000/mcp. To expose publicly, put it behind a TLS-terminating reverse proxy (Caddy, nginx, Traefik) and pass --allowed-host <hostname>.

Note: in HTTP mode the in-memory caches are shared across all connected clients.

Usage examples

Once connected, you can ask Claude things like:

  • “Search Arquivo.pt for ‘eleições 2005’ and show me the first three results.”
  • “Show me how publico.pt’s homepage looked on January 1, 2010.”
  • “How many times was expresso.pt archived in 2008?”
  • “Extract the text from the earliest snapshot of sapo.pt.”
  • ”Search for archived images of Terreiro do Paço from before 2010.”
  • ”Show me a screenshot of Público's homepage on Jan 1, 2010.”

API endpoints used

  • https://arquivo.pt/textsearch — text search
  • https://arquivo.pt/imagesearch — image search
  • https://arquivo.pt/wayback/cdx — capture index (CDX)
  • https://arquivo.pt/wayback/{timestamp}/{url} — snapshot retrieval
  • https://arquivo.pt/wayback/noFrame/{timestamp}/{url} — clean snapshot for text extraction
  • https://arquivo.pt/screenshot?url=... — PNG screenshot render
  • https://arquivo.pt/noFrame/replay/{timestamp}/{url} — frameless replay for screenshot

Official docs: https://github.com/arquivo/pwa-technologies/wiki/Arquivo.pt-API

Development

git clone https://github.com/thaenor/arquivo-pt-mcp.git
cd arquivo-pt-mcp
uv sync --extra dev
pytest -q

The project uses:

  • pytest + pytest-asyncio for testing
  • pytest-cov for coverage
  • ruff for linting and formatting
ruff check src tests
ruff format src tests
pytest --cov=arquivo_pt_mcp

Integration tests

Optionally, run integration tests against the live Arquivo.pt API:

RUN_INTEGRATION=1 pytest -m integration -v

These tests are marked with @pytest.mark.integration and skipped by default. They only run when the RUN_INTEGRATION=1 environment variable is set. On GitHub Actions they can be triggered manually via workflow_dispatch.

GitHub Actions note: Standard GitHub-hosted runners are US-based. Transatlantic TCP connectivity to arquivo.pt (hosted in Portugal) is sometimes unreliable — tests may fail with httpx.ConnectError or httpx.ConnectTimeout regardless of code quality. For this reason, the integration tests have been removed from the nightly CI schedule. They can still be triggered manually via workflow_dispatch when needed. When run locally from a European location the tests pass consistently.

Prémio Arquivo.pt

This project was built for the Prémio Arquivo.pt, a Portuguese contest that encourages the creation of tools and applications leveraging the Portuguese Web Archive for educational, scientific, cultural, and technical purposes.

Roadmap

  • [ ] Save Page Now — requires API credentials
  • [x] TTL caching to reduce calls to Arquivo.pt
  • [x] Rate-limit handling with exponential backoff
  • [x] Input validation with Pydantic models
  • [x] Advanced search by domain and collection
  • [x] Integration with additional MCP clients (Zed, Cline, Windsurf)

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选