wienerberliner-pi-arxiv

内容来源:README.md(说明文档) · 原始地址 · 查看安装指南

原始内容

pi-arxiv

One Pi extension for arXiv discovery, exact lookup, and full-paper Markdown fetching.

Tools

  • arxiv_search — search arXiv by query/category with sorting and pagination.
  • arxiv_paper — exact metadata lookup by arXiv ID or URL.
  • arxiv_fetch2md — fetch a paper body as Markdown via arxiv2md and save it locally.

arxiv_fetch2md uses arxiv2md instead of PDF scraping. arxiv2md parses arXiv's structured HTML when available, which generally preserves sections and math better than PDF extraction.

Library folder setup

Fetched Markdown papers are saved to a local library folder.

On first fetch, if no folder is configured, the extension searches for likely paper/library folders such as ~/Documents/Papers, ~/Papers, ~/Research, ~/Documents/Arxiv, and Obsidian vaults. In interactive Pi it asks which one to use, offers a recommended ~/Documents/Arxiv folder when none exists, and saves the choice in:

~/.pi/agent/pi-arxiv/config.json

You can reconfigure any time:

/arxiv-library

For non-interactive runs, set:

export PI_ARXIV_LIBRARY="$HOME/Documents/Arxiv"

or let the extension create/use ~/Documents/Arxiv.

Install

After publishing:

pi install npm:@wienerberliner/pi-arxiv

Install while developing

From this package directory:

npm install
pi -e .

Or add this local package to project/user Pi settings.

Credits

This extension was inspired by the Pi arXiv plugins that came before it:

Notes

  • arXiv API calls are throttled with a 3 second delay between requests, following arXiv's guidance for repeated API calls.
  • arxiv_fetch2md depends on arxiv2md's public API and its rate limits.
  • PDF-only arXiv papers may not have structured HTML; in that case arxiv2md may fail and a dedicated PDF extraction fallback would be needed.