Files
fundi-scraper-api/api/README.md
T
ebbe.bassandClaude Sonnet 5 eaed7d4ce6 Add archive scraping, webcal feed, and scrape schedule reporting
- scraper.py: --archive flag to scrape https://fundbureau.de/archiv.html
  (past events), fixing a container-specific wait selector and a crash
  on events with no door time that this surfaced.
- api: EventStore now merges an optional archive JSON file into the
  main event list (deduped by date+name); adds GET /events/archive.
- api: GET /calendar.ics serves an RFC 5545 feed of all events for
  webcal subscriptions, with an all-day fallback when no door time is
  parseable.
- api: GET /health reports last-scrape time (from file mtime) and, via
  new SCRAPE_CRON/ARCHIVE_SCRAPE_CRON env vars, the next scheduled run
  and seconds until it, for both the regular and archive scrape.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 13:42:51 +02:00

2.7 KiB

Fundi Scraper API

A small REST API (FastAPI) that serves the events scraped by scraper/scraper.py from a JSON file.

Setup

cd api
pip install -r requirements.txt

Run

cd api
uvicorn main:app --reload

Interactive docs: http://127.0.0.1:8000/docs

By default the API reads ../fundi-scraped-output.json (upcoming events) and ../fundi-archive-output.json (past events, optional — produced by scraper.py --archive), merging both into one in-memory list. Point either elsewhere with an env var:

DATA_FILE=/path/to/events.json ARCHIVE_DATA_FILE=/path/to/archive.json uvicorn main:app

Scrape schedule reporting

/health reports when each scraper output file was last written (its mtime) and, if you tell it the cron schedule, when it's next due. This doesn't read your crontab — set the same expression(s) you put there as env vars:

SCRAPE_CRON="0 * * * *" ARCHIVE_SCRAPE_CRON="0 4 * * *" uvicorn main:app

SCRAPE_CRON covers the plain scrape (fundi-scraped-output.json); ARCHIVE_SCRAPE_CRON covers --archive (fundi-archive-output.json) and falls back to SCRAPE_CRON if unset — set it separately only if the archive scrape runs on its own cron line. Leave both unset to just get the last-scrape info with next_scrape_at / seconds_until_next_scrape as null.

Endpoints

Method Path Description
GET /health Status, event counts, and last/next scrape timing (see below)
GET /events List events, with optional filters (see below)
GET /events/archive Past events only — shorthand for /events?upcoming=false
GET /events/{index} Single event by its position in the merged list (0-based)
GET /artists Deduplicated, sorted list of all artist names
GET /calendar.ics iCalendar/webcal feed of all events — subscribe with webcal://<host>/calendar.ics
POST /reload Re-read the data file(s) from disk (after a fresh scrape)

/events query parameters

Param Type Meaning
free bool Only free / only paid events
name string Case-insensitive substring match on the event name
artist string Case-insensitive substring match on any artist name
date string Exact match on the raw date string (dd.mm.yy)
upcoming bool true = today or later, false = past events

Examples:

curl 'http://127.0.0.1:8000/events?free=true'
curl 'http://127.0.0.1:8000/events?artist=randali&upcoming=true'
curl 'http://127.0.0.1:8000/events/0'